Multi-scale neural distribution prediction and hierarchical migration early warning method

By employing a multi-scale neural distribution prediction and hierarchical migration early warning method, and utilizing neural stochastic differential equations and hierarchical Bayesian statistical learning, multi-scale carbon emission path samples are generated and early warning thresholds are calibrated. This solves the problem of the inherent correlation mechanism between complex stochastic dynamics and multi-scale temporal coupling in the carbon emission system, and achieves highly reliable carbon emission distribution prediction and early warning.

CN121936643APending Publication Date: 2026-04-28HUBEI UNIV OF ECONOMICS +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HUBEI UNIV OF ECONOMICS
Filing Date
2025-11-20
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

Existing carbon emission prediction and early warning methods are unable to effectively characterize the complex stochastic dynamics and multi-scale temporal coupling relationships of carbon emission systems, resulting in the inability to achieve highly reliable distribution prediction and adaptive adjustment of early warning thresholds. In particular, they lack effective cross-condition information sharing mechanisms and statistical learning frameworks in scenarios with small sample conditions and condition switching.

Method used

A multi-scale neural distribution prediction and hierarchical migration early warning method is adopted. Through neural stochastic differential equation model and hierarchical Bayesian statistical learning method, a multi-scale carbon emission path sample set is generated. The path-level adaptability score is calculated by combining the exceedance risk integral, the first exceedance time and the path fluctuation variance. The early warning threshold of small sample operating conditions is regularized and calibrated by operating condition stratification mechanism and hierarchical Bayesian parameter sharing.

Benefits of technology

It improves the accuracy of carbon emission distribution prediction and the robustness of the early warning system, enhances the generalization ability under operating condition switching scenarios, effectively alleviates the problem of unstable early warning threshold under small sample operating conditions, and improves the adaptive adjustment capability of the early warning system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121936643A_ABST
    Figure CN121936643A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of carbon emission prediction and early warning, and provides a multi-scale neural distribution prediction and hierarchical migration early warning method, which comprises the following steps: acquiring historical carbon emission data, respectively inputting a historical sequence and a to-be-predicted sequence into an energy consumption stochastic differential equation model and a carbon factor stochastic differential equation model, generating a multi-scale carbon emission path sample set through an independent random disturbance term; calculating a path-level suitability score of each path sample based on the standard-exceeding risk integral, the first standard-exceeding moment and the path fluctuation variance; layering the calibration data set into a plurality of working condition layers according to working condition labels, sharing distribution shape parameters among the working condition layers through a hierarchical Bayesian method, and regularizing quantiles of small sample working condition layers to obtain an early warning threshold value of each working condition layer; and selecting a corresponding early warning threshold value according to the current working condition label to compare and trigger early warning. According to the method, the accuracy of carbon emission distribution prediction and the robustness of an early warning system are improved, and the problem that the early warning threshold value is unstable under the small sample working condition is relieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of carbon emission prediction and early warning technology, and in particular to a multi-scale neural distribution prediction and hierarchical migration early warning method. Background Technology

[0002] Carbon emission forecasting and early warning are key technological supports for green industrial parks, industrial enterprises, and power systems to achieve carbon neutrality goals, and are widely used in important fields such as energy management, energy conservation and emission reduction, and carbon asset management. Carbon emission data exhibits typical multi-timescale characteristics, operating condition dependence, and stochastic fluctuations. In actual forecasting, it is affected by multiple factors such as changes in energy consumption, dynamic adjustments of carbon factors, and operating condition switching, resulting in complex nonlinear dynamic characteristics. Traditional carbon emission forecasting methods, as the main technical means of carbon emission management, predict future carbon emission levels and conduct early warning analysis by establishing a mapping relationship between energy consumption and carbon factors. However, the unique multi-scale time dynamics, operating condition heterogeneity, and small sample operating conditions of carbon emission systems bring many challenges to distributed forecasting and intelligent early warning. The key lies in how to effectively characterize the stochastic dynamic process of carbon emissions, accurately model multi-timescale coupling relationships, and achieve intelligent early warning threshold calibration across operating conditions.

[0003] In existing technologies, carbon emission prediction and early warning mainly employ traditional time series analysis and basic machine learning methods to achieve basic carbon emission prediction functions. However, existing methods do not adequately consider the inherent physical correlation mechanism between the complex stochastic dynamics and multi-scale time coupling of carbon emission systems. They struggle to organically integrate the objective physical constraints of stochastic processes with the actual dynamic characteristics of operating conditions, resulting in an inability to achieve highly reliable carbon emission distribution prediction and adaptive adjustment of early warning thresholds. Particularly when facing complex scenarios such as small sample operating conditions, operating condition switching, and distribution drift, existing methods lack effective cross-operating condition information sharing mechanisms and statistical learning frameworks, making it difficult to guarantee the robustness and generalization ability of the early warning system. Summary of the Invention

[0004] In view of this, the present invention proposes a multi-scale neural distribution prediction and hierarchical migration early warning method, which solves the problem that existing methods do not adequately consider the inherent physical correlation mechanism between the complex stochastic dynamics and multi-scale temporal coupling of carbon emission systems, and are unable to organically integrate the objective physical constraints of stochastic processes with the actual dynamic characteristics of operating conditions, resulting in the inability to achieve highly reliable carbon emission distribution prediction and adaptive adjustment of early warning thresholds.

[0005] The technical solution of this invention is implemented as follows: This invention provides a method for multi-scale neural distribution prediction and hierarchical migration early warning, comprising the following steps: Acquire historical carbon emission data, which includes historical energy consumption sequences, historical carbon factor sequences, and operating condition labels; Historical energy consumption sequences and historical carbon factor sequences are input into the energy consumption stochastic differential equation model and the carbon factor stochastic differential equation model, respectively. Energy consumption paths and carbon factor paths are generated through independent random perturbation terms, and a multi-scale carbon emission path sample set is calculated. The energy consumption sequence and carbon factor sequence to be predicted are input into the energy consumption stochastic differential equation model and the carbon factor stochastic differential equation model, respectively, to generate a multi-scale carbon emission path sample set for future periods. Based on the exceedance risk score, the first exceedance time, and the path fluctuation variance, the path-level fit score of each path sample in the multi-scale carbon emission path sample set is calculated. The calibration dataset is divided into multiple working condition layers according to the working condition label. The quantile of the path-level adaptability score is calculated in each working condition layer. The distribution shape parameter is shared between working condition layers through the hierarchical Bayesian method, and the quantile of the small sample working condition layer is regularized to obtain the warning threshold of each working condition layer. Select the corresponding warning threshold based on the current working condition label, compare the path-level adaptability score with the warning threshold, and trigger a warning when the score is exceeded.

[0006] Based on the above technical solutions, preferably, the step of inputting historical energy consumption sequences and historical carbon factor sequences into the energy consumption stochastic differential equation model and the carbon factor stochastic differential equation model, respectively, generating energy consumption paths and carbon factor paths through independent random perturbation terms, and calculating a multi-scale carbon emission path sample set, including: The historical energy consumption sequence is input into the encoder network, which is used to extract the temporal features of the historical energy consumption sequence, convert the operating condition label into an operating condition embedding vector, and concatenate the temporal features with the operating condition embedding vector to obtain the energy consumption condition features. Construct an energy consumption drift network and an energy consumption diffusion network, input the energy consumption condition features into the energy consumption drift network and the energy consumption diffusion network, the energy consumption drift network outputs the energy consumption drift coefficient, the energy consumption diffusion network outputs the energy consumption diffusion coefficient, and based on the energy consumption drift coefficient, the energy consumption diffusion coefficient and the first independent random perturbation term, generate multiple energy consumption path samples through a numerical solver; The historical carbon factor sequence is decomposed at multiple time scales to obtain carbon factor scale components at each time scale. Fast fluctuation features and slow trend features are extracted for each carbon factor scale component. The fast fluctuation features and the slow trend features are input into a carbon factor drift network and a carbon factor diffusion network. The carbon factor drift network outputs a carbon factor drift coefficient, and the carbon factor diffusion network outputs a carbon factor diffusion coefficient. Based on the carbon factor drift coefficient, the carbon factor diffusion coefficient, and a second independent random perturbation term, multiple carbon factor path samples are generated by a numerical solver. Each energy consumption path sample is multiplied point by point with the corresponding carbon factor path sample at each time point to obtain the multi-scale carbon emission path sample set.

[0007] Based on the above technical solutions, preferably, the step of decomposing the historical carbon factor sequence across multiple time scales includes: The historical carbon factor sequence is aggregated using a sliding window at a first time scale, a second time scale, and a third time scale to obtain a first-scale carbon factor sequence, a second-scale carbon factor sequence, and a third-scale carbon factor sequence. The first-scale carbon factor sequence, the second-scale carbon factor sequence, and the third-scale carbon factor sequence constitute the carbon factor scale components at each time scale.

[0008] Based on the above technical solutions, preferably, the extraction of rapid fluctuation features and slow trend features for each carbon factor scale component includes: High-pass filtering is applied to the scale components of each carbon factor to extract high-frequency components and obtain rapid fluctuation characteristics. Low-pass filtering is applied to the scale components of each carbon factor to extract low-frequency components and obtain slow trend characteristics. The step of inputting the rapid fluctuation feature and the slow trend feature into the carbon factor drift network includes: The rapid fluctuation features corresponding to the current time scale are weighted and fused with the slow trend features corresponding to other time scales to obtain cross-scale coupling features. The cross-scale coupling features are then input into the carbon factor drift network, with higher weights assigned to the rapid fluctuation features of shorter time scales and higher weights assigned to the slow trend features of longer time scales.

[0009] Based on the above technical solutions, preferably, the calculation formulas for the energy consumption drift coefficient and the energy consumption diffusion coefficient are as follows: ; ; in, for and Energy drift coefficient; for and The energy dissipation diffusion coefficient; For a moment Energy path value; This is the energy consumption condition feature vector; For reference energy consumption levels; The mean recovery strength parameter; This is a vector of linear influence coefficients for the operating conditions; This refers to the nonlinear interaction strength parameter; Map weight vectors to working condition features; Basic diffusion intensity; This is the horizontal dependence coefficient; The modulation coefficient is the operating condition coefficient. Let L2 be the eigenvector of the energy consumption condition; For activation functions; It is an exponential function.

[0010] Based on the above technical solutions, preferably, the calibration dataset is divided into multiple working condition layers according to the working condition label. Within each working condition layer, the quantiles of the path-level adaptability score are calculated. A hierarchical Bayesian method is used to share distribution shape parameters between working condition layers, and the quantiles of small sample working condition layers are regularized to obtain the warning thresholds for each working condition layer. Obtain a calibration dataset, which includes historical path-level adaptability scores and corresponding historical operating condition labels. Based on the historical operating condition labels, the calibration dataset is divided into multiple operating condition layers according to operating condition type. Each operating condition layer includes historical path-level adaptability scores corresponding to the same operating condition label. The historical path-level adaptability scores are sorted within each operating condition layer. The quantile value corresponding to a preset quantile level is calculated as the initial threshold of the operating condition layer. The number of samples in each operating condition layer is counted. Based on the number of samples, small sample operating condition layers and large sample operating condition layers are identified. A hierarchical Bayesian model is constructed, and the historical path-level fit scores of each working condition layer are used as observation data and input into the hierarchical Bayesian model. The hierarchical Bayesian model predicts the shared distribution shape parameters among the large sample working condition layers. Based on the shared distribution shape parameters, Bayesian regularization is applied to the quantiles of the small sample working condition layers. The quantile values ​​of each working condition layer are updated by maximum a posteriori estimation, and the regularized quantiles are used as the warning thresholds for each working condition layer.

[0011] Based on the above technical solutions, preferably, the construction of the hierarchical Bayesian model includes: A distribution shape parameter is set as a hierarchical parameter, the distribution shape parameter including shape hyperparameter and scale hyperparameter, and a super-prior distribution is set for the shape hyperparameter and the scale hyperparameter respectively; The step of inputting the historical path-level adaptability scores of each working condition layer as observation data into the hierarchical Bayesian model includes: The historical path-level adaptability score of each working condition layer follows a parameterized distribution. The distribution parameters of the parameterized distribution are jointly determined by the local parameters of the corresponding working condition layer and the shared distribution shape parameters. The observation likelihood function of the working condition layer is constructed, and the shape hyperparameter, the scale hyperparameter, and the local parameters of each working condition layer are used as model parameters.

[0012] Based on the above technical solutions, preferably, the step of performing Bayesian regularization on the quantiles of the small sample working condition layer based on the shared distribution shape parameter includes: The hierarchical Bayesian model is inferred using the Markov chain Monte Carlo method. The posterior distribution of the shared distribution shape parameter is estimated from the observation data of the large-sample working condition layer. The posterior distribution of the shared distribution shape parameter is passed as prior information to the small-sample working condition layer. The regularized local parameters are calculated by combining the limited observation data of the small-sample working condition layer. Based on the regularized local parameters and the shared distribution shape parameter, the regularized quantile value of the small-sample working condition layer at the preset quantile level is calculated.

[0013] Based on the above technical solutions, preferably, the calculation of the path-level fit score for each path sample in the multi-scale carbon emission path sample set based on the exceedance risk score, the first exceedance time, and the path fluctuation variance includes: A warning threshold baseline is set, and the time values ​​of each path sample in the multi-scale carbon emission path sample set are compared with the warning threshold baseline to count the set of times when each path sample exceeds the standard. The excess duration percentage is calculated based on the set of excess times to obtain the excess risk score. The earliest time in the set of excess times is extracted to obtain the first excess time. The standard deviation of the values ​​at each time for each path sample is calculated to obtain the path fluctuation variance. The excess risk score, the first excess time, and the path fluctuation variance are combined to obtain the path risk index combination corresponding to each path sample. A path adaptability scoring function is constructed. The path risk indicators corresponding to each path sample are combined and input into the path adaptability scoring function. Weight coefficients are set for the excess risk score, the first excess time and the path fluctuation variance, respectively. The path-level adaptability score of each path sample is obtained by weighted summation.

[0014] Based on the above technical solutions, preferably, the step of selecting a corresponding early warning threshold according to the current working condition label, comparing the path-level adaptability score with the early warning threshold, and triggering an early warning when the threshold is exceeded includes: Obtain the current working condition label as the current working condition label, match the current working condition label with the working condition type of each working condition layer, and when the current working condition label completely matches the working condition type of the first working condition layer, select the warning threshold corresponding to the first working condition layer as the current warning threshold. When the current working condition label cannot be completely matched with any working condition layer, the similarity between the current working condition label and the working condition type of each working condition layer is calculated, the warning threshold corresponding to the working condition layer with the highest similarity is selected as the current warning threshold, and it is marked as an approximate matching state. Obtain the path-level adaptability score corresponding to the current time, compare the path-level adaptability score with the current warning threshold, and generate a warning signal when the path-level adaptability score is greater than the current warning threshold. The warning signal includes the warning time, current operating condition label, path-level adaptability score, current warning threshold and matching status, and output the warning signal as the warning result.

[0015] The multi-scale neural distribution prediction and hierarchical migration early warning method of the present invention has the following advantages over the prior art: (1) By using the neural stochastic differential equation model and hierarchical Bayesian statistical learning method, the stochastic dynamic process of energy consumption and carbon factor is characterized by independent stochastic perturbation terms, and a multi-scale carbon emission path sample set is generated. The path-level adaptability score is calculated by combining the three-dimensional indexes of the risk integral of exceeding the standard, the first time of exceeding the standard and the path fluctuation variance. The warning threshold of small sample working conditions is regularized and calibrated by the working condition stratification mechanism and hierarchical Bayesian parameter sharing, which improves the accuracy of carbon emission distribution prediction and the robustness of the warning system. At the same time, the problem of unstable warning threshold under small sample working conditions is effectively alleviated by the cross-working condition information sharing mechanism, which enhances the generalization ability of the warning system in the working condition switching scenario. (2) By integrating the energy consumption stochastic differential equation modeling with the carbon factor stochastic differential equation modeling with the conditional working conditions, the energy consumption time series features are extracted by the encoder network and concatenated with the working condition embedding vector to form conditional features. The rapid fluctuation features and slow trend features of carbon factor are extracted by combining multi-scale sliding window aggregation and high-pass and low-pass filtering. The coupling modeling between different time scales is realized through the cross-scale weighted fusion mechanism, which improves the accuracy of carbon emission path sample generation and the ability to depict dynamics at multiple scales. At the same time, the independence of energy consumption and carbon factor random fluctuations is ensured by the independent random perturbation term mechanism, which enhances the reliability of distribution prediction. (3) By integrating the working condition stratification mechanism with the hierarchical Bayesian statistical learning framework, the Markov chain Monte Carlo method is used to predict the posterior distribution of the shared distribution shape parameter from the large sample working condition layer. The hierarchical parameterized model is constructed by combining the super-prior distribution of shape hyperparameter and scale hyperparameter. By passing the posterior distribution of the large sample working condition as prior information to the small sample working condition layer for Bayesian regularization, the robustness and statistical reliability of the warning threshold setting are improved. At the same time, the hierarchical parameter sharing mechanism effectively alleviates the threshold overfitting problem in the data imbalance scenario, and enhances the generalization ability and adaptive adjustment ability of the warning system in new and rare working conditions. Attached Figure Description

[0016] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0017] Figure 1 This is a flowchart of a multi-scale neural distribution prediction and hierarchical migration early warning method according to the present invention. Detailed Implementation

[0018] The technical solutions of the present invention will be clearly and completely described below with reference to the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.

[0019] Please see Figure 1 This invention provides a method for multi-scale neural distribution prediction and hierarchical migration early warning, comprising the following steps: Acquire historical carbon emission data, which includes historical energy consumption sequences, historical carbon factor sequences, and operating condition labels; Historical energy consumption sequences and historical carbon factor sequences are input into the energy consumption stochastic differential equation model and the carbon factor stochastic differential equation model, respectively. Energy consumption paths and carbon factor paths are generated through independent random perturbation terms, and a multi-scale carbon emission path sample set is calculated. The energy consumption sequence and carbon factor sequence to be predicted are input into the energy consumption stochastic differential equation model and the carbon factor stochastic differential equation model, respectively, to generate a multi-scale carbon emission path sample set for future periods. Based on the exceedance risk score, the first exceedance time, and the path fluctuation variance, the path-level fit score of each path sample in the multi-scale carbon emission path sample set is calculated. The calibration dataset is divided into multiple working condition layers according to the working condition label. The quantile of the path-level adaptability score is calculated in each working condition layer. The distribution shape parameter is shared between working condition layers through the hierarchical Bayesian method, and the quantile of the small sample working condition layer is regularized to obtain the warning threshold of each working condition layer. Select the corresponding warning threshold based on the current working condition label, compare the path-level adaptability score with the warning threshold, and trigger a warning when the score is exceeded.

[0020] Specifically, this embodiment utilizes a neural stochastic differential equation model and hierarchical Bayesian statistical learning method to characterize the stochastic dynamic processes of energy consumption and carbon factors using independent stochastic perturbation terms, generating a multi-scale carbon emission path sample set. It then calculates the path-level adaptability score by combining three-dimensional indices: exceedance risk integral, first exceedance time, and path fluctuation variance. Furthermore, it regularizes and calibrates the warning threshold for small-sample operating conditions through a hierarchical operating condition mechanism and hierarchical Bayesian parameter sharing. This embodiment solves the problems of complex stochastic dynamic modeling of carbon emission systems, multi-timescale coupled representation, and adaptive setting of intelligent warning thresholds under heterogeneous operating conditions. It improves the accuracy of carbon emission distribution prediction and the robustness of the warning system. Simultaneously, the cross-operating condition information sharing mechanism effectively alleviates the problem of unstable warning thresholds under small-sample operating conditions, enhancing the generalization ability of the warning system in operating condition switching scenarios.

[0021] The acquisition of historical carbon emission data includes historical energy consumption sequences, historical carbon factor sequences, and operating condition labels, including: Raw energy consumption data and raw carbon factor data are collected from multiple equipment monitoring points and the power grid dispatching platform. The raw energy consumption data is aggregated at multiple time scales, and the raw carbon factor data is time-aligned at the same time scale to obtain aligned energy consumption sequences and aligned carbon factor sequences. Abnormal data points in the aligned energy consumption sequences and aligned carbon factor sequences are detected, and abnormal data points are marked and interpolated for repair to obtain historical energy consumption sequences and historical carbon factor sequences.

[0022] In one specific embodiment, the aggregation of the raw energy consumption data across multiple time scales includes: The raw energy consumption data is grouped according to multiple equipment monitoring points. The raw energy consumption data of each equipment monitoring point is summed according to the first time scale, the second time scale, and the third time scale to obtain the equipment energy consumption subsequence of each equipment monitoring point at each time scale. The aligned energy consumption sequence is obtained by summing the energy consumption subsequences of all equipment monitoring points at the same time scale.

[0023] In one specific embodiment, the step of time-aligning the original carbon factor data on the same time scale includes: The timestamp sequence of the original carbon factor data is obtained, and the timestamp sequence is matched with the timestamp sequence of the aligned energy consumption sequence. For the unmatched time points, linear interpolation is used to generate carbon factor values, thus obtaining the aligned carbon factor sequence.

[0024] Time features are extracted based on the timestamps of the historical energy consumption sequence. The time features include hour identifiers, weekday identifiers, and holiday identifiers. Load statistical features of the historical energy consumption sequence within a preset time window are calculated. The load statistical features include mean, variance, and peak-to-valley difference. The time features and the load statistical features are input into the operating condition classifier to obtain the operating condition label.

[0025] In one specific embodiment, inputting the time features and the load statistics features into the operating condition classifier includes: Construct a working condition feature vector, which includes the time feature and the load statistical feature; The operating condition feature vector is input into a decision tree classifier. The decision tree classifier performs a first-level classification based on the hour identifier and the week identifier, and a second-level classification based on the mean and peak-valley difference of the load statistical features. The decision tree classifier outputs a work condition category identifier, which includes day shift, night shift, weekend, and holiday categories. The consistency of the operating condition category identifiers within a continuous time period is checked. When the operating condition category identifier at a single time point is inconsistent with the operating condition category identifiers at the preceding and following time points, a majority voting rule is used to correct it, thus obtaining the operating condition label.

[0026] Specifically, this embodiment integrates multi-source heterogeneous data acquisition with a multi-timescale aggregation mechanism. It utilizes a linear interpolation method combining device-level energy consumption summation and timestamp matching to achieve time alignment between energy consumption and carbon factor sequences. This is combined with anomaly point marking and repair, extraction of time features and load statistical features, and a decision tree classifier for operating condition category identification and consistency verification using majority voting rules. This embodiment solves the challenges of multi-device monitoring point data fusion, data alignment across different time scales, and automatic operating condition identification, improving the quality of historical carbon emission data and the accuracy of operating condition labels. Furthermore, the consistency verification mechanism across consecutive time periods effectively eliminates temporal jumps in operating condition labels, enhancing the temporal stability of the dataset.

[0027] The process involves inputting historical energy consumption sequences and historical carbon factor sequences into the energy consumption stochastic differential equation model and the carbon factor stochastic differential equation model, respectively. Energy consumption paths and carbon factor paths are generated through independent random perturbation terms, and a multi-scale carbon emission path sample set is calculated, including: The historical energy consumption sequence is input into the encoder network, which is used to extract the temporal features of the historical energy consumption sequence, convert the operating condition label into an operating condition embedding vector, and concatenate the temporal features with the operating condition embedding vector to obtain the energy consumption condition features. Construct an energy consumption drift network and an energy consumption diffusion network, input the energy consumption condition features into the energy consumption drift network and the energy consumption diffusion network, the energy consumption drift network outputs the energy consumption drift coefficient, the energy consumption diffusion network outputs the energy consumption diffusion coefficient, and based on the energy consumption drift coefficient, the energy consumption diffusion coefficient and the first independent random perturbation term, generate multiple energy consumption path samples through a numerical solver; The historical carbon factor sequence is decomposed at multiple time scales to obtain carbon factor scale components at each time scale. Fast fluctuation features and slow trend features are extracted for each carbon factor scale component. The fast fluctuation features and the slow trend features are input into a carbon factor drift network and a carbon factor diffusion network. The carbon factor drift network outputs a carbon factor drift coefficient, and the carbon factor diffusion network outputs a carbon factor diffusion coefficient. Based on the carbon factor drift coefficient, the carbon factor diffusion coefficient, and a second independent random perturbation term, multiple carbon factor path samples are generated by a numerical solver. Each energy consumption path sample is multiplied point by point with the corresponding carbon factor path sample at each time point to obtain the multi-scale carbon emission path sample set.

[0028] The step of decomposing the historical carbon factor sequence across multiple time scales includes: The historical carbon factor sequence is aggregated using a sliding window at a first time scale, a second time scale, and a third time scale to obtain a first-scale carbon factor sequence, a second-scale carbon factor sequence, and a third-scale carbon factor sequence. The first-scale carbon factor sequence, the second-scale carbon factor sequence, and the third-scale carbon factor sequence constitute the carbon factor scale components at each time scale.

[0029] The extraction of rapid fluctuation features and slow trend features for each carbon factor scale component includes: High-pass filtering is applied to the scale components of each carbon factor to extract high-frequency components and obtain rapid fluctuation characteristics. Low-pass filtering is applied to the scale components of each carbon factor to extract low-frequency components and obtain slow trend characteristics.

[0030] The step of inputting the rapid fluctuation feature and the slow trend feature into the carbon factor drift network includes: The rapid fluctuation features corresponding to the current time scale are weighted and fused with the slow trend features corresponding to other time scales to obtain cross-scale coupling features. The cross-scale coupling features are then input into the carbon factor drift network, with higher weights assigned to the rapid fluctuation features of shorter time scales and higher weights assigned to the slow trend features of longer time scales.

[0031] In one specific embodiment, the calculation formula for the cross-scale coupling feature is: ; in, For the first Cross-scale coupling characteristics across time scales; For the first Rapid fluctuations over time; For the first Slow trend characteristics over a time scale; For the first The nonlinear modulation intensity coefficient on the time scale, with a value range of: ; For the first The basic weighting coefficient for the time scale is a positive number; This is a scale-distance sensitivity parameter, and it is a positive number. For the first Time scale and the first The volatility correlation coefficient over time scales, with values ​​ranging from... ; For the first Standard deviation of slow trend characteristics over a time scale; For the first Standard deviation of rapid fluctuation characteristics over a time scale; This is the upper limit of the variance ratio pruning threshold; This is the first numerically stable term to prevent division by zero; It is the hyperbolic tangent activation function.

[0032] The formulas for calculating the energy consumption drift coefficient and the energy consumption diffusion coefficient are as follows: ; ; in, for and The energy drift coefficient; for and The energy dissipation diffusion coefficient; For a moment Energy path value; This is the energy consumption condition feature vector; As a reference energy consumption level, it is a positive number; The mean recovery strength parameter; This is a vector of linear influence coefficients for the operating conditions; This refers to the nonlinear interaction strength parameter; Map weight vectors to working condition features; The basic diffusion intensity is a positive number; This is the horizontal dependence coefficient, and it is a non-negative number. The modulation coefficient is the operating condition coefficient. Let L2 be the eigenvector of the energy consumption condition; For activation functions; It is an exponential function.

[0033] Specifically, this embodiment integrates energy consumption stochastic differential equation modeling with operating condition-conditional modeling and carbon factor stochastic differential equation modeling with multi-timescale decomposition. It utilizes an encoder network to extract time-series energy consumption features and concatenates them with the operating condition embedding vector to form conditional features. Multi-scale sliding window aggregation and high-pass / low-pass filtering are combined to extract rapid fluctuation and slow trend features of the carbon factor. A cross-scale weighted fusion mechanism achieves coupled modeling across different time scales. The independent random perturbation term and drift-diffusion network address the issues of joint dynamic modeling of energy consumption and carbon factor under heterogeneous operating conditions and coupled multi-timescale representation, improving the accuracy of carbon emission path sample generation and multi-scale dynamic characterization capabilities. Simultaneously, the independent random perturbation term mechanism ensures the independence of random fluctuations in energy consumption and carbon factor, enhancing the reliability of distribution prediction.

[0034] The step of inputting the energy consumption sequence to be predicted and the carbon factor sequence to be predicted into the energy consumption stochastic differential equation model and the carbon factor stochastic differential equation model, respectively, to generate a multi-scale carbon emission path sample set for future periods includes: The historical observation window data for the prediction start time is obtained. The historical observation window data includes historical observation energy consumption subsequence and historical observation carbon factor subsequence. The historical observation energy consumption subsequence is input into the encoder of the energy consumption stochastic differential equation model to extract the initial energy consumption hidden state. The historical observation carbon factor subsequence is decomposed at multiple time scales to obtain the initial carbon factor scale state. The starting condition label corresponding to the prediction start time is obtained and the starting condition label is converted into the starting condition embedding vector.

[0035] In one specific embodiment, obtaining the historical observation window data for the prediction start time includes: Using the prediction start time as the cutoff point, backtracking by a preset observation window length, the energy consumption sequence to be predicted within this time period is extracted as the historical observation energy consumption subsequence, and the carbon factor sequence to be predicted within this time period is extracted as the historical observation carbon factor subsequence.

[0036] In one specific embodiment, the step of inputting the historical observed energy consumption subsequence into the encoder of the energy consumption stochastic differential equation model to extract the initial energy consumption hidden state includes: The historical energy consumption subsequence is input into the encoder network of the trained energy consumption stochastic differential equation model. The encoder network performs feature extraction and temporal encoding on the historical energy consumption subsequence and outputs the hidden state vector of the encoder at the last moment as the initial energy consumption hidden state.

[0037] In one specific embodiment, the step of decomposing the historical observed carbon factor subsequence across multiple time scales to obtain the initial carbon factor scale state includes: The historical carbon factor subsequences are aggregated using a sliding window at multiple time scales to obtain initial carbon factor scale sequences at each time scale. The last time value of each initial carbon factor scale sequence is extracted as the initial carbon factor scale state.

[0038] Based on the initial energy consumption hidden state and the initial operating condition embedding vector, the predicted energy consumption drift coefficient and the predicted energy consumption diffusion coefficient are generated through the drift network and diffusion network of the energy consumption stochastic differential equation model, respectively. Multiple independent random seeds are set, and a first predicted random perturbation term is generated based on each independent random seed. Multiple energy consumption paths to be predicted are generated according to the predicted energy consumption drift coefficient, the predicted energy consumption diffusion coefficient and each first predicted random perturbation term. Based on the initial carbon factor scale state, the predicted carbon factor drift coefficient and the predicted carbon factor diffusion coefficient are generated through the drift network and diffusion network of the carbon factor stochastic differential equation model, respectively. A second predicted random perturbation term is generated based on each random seed. Multiple carbon factor paths to be predicted are generated according to the predicted carbon factor drift coefficient, the predicted carbon factor diffusion coefficient and each second predicted random perturbation term. Each predicted energy consumption path is multiplied point by point with its corresponding predicted carbon factor path to obtain a multi-scale carbon emission path sample set for the future time period.

[0039] In one specific embodiment, setting multiple independent random seeds includes: Based on the preset number of path samples, multiple distinct random seed values ​​are generated, with each random seed value corresponding to an independent prediction path.

[0040] In one specific embodiment, generating a first predicted random perturbation term and a second predicted random perturbation term based on each random seed includes: For each random seed value, a first random number generator and a second random number generator are initialized respectively. The first random number generator generates a random sequence that follows a standard normal distribution as the first predicted random perturbation term, and the second random number generator generates a random sequence that follows a standard normal distribution as the second predicted random perturbation term. The first predicted random perturbation term and the second predicted random perturbation term are independent of each other.

[0041] In one specific embodiment, generating multiple energy consumption paths to be predicted based on the predicted energy consumption drift coefficient, the predicted energy consumption diffusion coefficient, and each of the first predicted random disturbance terms includes: For the first predicted random perturbation term corresponding to each random seed, numerical integration is performed using the Euler-Maruyama method or the Milstein method. Starting from the initial energy consumption hidden state, a complete energy consumption path to be predicted is generated by progressively advancing according to the predicted energy consumption drift coefficient and the predicted energy consumption diffusion coefficient. Multiple random seeds generate multiple energy consumption paths to be predicted.

[0042] Specifically, this embodiment integrates historical observation window data extraction with a stochastic differential equation model recursion mechanism. It utilizes a trained encoder network to extract the initial energy consumption hidden state and multi-scale decomposition to extract the initial carbon factor scale state. Independent random seeds are used to initialize the first and second random number generators respectively, ensuring the independence of energy consumption and carbon factor perturbation terms. Multiple independent prediction paths are then generated step-by-step through numerical integration using the Euler-Maruyama method or the Milstein method. This embodiment addresses the issues of quantifying future carbon emission uncertainty and generating diverse path samples, improving the accuracy of distribution prediction and the diversity coverage of prediction paths. Furthermore, the parallel generation mechanism of multiple independent random seeds effectively constructs a complete prediction distribution sample space.

[0043] The method calculates the path-level fit score for each path sample in the multi-scale carbon emission path sample set based on the exceedance risk score, the time of the first exceedance, and the path fluctuation variance, including: A warning threshold baseline is set, and the time values ​​of each path sample in the multi-scale carbon emission path sample set are compared with the warning threshold baseline to count the set of times when each path sample exceeds the standard. The excess duration percentage is calculated based on the set of excess times to obtain the excess risk score. The earliest time in the set of excess times is extracted to obtain the first excess time. The standard deviation of the values ​​at each time for each path sample is calculated to obtain the path fluctuation variance. The excess risk score, the first excess time, and the path fluctuation variance are combined to obtain the path risk index combination corresponding to each path sample.

[0044] In one specific embodiment, comparing the time-varying values ​​of each path sample in the multi-scale carbon emission path sample set with the warning threshold baseline includes: Obtain the statistical characteristics of historical carbon emission data, calculate the mean and standard deviation of historical carbon emission data, and set the warning threshold baseline based on the mean and standard deviation; Iterate through the time values ​​of each path sample, compare each time value with the warning threshold baseline, and mark the time value as exceeding the warning threshold baseline. Collect all the time values ​​exceeding the threshold to form the set of time values ​​exceeding the threshold.

[0045] In one specific embodiment, calculating the percentage of time exceeding the standard based on the set of times exceeding the standard as the risk score for exceeding the standard includes: Calculate the number of times exceeding the standard in the set of times exceeding the standard, divide the number by the total number of times in the path samples, and obtain the percentage of time exceeding the standard as the risk score of exceeding the standard.

[0046] In one specific embodiment, calculating the standard deviation of the values ​​of each path sample at each time point as the path fluctuation variance includes: Calculate the arithmetic mean of the values ​​at each time step for each path sample, calculate the sum of squared deviations between the values ​​at each time step and the arithmetic mean, divide the sum of squared deviations by the number of time steps minus one, and take the square root to obtain the path fluctuation variance.

[0047] A path adaptability scoring function is constructed. The path risk indicators corresponding to each path sample are combined and input into the path adaptability scoring function. Weight coefficients are set for the excess risk score, the first excess time and the path fluctuation variance, respectively. The path-level adaptability score of each path sample is obtained by weighted summation.

[0048] In one specific embodiment, the construction path adaptability scoring function includes: The excess risk weight, time weight, and variance weight are set as the weight coefficients, wherein the excess risk weight corresponds to the excess risk integral, the time weight corresponds to the first excess time, and the variance weight corresponds to the path fluctuation variance.

[0049] In one specific embodiment, the step of inputting the combined path risk indicators corresponding to each path sample into the path suitability scoring function includes: The time reciprocal transformation is performed on the first time exceeding the standard. When there is no time exceeding the standard in the path sample, the first time exceeding the standard is set as a preset maximum time value, and the reciprocal of the preset maximum time value is calculated as the time reciprocal component. When there is a time exceeding the standard in the path sample, the reciprocal of the first time exceeding the standard is calculated as the time reciprocal component. The excess risk integral, the reciprocal component of the time, and the path fluctuation variance are normalized and multiplied by the corresponding excess risk weight, the time weight, and the variance weight, respectively. The weighted results are then summed to obtain the path-level adaptability score.

[0050] Specifically, this embodiment integrates a three-dimensional risk index—including excess risk score, first excess time, and path fluctuation variance—with a weighted scoring mechanism. It quantifies path excess characteristics using a baseline comparison of warning thresholds and statistical methods for excess time sets. It extracts path risk features by combining the calculation of excess duration percentage, time reciprocal transformation, and standard deviation calculation. Finally, it constructs a comprehensive path suitability scoring function through normalization and weighted summation. This embodiment solves the problems of multi-dimensional risk quantification of carbon emission path samples and comprehensive risk assessment of heterogeneous paths, improving the comprehensiveness and discriminative power of path risk assessment. Furthermore, by employing a preset maximum time value processing mechanism for paths without excess emissions, it effectively avoids scoring failure in extreme cases, enhancing the early warning system's ability to identify paths with different risk patterns and the reliability of risk ranking.

[0051] The calibration dataset is divided into multiple working condition layers according to the working condition label. Within each working condition layer, the quantiles of the path-level adaptability score are calculated. A hierarchical Bayesian method is used to share distribution shape parameters between working condition layers, and the quantiles of small sample working condition layers are regularized to obtain the warning thresholds for each working condition layer. Obtain a calibration dataset, which includes historical path-level adaptability scores and corresponding historical operating condition labels. Based on the historical operating condition labels, the calibration dataset is divided into multiple operating condition layers according to operating condition type. Each operating condition layer includes historical path-level adaptability scores corresponding to the same operating condition label. The historical path-level adaptability scores are sorted within each operating condition layer. The quantile value corresponding to a preset quantile level is calculated as the initial threshold of the operating condition layer. The number of samples in each operating condition layer is counted. Based on the number of samples, small sample operating condition layers and large sample operating condition layers are identified. A hierarchical Bayesian model is constructed, and the historical path-level fit scores of each working condition layer are used as observation data and input into the hierarchical Bayesian model. The hierarchical Bayesian model predicts the shared distribution shape parameters among the large sample working condition layers. Based on the shared distribution shape parameters, Bayesian regularization is applied to the quantiles of the small sample working condition layers. The quantile values ​​of each working condition layer are updated by maximum a posteriori estimation, and the regularized quantiles are used as the warning thresholds for each working condition layer.

[0052] The construction of the hierarchical Bayesian model includes: The distribution shape parameter is set as a hierarchical parameter, which includes shape hyperparameter and scale hyperparameter. A super-prior distribution is set for the shape hyperparameter and the scale hyperparameter respectively.

[0053] The step of inputting the historical path-level adaptability scores of each working condition layer as observation data into the hierarchical Bayesian model includes: The historical path-level adaptability score of each working condition layer follows a parameterized distribution. The distribution parameters of the parameterized distribution are jointly determined by the local parameters of the corresponding working condition layer and the shared distribution shape parameters. The observation likelihood function of the working condition layer is constructed, and the shape hyperparameter, the scale hyperparameter, and the local parameters of each working condition layer are used as model parameters.

[0054] The Bayesian regularization of the quantiles of the small sample working condition layer based on the shared distribution shape parameter includes: The hierarchical Bayesian model is inferred using the Markov chain Monte Carlo method. The posterior distribution of the shared distribution shape parameter is estimated from the observation data of the large-sample working condition layer. The posterior distribution of the shared distribution shape parameter is passed as prior information to the small-sample working condition layer. The regularized local parameters are calculated by combining the limited observation data of the small-sample working condition layer. Based on the regularized local parameters and the shared distribution shape parameter, the regularized quantile value of the small-sample working condition layer at the preset quantile level is calculated.

[0055] In one specific embodiment, the formula for calculating the regularized quantile value is: ; ; in, For the first Regularized quantile values ​​for the working condition layer; For the first Empirical quantile values ​​for the working condition layer; For the first Migration correction amount for the working condition layer; These are the prior quantiles calculated based on the shared distribution shape parameters; Shape hyperparameters shared globally; These are scale hyperparameters that are shared globally. It is the basic regularization strength parameter, and it is a positive number; For the first Number of samples in the working condition layer; For the first KL divergence between local distribution and global prior distribution in the working condition layer; This is the tolerance parameter for distributional variability, and it is a positive number. It is an exponential function.

[0056] Specifically, this embodiment integrates a working condition stratification mechanism with a hierarchical Bayesian statistical learning framework. It uses the Markov chain Monte Carlo method to predict the posterior distribution of shared shape parameters from a large-sample working condition layer. A hierarchical parameterized model is constructed by combining the prior distributions of shape and scale hyperparameters. The posterior distribution of the large-sample working conditions is then passed as prior information to the small-sample working condition layer for Bayesian regularization. This embodiment addresses the issues of unstable early warning thresholds for small-sample working conditions and cross-working condition information sharing under heterogeneous working conditions, improving the robustness and statistical reliability of early warning threshold setting. Simultaneously, the hierarchical parameter sharing mechanism effectively alleviates the threshold overfitting problem in imbalanced data scenarios, enhancing the generalization and adaptive adjustment capabilities of the early warning system under new and rare working conditions.

[0057] The step of selecting the corresponding early warning threshold based on the current working condition label, comparing the path-level adaptability score with the early warning threshold, and triggering an early warning when the threshold is exceeded includes: Obtain the current working condition label as the current working condition label, match the current working condition label with the working condition type of each working condition layer, and when the current working condition label completely matches the working condition type of the first working condition layer, select the warning threshold corresponding to the first working condition layer as the current warning threshold. When the current working condition label cannot be completely matched with any working condition layer, the similarity between the current working condition label and the working condition type of each working condition layer is calculated, the warning threshold corresponding to the working condition layer with the highest similarity is selected as the current warning threshold, and it is marked as an approximate matching state.

[0058] In one specific embodiment, matching the current operating condition label with the operating condition type of each operating condition layer includes: Extract the first working condition feature vector of the current working condition label, extract the second working condition feature vector of each working condition layer working condition type, and calculate the Euclidean distance between the first working condition feature vector and the second working condition feature vector as the matching distance.

[0059] In one specific embodiment, the step of selecting the warning threshold corresponding to the first working condition layer when the current working condition label completely matches the working condition type of the first working condition layer includes: A complete match threshold is set. When the matching distance is less than the complete match threshold, it is determined that the current working condition label is completely matched with the working condition type of the first working condition layer. The warning threshold is then selected from the successfully matched working condition layers.

[0060] In one specific embodiment, calculating the similarity between the current working condition label and the working condition types of each working condition layer includes: The matching distance is converted into a similarity score, which is inversely proportional to the matching distance. The similarity score of all working condition layers is calculated, and the working condition layer with the highest similarity score is selected. The warning threshold corresponding to the working condition layer with the highest similarity score is extracted as the current warning threshold, and the similarity score is recorded as the matching confidence.

[0061] Obtain the path-level adaptability score corresponding to the current time, compare the path-level adaptability score with the current warning threshold, and generate a warning signal when the path-level adaptability score is greater than the current warning threshold. The warning signal includes the warning time, current operating condition label, path-level adaptability score, current warning threshold and matching status, and output the warning signal as the warning result.

[0062] In one specific embodiment, comparing the path-level adaptability score with the current warning threshold includes: A set of warning level threshold combinations is set, which includes a mild warning threshold, a moderate warning threshold, and a severe warning threshold. The mild warning threshold is equal to the current warning threshold, the moderate warning threshold is equal to the current warning threshold multiplied by a moderate coefficient, and the severe warning threshold is equal to the current warning threshold multiplied by a severe coefficient. The warning level is determined based on the comparison between the path-level adaptability score and the thresholds for each level. When the path-level adaptability score is between the mild warning threshold and the moderate warning threshold, it is set to a mild warning level; when the path-level adaptability score is between the moderate warning threshold and the severe warning threshold, it is set to a moderate warning level; when the path-level adaptability score exceeds the severe warning threshold, it is set to a severe warning level.

[0063] In one specific embodiment, generating the warning signal includes: A warning information structure is constructed, which encapsulates the warning time, the current working condition label, the path-level adaptability score, the current warning threshold, the matching status, and the warning level into structured warning information. The proportion by which the path-level adaptability score exceeds the current warning threshold is calculated as a risk intensity index. The risk intensity index is added to the structured warning information, and the structured warning information is output as the warning result.

[0064] Specifically, this embodiment integrates a dual mechanism of precise and approximate matching of operating conditions and a multi-level early warning threshold system. It uses Euclidean distance to calculate the matching degree of operating condition feature vectors and converts it into a similarity score for operating condition layer selection. It combines a three-level early warning threshold system (mild, moderate, and severe) with risk intensity index calculation, and encapsulates multi-dimensional information such as early warning time, operating condition label, adaptability score, matching status, and early warning level through structured early warning information. This embodiment solves the problems of adaptive threshold selection in scenarios with dynamic switching of operating conditions, robust handling of unknown operating conditions, and refined risk level classification. It improves the accuracy of early warning triggering and the precision of risk assessment. Simultaneously, through the approximate matching mechanism and matching confidence recording, it effectively addresses new and rare operating condition scenarios, enhancing the generalization ability and traceability of the early warning system in complex and changing operating conditions.

[0065] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A method for multi-scale neural distribution prediction and hierarchical migration early warning, characterized in that, Includes the following steps: Acquire historical carbon emission data, which includes historical energy consumption sequences, historical carbon factor sequences, and operating condition labels; Historical energy consumption sequences and historical carbon factor sequences are input into the energy consumption stochastic differential equation model and the carbon factor stochastic differential equation model, respectively. Energy consumption paths and carbon factor paths are generated through independent random perturbation terms, and a multi-scale carbon emission path sample set is calculated. The energy consumption sequence and carbon factor sequence to be predicted are input into the energy consumption stochastic differential equation model and the carbon factor stochastic differential equation model, respectively, to generate a multi-scale carbon emission path sample set for future periods. Based on the exceedance risk score, the first exceedance time, and the path fluctuation variance, the path-level fit score of each path sample in the multi-scale carbon emission path sample set is calculated. The calibration dataset is divided into multiple working condition layers according to the working condition label. The quantile of the path-level adaptability score is calculated in each working condition layer. The distribution shape parameter is shared between working condition layers through the hierarchical Bayesian method, and the quantile of the small sample working condition layer is regularized to obtain the warning threshold of each working condition layer. Select the corresponding warning threshold based on the current working condition label, compare the path-level adaptability score with the warning threshold, and trigger a warning when the score is exceeded.

2. The method for multi-scale neural distribution prediction and hierarchical migration early warning as described in claim 1, characterized in that, The process involves inputting historical energy consumption sequences and historical carbon factor sequences into the energy consumption stochastic differential equation model and the carbon factor stochastic differential equation model, respectively. Energy consumption paths and carbon factor paths are generated through independent random perturbation terms, and a multi-scale carbon emission path sample set is calculated, including: The historical energy consumption sequence is input into the encoder network, which is used to extract the temporal features of the historical energy consumption sequence, convert the operating condition label into an operating condition embedding vector, and concatenate the temporal features with the operating condition embedding vector to obtain the energy consumption condition features. Construct an energy consumption drift network and an energy consumption diffusion network, input the energy consumption condition features into the energy consumption drift network and the energy consumption diffusion network, the energy consumption drift network outputs the energy consumption drift coefficient, the energy consumption diffusion network outputs the energy consumption diffusion coefficient, and based on the energy consumption drift coefficient, the energy consumption diffusion coefficient and the first independent random perturbation term, generate multiple energy consumption path samples through a numerical solver; The historical carbon factor sequence is decomposed at multiple time scales to obtain carbon factor scale components at each time scale. Fast fluctuation features and slow trend features are extracted for each carbon factor scale component. The fast fluctuation features and the slow trend features are input into a carbon factor drift network and a carbon factor diffusion network. The carbon factor drift network outputs a carbon factor drift coefficient, and the carbon factor diffusion network outputs a carbon factor diffusion coefficient. Based on the carbon factor drift coefficient, the carbon factor diffusion coefficient, and a second independent random perturbation term, multiple carbon factor path samples are generated by a numerical solver. Each energy consumption path sample is multiplied point by point with the corresponding carbon factor path sample at each time point to obtain the multi-scale carbon emission path sample set.

3. The method for multi-scale neural distribution prediction and hierarchical migration early warning as described in claim 2, characterized in that, The step of decomposing the historical carbon factor sequence across multiple time scales includes: The historical carbon factor sequence is aggregated using a sliding window at a first time scale, a second time scale, and a third time scale to obtain a first-scale carbon factor sequence, a second-scale carbon factor sequence, and a third-scale carbon factor sequence. The first-scale carbon factor sequence, the second-scale carbon factor sequence, and the third-scale carbon factor sequence constitute the carbon factor scale components at each time scale.

4. The method for multi-scale neural distribution prediction and hierarchical migration early warning as described in claim 3, characterized in that, The extraction of rapid fluctuation features and slow trend features for each carbon factor scale component includes: High-pass filtering is applied to the scale components of each carbon factor to extract high-frequency components and obtain rapid fluctuation characteristics. Low-pass filtering is applied to the scale components of each carbon factor to extract low-frequency components and obtain slow trend characteristics. The step of inputting the rapid fluctuation characteristics and the slow trend characteristics into the carbon factor drift network includes: The rapid fluctuation features corresponding to the current time scale are weighted and fused with the slow trend features corresponding to other time scales to obtain cross-scale coupling features. The cross-scale coupling features are then input into the carbon factor drift network, with higher weights assigned to the rapid fluctuation features of shorter time scales and higher weights assigned to the slow trend features of longer time scales.

5. The method for multi-scale neural distribution prediction and hierarchical migration early warning as described in claim 4, characterized in that, The formulas for calculating the energy consumption drift coefficient and the energy consumption diffusion coefficient are as follows: ; ; in, for and The energy drift coefficient; for and The energy dissipation diffusion coefficient; For a moment Energy path value; This is the energy consumption condition feature vector; For reference energy consumption levels; The mean recovery strength parameter; This is a vector of linear influence coefficients for the operating conditions; This refers to the nonlinear interaction strength parameter; Map weight vectors to working condition features; Basic diffusion intensity; This is the horizontal dependence coefficient; The modulation coefficient is the operating condition coefficient. Let L2 be the eigenvector of the energy consumption condition; For activation functions; It is an exponential function.

6. The method for multi-scale neural distribution prediction and hierarchical migration early warning as described in claim 1, characterized in that, The calibration dataset is divided into multiple working condition layers according to the working condition label. Within each working condition layer, the quantiles of the path-level adaptability score are calculated. A hierarchical Bayesian method is used to share distribution shape parameters between working condition layers, and the quantiles of small sample working condition layers are regularized to obtain the warning thresholds for each working condition layer. Obtain a calibration dataset, which includes historical path-level adaptability scores and corresponding historical operating condition labels. Based on the historical operating condition labels, the calibration dataset is divided into multiple operating condition layers according to operating condition type. Each operating condition layer includes historical path-level adaptability scores corresponding to the same operating condition label. The historical path-level adaptability scores are sorted within each operating condition layer. The quantile value corresponding to a preset quantile level is calculated as the initial threshold of the operating condition layer. The number of samples in each operating condition layer is counted. Based on the number of samples, small sample operating condition layers and large sample operating condition layers are identified. A hierarchical Bayesian model is constructed, and the historical path-level fit scores of each working condition layer are used as observation data and input into the hierarchical Bayesian model. The hierarchical Bayesian model predicts the shared distribution shape parameters among the large sample working condition layers. Based on the shared distribution shape parameters, Bayesian regularization is applied to the quantiles of the small sample working condition layers. The quantile values ​​of each working condition layer are updated by maximum a posteriori estimation, and the regularized quantiles are used as the warning thresholds for each working condition layer.

7. The method for multi-scale neural distribution prediction and hierarchical migration early warning as described in claim 6, characterized in that, The construction of the hierarchical Bayesian model includes: A distribution shape parameter is set as a hierarchical parameter, the distribution shape parameter including shape hyperparameter and scale hyperparameter, and a super-prior distribution is set for the shape hyperparameter and the scale hyperparameter respectively; The step of inputting the historical path-level adaptability scores of each working condition layer as observation data into the hierarchical Bayesian model includes: The historical path-level adaptability score of each working condition layer follows a parameterized distribution. The distribution parameters of the parameterized distribution are jointly determined by the local parameters of the corresponding working condition layer and the shared distribution shape parameters. The observation likelihood function of the working condition layer is constructed, and the shape hyperparameter, the scale hyperparameter, and the local parameters of each working condition layer are used as model parameters.

8. The method for multi-scale neural distribution prediction and hierarchical migration early warning as described in claim 7, characterized in that, The Bayesian regularization of the quantiles of the small sample working condition layer based on the shared distribution shape parameter includes: The hierarchical Bayesian model is inferred using the Markov chain Monte Carlo method. The posterior distribution of the shared distribution shape parameter is estimated from the observation data of the large-sample working condition layer. The posterior distribution of the shared distribution shape parameter is passed as prior information to the small-sample working condition layer. The regularized local parameters are calculated by combining the limited observation data of the small-sample working condition layer. Based on the regularized local parameters and the shared distribution shape parameter, the regularized quantile value of the small-sample working condition layer at the preset quantile level is calculated.

9. The method for multi-scale neural distribution prediction and hierarchical migration early warning as described in claim 1, characterized in that, The method calculates the path-level fit score for each path sample in the multi-scale carbon emission path sample set based on the exceedance risk score, the time of the first exceedance, and the path fluctuation variance, including: A warning threshold baseline is set, and the time values ​​of each path sample in the multi-scale carbon emission path sample set are compared with the warning threshold baseline to count the set of times when each path sample exceeds the standard. The excess duration percentage is calculated based on the set of excess times to obtain the excess risk score. The earliest time in the set of excess times is extracted to obtain the first excess time. The standard deviation of the values ​​at each time for each path sample is calculated to obtain the path fluctuation variance. The excess risk score, the first excess time, and the path fluctuation variance are combined to obtain the path risk index combination corresponding to each path sample. A path adaptability scoring function is constructed. The path risk indicators corresponding to each path sample are combined and input into the path adaptability scoring function. Weight coefficients are set for the excess risk score, the first excess time and the path fluctuation variance, respectively. The path-level adaptability score of each path sample is obtained by weighted summation.

10. The method for multi-scale neural distribution prediction and hierarchical migration early warning as described in claim 1, characterized in that, The step of selecting the corresponding early warning threshold based on the current working condition label, comparing the path-level adaptability score with the early warning threshold, and triggering an early warning when the threshold is exceeded includes: Obtain the current working condition label as the current working condition label, match the current working condition label with the working condition type of each working condition layer, and when the current working condition label completely matches the working condition type of the first working condition layer, select the warning threshold corresponding to the first working condition layer as the current warning threshold. When the current working condition label cannot be completely matched with any working condition layer, the similarity between the current working condition label and the working condition type of each working condition layer is calculated, the warning threshold corresponding to the working condition layer with the highest similarity is selected as the current warning threshold, and it is marked as an approximate matching state. Obtain the path-level adaptability score corresponding to the current time, compare the path-level adaptability score with the current warning threshold, and generate a warning signal when the path-level adaptability score is greater than the current warning threshold. The warning signal includes the warning time, current operating condition label, path-level adaptability score, current warning threshold and matching status, and output the warning signal as the warning result.