Photovoltaic power ultra-short-term probability prediction method based on dynamic multi-scale permutation entropy and improved xLSTM

Through dynamic multi-scale arrangement entropy and improved xLSTM model, the problem of insufficient multi-scale feature extraction and environmental variable correlation in photovoltaic power prediction is solved, the adaptability to mutation scenarios is enhanced, and high-precision probability prediction and grid scheduling support is achieved.

CN120449120APending Publication Date: 2025-08-08JIANGSU OCEAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510505266.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-22
Publication Date
2025-08-08

AI Technical Summary

Technical Problem

The existing photovoltaic power prediction models are insufficient in multi-scale feature extraction, insufficient dynamic correlation of environmental variables, insufficient adaptability to mutation scenarios, limited probability output capabilities, and difficult to deploy in real-time on edge devices.

Method used

Dynamic multi-scale arrangement entropy extraction features are adopted, combined with the improved xLSTM model, through the environment-aware gating layer, quantile regression layer and lightweight deployment layer, the historical memory weight is dynamically adjusted, and the confidence interval is output to achieve high-precision probability prediction.

Benefits of technology

It realizes accurate extraction of multi-scale features, enhances its response ability to environmental mutations, provides high coverage probability prediction, supports grid scheduling risk pre-control, has in-depth multi-scale feature mining, sensitive environmental mutation response, strong timing modeling ability, high probability prediction reliability, and strong model adaptability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120449120A_ABST
    Figure CN120449120A_ABST
Patent Text Reader

Abstract

The invention discloses a photovoltaic power ultra-short-term probability prediction method based on a dynamic multi-scale permutation entropy and an improved xLSTM, and the method comprises the following steps: 1, carrying out the preprocessing of multi-source heterogeneous data related to photovoltaic prediction, and carrying out the dynamic multi-scale feature extraction; 2, the features in the step 1 are fused, and xLSTM modeling is improved; and step 3, carrying out reverse normalization on the prediction result in the step 2, and then evaluating the prediction result. According to the method, the multi-scale complexity features are accurately extracted through the dynamic multi-scale permutation entropy, the coupling problem of meteorological sequence local fluctuation and global trend in a traditional entropy method is solved, and the problem that association between the multi-scale features and environment variables is ignored is solved; through an improved xLSTM gating correction mechanism, the problem that a single model lags in response to sudden weather change is solved, and historical memory weight is dynamically adjusted to enhance adaptability; the confidence interval is directly output through the quantile regression layer, the problem that traditional deterministic prediction cannot quantify uncertainty is solved, high-coverage probability prediction is provided, and strategy balance precision and calculation efficiency are dynamically integrated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of new energy prediction technology for power systems, and in particular to a photovoltaic power ultra-short-term probability prediction method based on dynamic multi-scale permutation entropy and improved xLSTM. Background Art

[0002] The new energy sector has achieved remarkable growth, injecting new momentum into economic development, contributing to energy security, and offering new solutions for the global energy transition. As a key breakthrough in energy structural transformation, the share of photovoltaic power generation continues to rise, effectively reducing reliance on fossil fuels and generating green momentum for high-quality economic development. However, photovoltaic power output is highly sensitive to meteorological factors, with output power exhibiting significant fluctuations and intermittentity, influenced by factors such as light intensity and atmospheric transparency. This output characteristic creates technical bottlenecks in power forecasting, presenting new challenges for distribution network risk assessment and control, and posing a potential threat to the safe and stable operation of the power system. To address these challenges, it is necessary to develop a dynamic analysis system for photovoltaic output and a high-precision forecasting algorithm. This technological breakthrough will significantly enhance the intelligent level of grid dispatching, optimize the efficiency of energy resource allocation, and provide key technical support for the construction of a new power system. It is of great practical significance for ensuring national energy security and ensuring the steady development of new energy.

[0003] Due to differences in climatic conditions and topographical features, the output characteristics of photovoltaic power generation systems in different geographical regions are significantly scenario-specific. Prediction models that fail to fully account for regional power generation characteristics will lead to systematic deviations between predicted results and actual operating data. To improve prediction accuracy, it is necessary to construct a feature space that incorporates multiple meteorological factors. Variables such as atmospheric temperature, relative humidity, wind speed and direction have nonlinear effects on photovoltaic output, and complex relationships exist between these variables. Directly using raw meteorological data introduces a large amount of redundant information, reducing the model's generalization ability. Multivariate correlation analysis and feature screening techniques can be used to extract key meteorological factors that strongly correlate with photovoltaic output. Combining physical mechanism modeling with machine learning algorithms, a regionalized meteorological-power mapping model can effectively characterize the dynamic relationship between meteorological factors and power generation capacity.

[0004] Machine learning algorithms are being deeply integrated into the field of photovoltaic power generation prediction. In deep learning architectures, long short-term memory networks (LSTMs) have become a research hotspot due to their ability to capture temporal features. This model effectively processes long-term dependencies of meteorological data such as light intensity and ambient temperature through gated recurrent units, and can achieve a prediction accuracy of 92.5% in typical scenarios. However, a single prediction model is difficult to adapt to complex and changing prediction scenarios, lacks adaptability to sudden weather changes, and does not consider the need for probabilistic prediction. In terms of feature extraction, traditional entropy methods (such as permutation entropy) do not combine multi-scale analysis with dynamic correlations with environmental variables, making it difficult to capture complex fluctuation patterns. Therefore, there is an urgent need for a photovoltaic power prediction method that combines high precision, strong robustness, and real-time performance. Summary of the Invention

[0005] Purpose of the invention: The present invention provides an ultra-short-term probabilistic prediction method for photovoltaic power based on dynamic multi-scale permutation entropy and improved xLSTM, which is used to solve the problems of insufficient multi-scale feature extraction of photovoltaic power series, insufficient modeling of dynamic correlation between environmental variables (temperature, humidity) and power, limited adaptability and probability output capability of the prediction model to sudden change scenarios, and difficulty in real-time deployment of complex models on edge devices.

[0006] Technical solution: The present invention provides a method for ultra-short-term probability prediction of photovoltaic power based on dynamic multi-scale permutation entropy and improved xLSTM, comprising the following steps:

[0007] Step 1: Preprocess the multi-source heterogeneous data involved in photovoltaic prediction and perform dynamic multi-scale feature extraction;

[0008] Step 2: Fuse the features in step 1 and improve the xLSTM modeling;

[0009] Step 3: Denormalize the prediction results in step 2 and then evaluate the prediction results.

[0010] Furthermore, in step 1, preprocessing the multi-source heterogeneous data involved in photovoltaic prediction specifically includes the following steps:

[0011] Step 11: Use data normalization method to preprocess the data:

[0012]

[0013] Where, X norm is the normalized data, x is the original data, x min is the minimum value in the original data, x max is the maximum value in the original data;

[0014] Step 12: After normalization, a sliding window is used to construct a data set. The input window uses the past 24 hours (t-23 to t), and the output window outputs the power values of the next six hours (t+1 to t+6).

[0015] Step 13: Coarse-grain analysis of the power sequence P(t) at hourly scale s∈{1,2,3,4}. The multi-scale coarse-grained calculation formula is:

[0016]

[0017] Where s is the scale factor, is the coarse-grained subsequence, To round down, s∈[1,3] is usually selected. The range needs to be adjusted according to the data characteristics. N is the total length of the original data.

[0018] Furthermore, in step 1, dynamic multi-scale feature extraction is performed as follows: introducing local fluctuation entropy LFE: considering, dynamically selecting key scales to retain the scales with the top 2 LFE values, the calculation formula of LFE is:

[0019]

[0020] Where n k is the number of samples in the kth interval, N s is the total number of samples of the coarse-grained sequence of the current scale s;

[0021] The following weight design is adopted:

[0022]

[0023] Where T(t) and H(t) are the temperature and humidity time series at the current moment. The following rules are adopted according to the actual effect (α=0.7, β=0.3);

[0024] The weighted probability formula is as follows:

[0025]

[0026] The final temperature and humidity joint weighted permutation entropy is defined as follows:

[0027]

[0028] Where π represents the arrangement pattern and δ(·) is the indicator function.

[0029] Furthermore, in step 2, the improved xLSTM modeling model includes an environment-aware gating layer, a quantile regression layer, and a lightweight deployment layer; the environment-aware gating layer takes historical power series, temperature, humidity, and temperature and humidity joint weighted permutation entropy as input, and generates a time series feature vector through standardization and multi-scale feature extraction, embeds a temperature and humidity correction term in the traditional forget gate, and dynamically adjusts the historical memory weight; the quantile regression layer extracts time series dependency features through LSTM units and outputs a hidden state h t The output layer adopts a quantile regression structure and synchronously outputs the power mean and confidence interval; the lightweight deployment layer adopts a teacher-student architecture. The teacher model adopts a complete xLSTM with a 64-dimensional hidden layer and outputs high-precision probability predictions. The student model adopts a lightweight xLSTM with a 32-dimensional hidden layer and inherits the distribution characteristics of the teacher model through knowledge distillation.

[0030] Furthermore, the environment-aware gating layer converts the current WPE value, temperature T(t), and humidity H(t) into historical power sequences:

[0031] Input=[WPE(s1),WPE(s2),T(t),H(t),P(t-23),...,P(t)]

[0032] Introducing a method of temperature and humidity correction terms to dynamically adjust the fusion weight of historical memory and current input

[0033] f t =σ(W f ·[h t-1 ,x t ]+b f +α·[T t ,H t ]

[0034] Where, f t is the gated output, α is a learnable parameter used to adjust the intensity of the effect of temperature and humidity on memory retention, W f is the learnable weight matrix, b f is the learnable bias term, T t ,H t Normalized values for temperature and humidity at this time.

[0035] Furthermore, the quantile regression layer is added on top of the LSTM layer to directly output multiple quantile values, selecting the 10%, 50%, and 90% quantiles:

[0036] Quantile 0.1 ,Quantile 0.5 ,Quantile 0.9 =Linear(h t )

[0037] The loss function uses the joint optimization of quantile loss (PinballLoss) and mean square error weighting, the formula is as follows:

[0038]

[0039] In the formula, the initial value of λ is 0.5, and it decays by 0.1 every 10 rounds, so as to achieve the effect of gradually transitioning from point prediction to probability prediction, where the 50% quantile (Quantile 0.5 ) as deterministic output, the corresponding 10% (Quantile 0.1 ) and 90% (Quantile 0.9 ) quantiles provide uncertainty ranges, P t is the true value, is the predicted value.

[0040] Furthermore, the lightweight deployment layer adopts a teacher-student architecture. The teacher model is a full xLSTM with a hidden layer of 64 dimensions. The student model is a lightweight xLSTM with a hidden layer of 32 dimensions. The hidden layer dimension is compressed to 32 dimensions, and the quantile regression layer is retained.

[0041] The loss function combines KL divergence and MSE:

[0042] L distill =KL(Q teacher ||Q student )+β·MSE(P true ,P student )

[0043] Where MSE is the indicator of prediction error, i.e., mean square error, which is used to ensure the mean prediction accuracy (β = 0.5), and KL divergence is used to align the probability distribution of the teacher model and the student model, i.e., Q teacher ,Q student , P student Predicted values for the student model.

[0044] Furthermore, in step 3, four evaluation indicators are used to evaluate the performance of model prediction, namely root mean square error (RMSE), mean absolute error (MAE), prediction interval average width (PI normalized averaged width, PINAW) and prediction interval coverage probability (PICP) to evaluate the prediction effect, denoted as RMSE, MAE, PINAW and PICP respectively. The evaluation indicators consider the deterministic error and probability quality respectively, and introduce an adaptive optimization strategy:

[0045] Deterministic error:

[0046]

[0047] Probability mass:

[0048]

[0049] Where N is the number of prediction samples; P i is the predicted value; is the true value.

[0050] Furthermore, if the PICP is lower than the threshold, the quantile loss weight is increased; otherwise, the interval width is optimized;

[0051] L adaptive =L+γ·(1-PICP)

[0052] Where γ is the adaptive coefficient, which is adjusted according to the real-time coverage.

[0053] Beneficial effects: Compared with the prior art, the present invention has the following significant advantages: the present invention accurately extracts multi-scale complexity features through dynamic multi-scale permutation entropy, solves the coupling problem of local fluctuations and global trends of meteorological sequences in traditional entropy methods, and ignores the problem of correlation between multi-scale characteristics and environmental variables, and accurately extracts power fluctuation features under multiple resolutions; through the gated correction mechanism of the improved xLSTM, the problem of delayed response of a single model to sudden weather changes is solved, and the historical memory weight is dynamically adjusted to enhance adaptability; the confidence interval is directly output through the quantile regression layer, which solves the problem that traditional deterministic predictions cannot quantify uncertainty, provides high-coverage probabilistic predictions, and provides risk prediction basis for power grid dispatching to support power grid risk pre-control; the dynamic integration strategy balances accuracy and computational efficiency; therefore, the present invention has the advantages of in-depth multi-scale feature mining, sensitive response to environmental mutations, strong time series modeling capabilities, high reliability of probabilistic predictions, accurate uncertainty quantification, and strong model adaptability. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] Figure 1 This is a model architecture diagram of the present invention.

[0055] Figure 2 This is the structural diagram of the improved xLSTM of the present invention. DETAILED DESCRIPTION

[0056] like Figure 1 As shown in FIG, a photovoltaic power ultra-short-term probability prediction method based on dynamic multi-scale permutation entropy and improved xLSTM includes the following steps:

[0057] Step 1: Raw data preprocessing and dynamic multi-scale feature extraction.

[0058] Photovoltaic forecasting involves multi-source heterogeneous data. Meteorological elements (such as temperature and humidity) and power generation data (such as power and voltage) have significant dimensional differences. Unprocessed raw data can lead to an imbalance in feature weights during model training, affecting the direction of gradient updates and ultimately reducing forecast accuracy. This paper uses data normalization methods to preprocess the data to eliminate dimensional effects, making features of different physical meanings comparable, increasing the integrity and continuity of photovoltaic data, and reducing the impact of weather fluctuations on forecast results. The basic principles of the normalization method are as follows:

[0059]

[0060] Where, X norm is the normalized data, x is the original data, x min is the minimum value in the original data, x max is the maximum value in the original data.

[0061] After normalization, the sliding window is used to construct the data set. The input window uses the past 24 hours (t-23 to t), and the output window outputs the power values of the next six hours (t+1 to t+6).

[0062] Next, the power sequence P(t) is coarse-grained according to the hourly scale s∈{1,2,3,4}. The multi-scale coarse-grained calculation formula is:

[0063]

[0064] Where s is the scale factor, is the coarse-grained subsequence, To round down, s∈[1,3] is usually chosen, and the range needs to be adjusted according to the characteristics of the data.

[0065] In order to solve the problem that the original algorithm depends on the number of sub-signals and combine the characteristics of photovoltaic prediction, the local fluctuation entropy (LFE) is introduced: considering the dynamic selection of key scales, the scales with the top 2 LFE values are retained. The calculation formula of LFE is:

[0066]

[0067] Improvements were made to address the issue of photovoltaic prediction and the original algorithm only weighting according to weight factors. In this experiment, relative humidity and two-meter temperature have the greatest impact on photovoltaic power. Combining the characteristics of the two and simulating the influence of light, the following weight design was adopted:

[0068]

[0069] Where T(t) and H(t) are the temperature and humidity time series at the current moment. The following rules (α = 0.7, β = 0.3) are adopted according to the actual effect.

[0070] The weighted probability formula is as follows:

[0071]

[0072] The final temperature and humidity joint weighted permutation entropy is defined as follows:

[0073]

[0074] Where π represents the arrangement pattern and δ(·) is the indicator function.

[0075] Step 2: Feature fusion and improved xLSTM modeling.

[0076] This paper uses an improved xLSTM method to improve the adaptability to sudden weather changes and the overall prediction performance by introducing temperature and humidity correction gating and quantile regression layers. The improved xLSTM model proposed in this paper has three layers: the first is the environmental perception gating layer, which takes the historical power sequence, temperature, and humidity as input, and generates a time series feature vector through standardization and multi-scale feature extraction. The temperature and humidity correction term is embedded in the traditional forget gate to dynamically adjust the historical memory weight; the second layer is the quantile regression layer (probability output layer), and the middle layer extracts the time series dependency features through the LSTM unit and outputs the hidden state h t The output layer adopts a quantile regression structure, which outputs the power mean and confidence interval simultaneously. The third layer is a lightweight deployment layer, which adopts a teacher-student architecture. The teacher model adopts a full xLSTM (hidden layer 64 dimensions) to output high-precision probability predictions, and the student model adopts a lightweight xLSTM (hidden layer 32 dimensions) to inherit the distribution characteristics of the teacher model through knowledge distillation. This structure improves the prediction speed and reduces the hardware requirements. The improved xLSTM structure is as follows: Figure 2 shown.

[0077] (1) The historical power sequence of the current WPE value, temperature T(t), and humidity H(t) is:

[0078] Input=[WPE(s1),WPE(s2),T(t),H(t),P(t-23),...,P(t)]

[0079] (2) Dynamic gate control unit: Introducing temperature and humidity correction terms to optimize the forget gate:

[0080] The forget gate of the traditional xLSTM is only based on the historical hidden state h t-1 and the current input x tThe information retention ratio is determined without considering the direct impact of environmental variables (temperature, humidity) on photovoltaic power. In scenarios with sudden weather changes (such as cloud cover and heavy rain), traditional models have difficulty quickly adjusting memory weights, resulting in increased prediction errors. Therefore, this paper introduces a temperature and humidity correction term into the forget gate to dynamically adjust the fusion weight of historical memory and current input.

[0081] f t =σ(W f ·[h t-1 ,x t ]+b f +α·[T t ,H t ]

[0082] Where α is a learnable parameter used to adjust the intensity of the effect of temperature and humidity on memory retention, T t ,H t Normalized values for temperature and humidity at this time.

[0083] (3) Quantile regression layer:

[0084] Traditional xLSTMs only output deterministic predictions and are unable to quantify the uncertainty of PV power. Grid dispatchers require both the predicted mean and confidence interval to assess potential risks. Considering the probabilistic nature of this design, this paper adds a quantile regression layer to the top layer of the LSTM to directly output multiple quantile values (e.g., 10%, 50%, and 90% quantiles):

[0085] Quantile 0.1 ,Quantile 0.5 ,Quantile 0.9 =Linear(h t )

[0086] The loss function uses the joint optimization of quantile loss (PinballLoss) and weighted mean square error. The formula is as follows:

[0087]

[0088] In the formula, the initial value of λ is 0.5, and it decays by 0.1 every 10 rounds, so as to achieve the effect of gradually transitioning from point prediction to probability prediction, where the 50% quantile (Quantile 0.5 ) as deterministic output, the corresponding 10% (Quantile 0.1 ) and 90% (Quantile 0.9 ) quantiles provide uncertainty ranges.

[0089] (4) Knowledge Distillation:

[0090] Traditional xLSTM models have a large number of parameters (hidden layers have as many as 64 dimensions), resulting in long prediction times on devices with standard hardware. However, directly compressing the model can lead to a decrease in probabilistic prediction performance, resulting in poor prediction results, which does not meet the original intention of the improved method. Therefore, this paper adopts a teacher-student architecture. The teacher model is a full xLSTM (hidden layer with 64 dimensions); the student model is a lightweight xLSTM (hidden layer with 32 dimensions). The hidden layer dimensions are compressed to 32, while the quantile regression layer is retained.

[0091] The loss function combines KL divergence and MSE:

[0092] L distill =KL(Q teacher ||Q student )+β·MSE(P true ,P student )

[0093] Where MSE is the indicator of prediction error, i.e., mean square error, which is used to ensure the mean prediction accuracy (β = 0.5), and KL divergence is used to align the probability distribution of the teacher model and the student model.

[0094] (5)TensorRT optimization:

[0095] If you need to deploy the trained model to a resource-constrained device, you need to convert the model to FP16 precision and deploy it to an edge device such as NVIDIA Jetson.

[0096] Step 3: Probabilistic forecast evaluation.

[0097] First, the prediction results are denormalized and then evaluated. The present invention adopts four evaluation indicators to evaluate the performance of model prediction, namely, root mean square error (RMSE), mean absolute error (MAE), prediction interval average bandwidth (PI normalized averaged width, PINAW) and prediction interval coverage (PICP) to evaluate the prediction effect, which are respectively denoted as RMSE, MAE, PINAW and PICP. The evaluation indicators are considered from two aspects: deterministic error and probability quality. At the same time, an adaptive optimization strategy is introduced.

[0098] (1) Deterministic error:

[0099]

[0100] (2) Probability mass:

[0101]

[0102] Where N is the number of prediction samples; P i is the predicted value; is the true value.

[0103] (3) Adaptive adjustment strategy:

[0104] In consideration of prediction accuracy, the present invention introduces a dynamic balance formula. If the PICP is lower than a threshold (such as 90%), the quantile loss weight is increased; otherwise, the interval width is optimized.

[0105] L adaptive =L+γ·(1-PICP)

[0106] Where γ is the adaptive coefficient, which is adjusted according to the real-time coverage.

[0107] To verify the performance of the present invention, a data set from the 2014 Global Energy Forecasting Competition was used for simulation verification. The data set includes solar energy data and meteorological data from April to December 2012 in a certain place. It is a 24-hour all-weather data set with a data resolution of 1 hour, which can predict the actual photovoltaic power in the next six hours. Then, the Spearman correlation coefficient of photovoltaic output and 12 types of meteorological data was calculated, and relative humidity and two-meter temperature were finally selected as meteorological characteristic data. The DMS-AMWPR method was used to extract data features and construct a new data set. The prediction time scale is 1 hour ahead, the test set is the data of the last 8 days, and the training set is the data of the first 22 days. The historical power and numerical weather forecast data are normalized before use.

[0108] To verify the effectiveness of the proposed prediction model, the method of the present invention is compared with traditional LSTM, ELM, and Stacking models. To ensure the accuracy of the experimental results, each experiment is conducted five times. Finally, the prediction result with the median prediction accuracy is selected for subsequent analysis and comparison. The short-term photovoltaic power prediction results are summarized in Table 1.

[0109] Table 1 Summary of various photovoltaic power short-term prediction indicators

[0110]

[0111] As can be seen from Table 1, the model proposed in this invention outperforms other models in every evaluation metric, achieving excellent prediction results. The MAE value is 0.0259, representing reductions of 248.6% and 81.3% compared to the other two models, respectively. The PINAW value is 0.2984, representing reductions of 42.6%, 28.6%, 42.4%, 25.5%, and 4.6% compared to the other three models, respectively. In summary, the prediction model proposed in this invention significantly improves prediction accuracy and decreases prediction error compared to other models, and the prediction results have a higher degree of fit with the true values, indicating that the ultra-short-term probabilistic photovoltaic power prediction model proposed in this invention significantly improves performance compared to other compared prediction models.

[0112] In response to the problems of insufficient multi-scale modeling and low reliability of probabilistic prediction in commonly used ultra-short-term photovoltaic power prediction, the present invention proposes a method for ultra-short-term probabilistic prediction of photovoltaic power based on dynamic multi-scale permutation entropy and improved xLSTM. Through the dynamic multi-scale permutation entropy, multi-scale complexity features are accurately extracted, which solves the coupling problem of local fluctuations and global trends of meteorological series in traditional entropy methods and the problem of ignoring the correlation between multi-scale characteristics and environmental variables, and accurately extracts power fluctuation characteristics under multiple resolutions; through the gated correction mechanism of the improved xLSTM, the problem of delayed response of a single model to sudden weather changes is solved, and the historical memory weight is dynamically adjusted to enhance adaptability; through the quantile regression layer, the confidence interval is directly output, which solves the problem that traditional deterministic prediction cannot quantify uncertainty, provides high-coverage probabilistic prediction, provides risk prediction basis for power grid dispatching and supports power grid risk pre-control; dynamic integration strategy balances accuracy and computational efficiency. Therefore, the present invention has the advantages of in-depth multi-scale feature mining, sensitive response to environmental mutations, strong time series modeling capability, high reliability of probabilistic prediction, accurate uncertainty quantification, and strong model adaptability.

Claims

1. A photovoltaic power ultra-short-term probability prediction method based on dynamic multi-scale permutation entropy and improved xLSTM, characterized by: The steps include: Step 1: Preprocess the multi-source heterogeneous data involved in photovoltaic prediction and perform dynamic multi-scale feature extraction; Step 2: Fuse the features in step 1 and improve the xLSTM modeling; Step 3: Denormalize the prediction results in step 2 and then evaluate the prediction results.

2. The photovoltaic power ultra-short-term probability prediction method based on dynamic multi-scale permutation entropy and improved xLSTM according to claim 1 is characterized in that: In step 1, preprocessing of multi-source heterogeneous data involved in photovoltaic prediction specifically includes the following steps: Step 11: Use data normalization method to preprocess the data: Where, X norm is the normalized data, x is the original data, x min is the minimum value in the original data, x max is the maximum value in the original data; Step 12: After normalization, a sliding window is used to construct a data set. The input window uses the past 24 hours (t-23 to t), and the output window outputs the power values of the next six hours (t+1 to t+6). Step 13: Coarse-grain analysis of the power sequence P(t) at hourly scale s∈{1,2,3,4}. The multi-scale coarse-grained calculation formula is: Where s is the scale factor, is the coarse-grained subsequence, To round down, s∈[1,3] is usually selected. The range needs to be adjusted according to the data characteristics. N is the total length of the original data.

3. The photovoltaic power ultra-short-term probability prediction method based on dynamic multi-scale permutation entropy and improved xLSTM according to claim 1 is characterized in that: In step 1, dynamic multi-scale feature extraction is performed as follows: Introducing local fluctuation entropy LFE: Considering, dynamically selecting key scales to retain the scales with the top 2 LFE values, the calculation formula of LFE is: Where n k is the number of samples in the kth interval, N s is the total number of samples of the coarse-grained sequence of the current scale s; The following weight design is adopted: Where T(t) and H(t) are the temperature and humidity time series at the current moment. The following rules are adopted according to the actual effect (α=0.7, β=0.3); The weighted probability formula is as follows: The final temperature and humidity joint weighted permutation entropy is defined as follows: Where π represents the arrangement pattern and δ(·) is the indicator function.

4. The photovoltaic power ultra-short-term probability prediction method based on dynamic multi-scale permutation entropy and improved xLSTM according to claim 1 is characterized in that: In step 2, the improved xLSTM modeling model includes an environment-aware gating layer, a quantile regression layer, and a lightweight deployment layer. The environment-aware gating layer takes historical power series, temperature, humidity, and the weighted permutation entropy of temperature and humidity as input, and generates a time series feature vector through standardization and multi-scale feature extraction, and dynamically adjusts the historical memory weight. The quantile regression layer extracts time series dependency features through LSTM units and outputs a hidden state h. t The output layer adopts a quantile regression structure and synchronously outputs the power mean and confidence interval; the lightweight deployment layer adopts a teacher-student architecture. The teacher model adopts a complete xLSTM with a 64-dimensional hidden layer and outputs high-precision probability predictions. The student model adopts a lightweight xLSTM with a 32-dimensional hidden layer and inherits the distribution characteristics of the teacher model through knowledge distillation.

5. The photovoltaic power ultra-short-term probability prediction method based on dynamic multi-scale permutation entropy and improved xLSTM as claimed in claim 4 is characterized in that: The environment-aware gating layer converts the current WPE value, temperature T(t), and humidity H(t) into historical power sequences: Input=[WPE(s1),WPE(s2),T(t),H(t),P(t-23),...,P(t)] Introducing a method of temperature and humidity correction terms to dynamically adjust the fusion weight of historical memory and current input f t =σ(W f ·[h t-1 ,x t ]+b f +α·[T t ,H t ] Where, f t is the gated output, α is a learnable parameter used to adjust the intensity of the effect of temperature and humidity on memory retention, W f is the learnable weight matrix, b f is the learnable bias term, T t ,H t Normalized values for temperature and humidity at this time.

6. The photovoltaic power ultra-short-term probability prediction method based on dynamic multi-scale permutation entropy and improved xLSTM according to claim 4 is characterized in that: The quantile regression layer is added on top of the LSTM layer to directly output multiple quantile values, selecting the 10%, 50%, and 90% quantiles: Quantile 0.1 ,Quantile 0.5 ,Quantile 0.9 =Linear(h t ) The loss function uses the joint optimization of quantile loss and mean square error weighting, and the formula is as follows: In the formula, the initial value of λ is 0.5, and it decays by 0.1 every 10 rounds, so as to achieve the effect of gradually transitioning from point prediction to probability prediction, where the 50% quantile (Quantile 0.5 ) as deterministic output, the corresponding 10% (Quantile 0.1 ) and 90% (Quantile 0.9 ) quantiles provide uncertainty ranges, P t is the true value, is the predicted value.

7. The photovoltaic power ultra-short-term probability prediction method based on dynamic multi-scale permutation entropy and improved xLSTM according to claim 4 is characterized in that: The lightweight deployment layer adopts a teacher-student architecture. The teacher model is a full xLSTM with a 64-dimensional hidden layer. The student model is a lightweight xLSTM with a 32-dimensional hidden layer. The hidden layer dimension is compressed to 32 dimensions, and the quantile regression layer is retained. The loss function combines KL divergence and MSE: L distill =KL(Q teacher ||Q student )+β·MSE(P true ,P student ) Where MSE is the indicator of prediction error, i.e., mean square error, which is used to ensure the mean prediction accuracy (β = 0.5), and KL divergence is used to align the probability distribution of the teacher model and the student model, i.e., Q teacher ,Q student , P student Predicted values for the student model.

8. The photovoltaic power ultra-short-term probability prediction method based on dynamic multi-scale permutation entropy and improved xLSTM according to claim 1 is characterized in that: In step 3, four evaluation indicators are used to evaluate the performance of model prediction, namely relative mean square error (RMSE), mean absolute error (MAE), average prediction interval bandwidth (PINAW), and prediction interval coverage (PICP). They are denoted as RMSE, MAE, PINAW, and PICP respectively. The evaluation indicators consider the deterministic error and probability quality respectively, and an adaptive optimization strategy is introduced at the same time: Deterministic error: Probability mass: Where N is the number of prediction samples; P i is the predicted value; is the true value.

9. The photovoltaic power ultra-short-term probability prediction method based on dynamic multi-scale permutation entropy and improved xLSTM according to claim 1 is characterized in that: If PICP is lower than the threshold, increase the quantile loss weight; otherwise, optimize the interval width; L adaptive =L+γ·(1-PICP) Where γ is the adaptive coefficient, which is adjusted according to the real-time coverage.