A method, equipment, and medium for predicting new energy output based on multidimensional index correlation analysis
By combining multidimensional index correlation analysis with a hybrid prediction model that integrates convolutional neural networks and long short-term memory networks, the problem of the comprehensive influence of factors in the prediction of new energy output has been solved, achieving high-precision and stable prediction results and meeting the precise needs of the power system.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HEFEI UNIV OF TECH
- Filing Date
- 2025-12-29
- Publication Date
- 2026-05-26
AI Technical Summary
Existing methods for predicting renewable energy output cannot fully consider the combined effects of various environmental factors and system parameters, resulting in large and unstable prediction errors. In particular, the prediction response is delayed under extreme weather conditions, which fails to meet the accuracy requirements of the power system.
This paper adopts a method based on multidimensional index correlation analysis. By acquiring historical renewable energy output data and preprocessing it, a multidimensional index system is established. A hybrid prediction model of convolutional neural network and long short-term memory network is used to extract the interaction features and time series features of the multidimensional indexes, and a renewable energy output prediction method based on multidimensional index correlation analysis is constructed.
It improves the accuracy and adaptability of new energy output forecasting, enabling accurate prediction of future output trends in dynamic environments and providing reliable technical support.
Smart Images

Figure CN122087307A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of new energy power output prediction technology, specifically a new energy power output prediction method, equipment, and medium based on multidimensional index correlation analysis. Background Technology
[0002] Due to the fluctuating and intermittent nature of renewable energy output, accurate forecasting has become a core challenge for the safe operation of the power grid. Accurate forecasting of renewable energy output is beneficial for improving the grid's absorption capacity and provides crucial support for the safe operation and optimal economic decision-making of the power system. However, renewable energy output is complex due to the coupled influence of multiple dimensions, including weather uncertainties, equipment status, and geographical information. Current traditional forecasting methods often rely on single-factor prediction models, making it difficult to comprehensively consider the combined impact of various environmental factors and system parameters on renewable energy generation, resulting in significant forecast errors and instability. For example, if coastal wind power clusters do not incorporate the meteorological transmission effects of neighboring stations, they may experience delayed forecast responses to sudden output changes under extreme weather conditions.
[0003] The current mainstream new energy output prediction technologies are mainly divided into physical principle-based models, statistical methods, and machine learning methods. However, these methods have certain limitations in practical applications: (1) They rely heavily on detailed geographical information and accurate meteorological data, and the physical formulas themselves have inherent errors. The parameter adjustment process is complex and difficult to adapt to rapidly changing environments; (2) Compared with physical methods, modeling is relatively simple, but a large amount of historical data is required as support. When encountering complex situations such as sudden weather changes, their prediction accuracy is prone to a significant drop, and it is difficult to process nonlinear and high-dimensional data; (3) Although they can provide better prediction results after feature engineering, they often lack the ability to capture long-term dependencies in time series.
[0004] In summary, existing forecasting methods are no longer sufficient to meet the precise requirements of current power systems for renewable energy output planning. Therefore, there is an urgent need to develop a renewable energy output forecasting method that can overcome reliance on a single data source, effectively integrate multi-source data, and accurately extract spatiotemporal features. Summary of the Invention
[0005] To address the technical problems existing in the prior art, this invention provides a new energy output prediction method based on multidimensional index correlation analysis. This method overcomes the dependence of traditional prediction models on a single data source and the insufficient ability to fuse and process multi-source data. It effectively extracts the spatiotemporal characteristics of photovoltaic power generation data, improves prediction accuracy, and provides reliable technical support for grid connection of high-penetration new energy sources.
[0006] To achieve the above objectives, the present invention provides the following technical solution: This invention discloses a method for predicting new energy output based on multidimensional index correlation analysis, comprising: S1. Obtain historical data on renewable energy output and preprocess the data; S2. Establish a multi-dimensional indicator system and conduct correlation analysis and screening of the multi-dimensional indicators to determine key influencing indicators; the correlation analysis includes analyzing the linear and non-linear relationships between the indicators and the output of new energy sources; S3. Construct a hybrid prediction model based on convolutional neural networks and long short-term memory networks, using preprocessed data corresponding to key influencing indicators as input to train and predict future new energy output data; wherein, convolutional neural networks are used to extract the interaction features between multidimensional indicators, and long short-term memory networks are used to capture time-series features.
[0007] As a further improvement to the above scheme, the specific steps for data preprocessing in step S1 include: Missing data is filled using linear interpolation, and the calculation formula is as follows:
[0008] In the formula, This indicates the interpolated data value at the specified time point. The estimated value; and These are the two known time points closest to the interpolation point. and They are time points respectively and The actual observed value; The data is normalized using the following formula:
[0009] In the formula, This represents the raw, unprocessed numerical data. and These are the minimum and maximum values in the original data, respectively. The data is normalized, and the range of values is within... .
[0010] As a further improvement to the above scheme, in step S2, the Pearson correlation coefficient is calculated. To perform linear relationship analysis, the formula is as follows:
[0011] In the formula, The correlation coefficient is used to characterize the correlation between indicators and new energy output. The absolute value of the correlation coefficient reflects the strength of the linear relationship: the closer it is to 1, the stronger the relationship. If the coefficient is positive, it is a positive correlation; if the coefficient is negative, it is a negative correlation. For the sample number, , The total number of samples; This refers to a specific indicator within the constructed multidimensional indicator system. This refers to the output of new energy sources themselves; and The first Variables in a sample and Preprocessed data; and Variables and The sample mean; numerator Describing covariance, express and The product of standard deviations; Among them, retaining the condition Indicators greater than 0.6.
[0012] As a further improvement to the above scheme, in step S2, the grey relational degree is calculated to perform nonlinear relationship analysis, expressed by the following formula:
[0013] In the formula, Represents the reference sequence Comparison sequence Grey relational degree; Indicates the data point sequence number. , This represents the total number of data points. Indicates the reference sequence at the 1st The value of each data point Indicates the first The comparison sequence in the th ... The value of each data point; This represents a two-level minimum difference, where the first level min represents the minimum difference for each... Find the minimum absolute difference among all comparison sequences. The second level min represents the minimum absolute difference among all comparison sequences. Find the global minimum difference in the middle; This represents the maximum difference between two levels. The first level `max` represents the maximum difference for each... Find the maximum absolute difference among all comparison sequences. The second-level max represents the maximum absolute difference among all comparison sequences. Find the global maximum difference; The resolving factor is used to adjust the sensitivity of the correlation coefficient to extreme values, and is set to 0.5. Among them, retain Indicators greater than 0.7.
[0014] As a further improvement to the above scheme, in step S2, the indicators selected after linear relationship analysis and nonlinear relationship analysis are subjected to significance test, and the indicators with p-values less than 0.05 are retained.
[0015] As a further improvement to the above scheme, in step S3, the convolutional neural network includes convolutional layers and pooling layers; wherein, the output calculation formula of the convolutional layer is:
[0016] In the formula, For the first The output of each convolutional kernel, ; Indicates the number of convolution kernels; For the first Input of each channel; Indicates the number of input channels; The activation function for the convolutional layer; and The first The weights and biases of each convolutional kernel; Pooling layers are used to aggregate interaction information between multi-dimensional metrics. The output calculation formula for pooling layers is as follows:
[0017] In the formula, To represent the first step after the pooling operation j The output feature values of each channel ; Indicates the total number of input channels; This indicates the number of data items in the pooling window.
[0018] As a further improvement to the above scheme, step S3, the construction process of the Long Short-Term Memory network includes: Constructing a forget gate to filter historical information, expressed by the following formula:
[0019] In the formula, Indicates the forget gate at the current moment. t The output vector; For the Sigmoid function; This represents the input information at the current moment; Indicates input information The weight of the Gate of Oblivion; Represents the hidden state at the previous time step. The weight of the Gate of Oblivion; For the bias term of the forget gate; The input gate controls the input of new information, expressed by the following formula:
[0020] In the formula, This represents the output vector of the input gate at the current moment; Indicates input information The weights to the input gate; Represents the hidden state at the previous time step. The weights to the input gate; This represents the bias term of the input gate; The candidate cell state is calculated using the following formula:
[0021] In the formula, This represents the candidate cell state at the current moment; tanh represents the hyperbolic tangent function. This indicates that information is input into the input gate. The weights for candidate states; Represents the hidden state in the input gate The weights for candidate states; Bias terms representing candidate states; The cell state at the previous time step is updated using candidate cell states, expressed by the following formula:
[0022] In the formula, This indicates the current, updated cell state. It indicates the cell state at the previous moment; This represents element-wise multiplication; This indicates historical information that has been retained. This indicates newly added current information; The output information is controlled by the output gate, and the current hidden layer state is calculated. The formula is as follows:
[0023]
[0024] In the formula, This represents the output of the cell unit at the current moment. This represents the current hidden state at the time step. Indicates input information The weights of the output gates; Represents hidden state The weights of the output gates; This represents the bias term of the output gate.
[0025] As a further improvement to the above scheme, in step S1, the relevant historical renewable energy output data includes: total solar irradiance, direct irradiance, global horizontal irradiance, temperature, atmospheric pressure, relative humidity, and wind speed and direction.
[0026] The present invention also discloses a computer terminal, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the steps of a new energy output prediction method based on multidimensional index correlation analysis as described above.
[0027] The present invention also discloses a computer-readable storage medium storing a computer program thereon, characterized in that, when the program is executed by a processor, it implements the steps of a new energy output prediction method based on multidimensional index correlation analysis as described above.
[0028] Compared with the prior art, the beneficial effects of the present invention are: This invention acquires historical renewable energy output data and performs multi-source data fusion preprocessing, effectively integrating information from different sources and improving the accuracy and comprehensiveness of the data. This facilitates the efficient utilization and processing of multi-dimensional data in subsequent analysis. By establishing a multi-dimensional indicator system and combining multi-dimensional indicator correlation analysis and screening, the invention quantifies the impact of each indicator on renewable energy output. By dynamically adjusting the screening criteria, it automatically removes non-critical indicators with weak correlations, thus overcoming the limitations of traditional single-factor prediction models and improving the accuracy and operability of the indicators. Furthermore, this invention constructs a renewable energy output prediction method based on an LSTM prediction model. LSTM processing of time series data can remember important past information and combine it with current input, which is crucial for capturing the temporal characteristics in photovoltaic output capture data. This effectively predicts future renewable energy output trends and improves the adaptability and accuracy of the prediction model in dynamic environments. Attached Figure Description
[0029] Figure 1 This is a structural diagram of the wind farm power generation system in Embodiment 1 of the present invention.
[0030] Figure 2 This is a structural diagram of the photovoltaic power generation system in Embodiment 1 of the present invention.
[0031] Figure 3 This is a flowchart of the new energy output prediction method based on multidimensional index correlation analysis in Embodiment 1 of the present invention.
[0032] Figure 4This is a schematic diagram of the LSTM unit structure in Embodiment 1 of the present invention.
[0033] Figure 5 This is a schematic diagram of the computer equipment in Embodiment 2 of the present invention. Detailed Implementation
[0034] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0035] Example 1
[0036] This embodiment provides a new energy output prediction method based on multidimensional index correlation analysis. This method can be applied to... Figure 1 The wind farm shown consists of wind turbines, rectifiers, combiner boxes, inverters, internal loads, and ultimately, the main power grid to which it is connected. It can also be applied to... Figure 2 The photovoltaic power plant system shown is similar to the wind farm power generation system.
[0037] like Figure 3 As shown, the new energy output prediction method based on multidimensional index correlation analysis provided in this embodiment includes steps S1 to S3.
[0038] S1. Obtain historical data on renewable energy output and preprocess the data.
[0039] Specifically, historical data on new energy output and related environmental variables are collected, including total solar irradiance, direct irradiance, global horizontal irradiance, temperature, atmospheric pressure, relative humidity, wind speed and direction, and integrated to form a dataset at a specific point in time.
[0040] In step S1, the specific steps for preprocessing the data include: In the data preprocessing stage, missing data is systematically processed, outliers are identified and corrected, and variable correlations are analyzed to screen key multidimensional influencing factors, thereby improving the completeness and reliability of the dataset. Missing values include common empty values (such as 0, null, 'NA'); outliers include long-term constant meteorological data and values that significantly exceed reasonable ranges. Therefore, linear interpolation is used to impute missing data, and the calculation formula is as follows:
[0041] In the formula, This indicates the interpolated data value at the specified time point. The estimated value; and These are the two known time points closest to the interpolation point. and They are time points respectively and The actual observed value.
[0042] In addition, if obvious outliers are reasonably removed, this dataset will not be used.
[0043] Data from different dimensions may have different units and dimensions, requiring standardization or normalization for subsequent analysis and modeling. Therefore, min-max normalization is used to standardize the acquired data from different dimensions; the calculation formula is as follows:
[0044] In the formula, This represents the raw, unprocessed numerical data. and These are the minimum and maximum values in the original data, respectively. After normalization, the range of the data is compressed to [a certain value]. .
[0045] After processing, the minimum value of the data is normalized to 0, the maximum value is normalized to 1, and other data are mapped to this interval proportionally.
[0046] S2. Establish a multidimensional indicator system and conduct correlation analysis and screening of the multidimensional indicators to determine key influencing indicators; the correlation analysis includes analyzing the linear and nonlinear relationships between the indicators and the output of new energy sources.
[0047] In step S2, the data corresponding to the indicator is the normalized value of the indicator at a specific time point after processing. The correlation between multiple multidimensional indicators affecting renewable energy output is analyzed to determine which variables have a significant impact on the prediction target (renewable energy output), and redundant and irrelevant indicators are removed. The Pearson correlation coefficient is used to calculate and quantify the correlation between input indicators (such as temperature, wind speed, humidity, solar irradiance, etc.) and renewable energy output. The closer the coefficient is to 1, the stronger the linear relationship; a positive sign indicates a positive correlation, and vice versa. The Pearson correlation coefficient is then calculated. To perform linear relationship analysis, the formula is as follows:
[0048] In the formula, The correlation coefficient is used to characterize the correlation between indicators and new energy output. The absolute value of the correlation coefficient reflects the strength of the linear relationship: the closer it is to 1, the stronger the relationship. If the coefficient is positive, it is a positive correlation; if the coefficient is negative, it is a negative correlation. For the sample number, , The total number of samples; This refers to a specific indicator within the constructed multidimensional indicator system. This refers to the output of new energy sources themselves; and The first Variables in a sample and Preprocessed data; and Variables and The sample mean; numerator Describing covariance, express and The product of standard deviations; Retain satisfaction Indicators greater than 0.6.
[0049] The grey relational degree is calculated for nonlinear relationship analysis, and the formula is as follows:
[0050] In the formula, Represents the reference sequence Comparison sequence Grey relational degree; Indicates the data point sequence number. , This represents the total number of data points. Indicates the reference sequence at the 1st The value of each data point Indicates the first The comparison sequence in the th ... The value of each data point; This represents a two-level minimum difference, where the first level min represents the minimum difference for each... Find the minimum absolute difference among all comparison sequences. The second level min represents the minimum absolute difference among all comparison sequences. Find the global minimum difference in the middle; This represents the maximum difference between two levels. The first level `max` represents the maximum difference for each... Find the maximum absolute difference among all comparison sequences. The second-level max represents the maximum absolute difference among all comparison sequences. Find the global maximum difference; The resolving factor is used to adjust the sensitivity of the correlation coefficient to extreme values, and is set to 0.5. Among them, retain Indicators greater than 0.7.
[0051] Finally, the indicators selected after linear and nonlinear relationship analysis were subjected to significance tests, and those with p-values less than 0.05 were retained. The smaller the p-value, the more significant the correlation coefficient, meaning the correlation between the two variables is less likely to be random. Generally, a p-value less than 0.05 is considered significant, indicating a significant linear correlation between the two variables. If the absolute value of the correlation coefficient is large and the p-value is small, then a significant linear correlation exists between the two variables, and the indicator is retained.
[0052] S3. Construct a hybrid prediction model based on convolutional neural networks and long short-term memory networks, using preprocessed data corresponding to key influencing indicators as input to train and predict future new energy output data; wherein, convolutional neural networks are used to extract the interaction features between multidimensional indicators, and long short-term memory networks are used to capture time-series features.
[0053] In step S3, a CNN-LSTM-based model is constructed: the introduction of a convolutional neural network (CNN) can effectively learn the interaction of multi-dimensional indicators, enabling the model to automatically extract the complex relationships and interaction effects between various filtered indicators from the input data; the long short-term memory network (LSTM) can remember important past information and combine it with the current input to predict future photovoltaic output.
[0054] The convolutional neural network includes convolutional layers and pooling layers. It can learn local information about power generation at a given time, such as temperature, through convolutional kernels. A convolutional kernel with certain connection weights, also called a filter, acts on different input channels without changing the weights, producing the filter's output.
[0055] In the formula, For the first The output of each convolutional kernel, ; Indicates the number of convolution kernels; For the first Input of each channel; Indicates the number of input channels; The activation function for the convolutional layer is usually tanh, i.e., the hyperbolic tangent function; and The first The weights and biases of each convolutional kernel; After feature extraction by convolutional layers, pooling layers (such as max pooling or average pooling) are added. In multi-dimensional indicator interactions, pooling layers help aggregate the interaction information between multi-dimensional indicators, further exploring the relationships between these indicators. The output calculation formula of the pooling layer is as follows:
[0056] In the formula, To represent the first step after the pooling operation j The output feature values of each channel ; Indicates the total number of input channels; This indicates the number of data items in the pooling window.
[0057] As the number of network layers increases, convolutional networks can transition from local to global. In high-level convolutions, the convolutional kernel learns complex interaction patterns between multiple metrics. These interactions are no longer limited to a single metric, but rather the combined effect of multiple metrics, revealing how these metrics jointly affect power generation.
[0058] LSTM cells include three gate structures: input gate, forget gate, and output gate. These three gate structures are as follows: Figure 4 As shown, it uses these three gate structures to achieve the functions of filtering, memorizing, and updating information within the entire unit. The LSTM input data is processed by a CNN to obtain new data, which is then used as input. The following formula is used to construct a forgetting gate to filter out useless information and determine how much historical information to retain:
[0059] In the formula, Indicates the forget gate at the current moment. t The output vector controls the proportion of the cell state retained from the previous moment; The Sigmoid function compresses the output to the [0,1] interval, which can be used as a gate signal; This represents the input information at the current moment (such as meteorological data and other features). Indicates input information The weight of the Gate of Oblivion; Represents the hidden state at the previous time step. The weight of the Gate of Oblivion; For the bias term of the forget gate; After the forget gate filtering, the hidden state of the previous moment. and the input information at the current moment The input gate then controls the input of new information, and the following formula is used to construct the input gate for filtering the input information:
[0060] In the formula, This represents the output vector of the input gate at the current moment, which controls the proportion of new information input at the moment. Indicates input information The weights to the input gate; Represents the hidden state at the previous time step. The weights to the input gate; This represents the bias term of the input gate; Candidate cell states represent the original or input information at the current moment. This information will be used to update the cell state in the next step. The candidate cell state is calculated using the following formula:
[0061] In the formula, This represents the current state of the candidate cell, including the current input and previous state information; tanh represents the hyperbolic tangent function, which compresses the output to the interval [-1, 1]. This indicates that information is input into the input gate. The weights for candidate states; Represents the hidden state in the input gate The weights for candidate states; Bias terms representing candidate states; Output of the Forgot Gate f t Adjusting the cell state at the previous moment The degree of forgetting, and its relation to Multiplication filters historical information, retaining only valid features for use in the current cell state. The update is performed by calculating the current cell state using the following formula:
[0062] In the formula, This represents the current cell state, i.e., the updated cell state (long-term memory). It indicates the cell state at the previous moment; This represents element-wise multiplication (Hadamard product). This indicates historical information that has been retained. This indicates newly added current information; Output gates control which information should be passed to the next layer of the network or used as the final output. The following formula is used to construct the output gate that controls the output information:
[0063] In the formula, This represents the output of the cell unit at the current moment; Indicates input information The weights of the output gates; Represents hidden state The weights of the output gates; This represents the bias term of the output gate.
[0064] The following formula is used to calculate the output that is passed to the next layer of the network or the final output of the network:
[0065] In the formula, This represents the hidden state at the current time step; the weights and biases in the model are obtained through the training process of the algorithm.
[0066] In this embodiment, the hybrid prediction model takes the data corresponding to the indicators after screening in step S2 as input. The multiple convolutional layers can not only capture the relationship between certain external factors (such as temperature, radiation intensity, etc.) and power generation of a single indicator in a short period of time (such as the past few hours or a day), but also learn the interaction effect between different indicators. The long short-term memory network can remember important information in the past and combine it with the current input. By continuously adjusting the prediction error and automatically updating the parameters, the prediction accuracy is improved.
[0067] After the model is trained, the input features of the period to be predicted are input into the trained model. The model makes a prediction by comprehensively considering historical data and current features, and outputs the predicted value of new energy output for the future period.
[0068] Example 2
[0069] This embodiment provides a computer terminal, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the steps of the new energy output prediction method based on multidimensional index correlation analysis as described in Embodiment 1.
[0070] like Figure 5 As shown, the computer terminal provided in this embodiment includes: at least one processor 101, and a memory 102 connected to at least one processor 101. This embodiment does not limit the specific connection medium between the processor 101 and the memory 102. Figure 5 The example shown is the connection between processor 101 and memory 102 via bus 100. Bus 100 is... Figure 5 The connections between other components are shown in bold lines and are for illustrative purposes only, not as limiting information. Bus 100 can be divided into address bus, data bus, control bus, etc., for ease of representation. Figure 5 The bus is represented by a single thick line, but this does not indicate that there is only one bus or one type of bus. Alternatively, the processor 101 may also be called a controller; there is no restriction on the name.
[0071] In this embodiment, the memory 102 stores instructions that can be executed by at least one processor 101. The at least one processor 101 can execute the aforementioned method by executing the instructions stored in the memory 102.
[0072] The processor 101 is the control center of the device. It can connect to various parts of the control device through various interfaces and lines. By running or executing instructions stored in memory 102 and calling data stored in memory 102, the processor can perform various functions and process data, thereby monitoring the device as a whole.
[0073] In one possible design, processor 101 may include one or more processing units. Processor 101 may integrate an application processor and a modem processor, wherein the application processor mainly handles the operating system, user interface, and applications, and the modem processor mainly handles wireless communication. It is understood that the modem processor may also not be integrated into processor 101. In some embodiments, processor 101 and memory 102 may be implemented on the same chip; in some embodiments, they may also be implemented on separate chips.
[0074] Processor 101 can be a general-purpose processor, such as a central processing unit (CPU), digital signal processor, application-specific integrated circuit, field-programmable gate array or other programmable logic device, discrete gate or transistor logic device, or discrete hardware component, capable of implementing or executing the methods, steps, and logic block diagrams disclosed in the embodiments. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the new energy output prediction method based on multi-dimensional index correlation analysis disclosed in Embodiment 1 can be directly manifested as being executed by a hardware processor, or executed by a combination of hardware and software modules in processor 101.
[0075] Memory 102, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules. Memory 102 may include at least one type of storage medium, such as flash memory, hard disk, multimedia card, card-type memory, random access memory (RAM), static random access memory (SRAM), programmable read-only memory (PROM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), magnetic storage, magnetic disk, optical disk, etc. Memory 102 can be any other medium capable of carrying or storing desired program code in the form of instructions or data structures that can be accessed by a computer, but is not limited thereto. In this embodiment, memory 102 can also be a circuit or any other device capable of implementing storage functions for storing program instructions and / or data.
[0076] By designing and programming the processor 101, the code corresponding to the new energy output prediction method based on multidimensional index correlation analysis described in the foregoing embodiments can be embedded into the chip, thereby enabling the chip to execute the code during runtime. Figure 3 The steps of the new energy output prediction method based on multidimensional index correlation analysis are shown. How to design and program the processor 101 is a technique well-known to those skilled in the art and will not be described further here.
[0077] Example 3
[0078] This embodiment provides a computer-readable storage medium storing a computer program thereon. When the program is executed by a processor, it implements the steps of the new energy output prediction method based on multidimensional index correlation analysis as described in Embodiment 1.
[0079] The computer-readable storage medium may include flash memory, hard disk, multimedia card, card-type memory (e.g., SD or DX memory), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the storage medium may be an internal storage unit of a computer device, such as the hard disk or memory of the computer device. In other embodiments, the storage medium may also be an external storage device of the computer device, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc., provided on the computer device. Of course, the storage medium may include both internal storage units and external storage devices of the computer device. In this embodiment, the memory is typically used to store the operating system and various application software installed on the computer device. In addition, the memory can also be used to temporarily store various types of data that have been output or will be output.
[0080] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.
Claims
1. A method for predicting new energy output based on multidimensional index correlation analysis, characterized in that, include: S1. Obtain historical data on renewable energy output and preprocess the data; S2. Establish a multi-dimensional indicator system and conduct correlation analysis and screening of the multi-dimensional indicators to determine key influencing indicators; the correlation analysis includes analyzing the linear and non-linear relationships between the indicators and the output of new energy sources; S3. Construct a hybrid prediction model based on convolutional neural networks and long short-term memory networks, using preprocessed data corresponding to key influencing indicators as input to train and predict future new energy output data; wherein, convolutional neural networks are used to extract the interaction features between multidimensional indicators, and long short-term memory networks are used to capture time-series features.
2. The new energy output prediction method based on multidimensional index correlation analysis according to claim 1, characterized in that, In step S1, the specific steps for preprocessing the data include: Missing data is filled using linear interpolation, and the calculation formula is as follows: In the formula, This indicates the interpolated data value at the specified time point. The estimated value; and These are the two known time points closest to the interpolation point. and They are time points respectively and The actual observed values; The data is normalized using the following formula: In the formula, This represents the raw, unprocessed numerical data. and These are the minimum and maximum values in the original data, respectively. The data is normalized, and the range of values is within... .
3. The new energy output prediction method based on multidimensional index correlation analysis according to claim 1, characterized in that, In step S2, the Pearson correlation coefficient is calculated. To perform linear relationship analysis, the formula is as follows: In the formula, The correlation coefficient is used to characterize the correlation between indicators and new energy output. The absolute value of the correlation coefficient reflects the strength of the linear relationship: the closer it is to 1, the stronger the relationship. If the coefficient is positive, it is a positive correlation; if the coefficient is negative, it is a negative correlation. For the sample number, , The total number of samples; This refers to a specific indicator within the constructed multidimensional indicator system. This refers to the output of new energy sources themselves; and The first Variables in a sample and Preprocessed data; and Variables and The sample mean; numerator Describing covariance, express and The product of standard deviations; Among them, retaining the condition Indicators greater than 0.
6.
4. The new energy output prediction method based on multidimensional index correlation analysis according to claim 3, characterized in that, In step S2, the grey relational degree is calculated for nonlinear relationship analysis, expressed by the following formula: In the formula, Represents the reference sequence Comparison sequence Grey relational degree; Indicates the data point sequence number. , This represents the total number of data points. Indicates the reference sequence at the 1st The value of each data point Indicates the first The comparison sequence in the th ... The value of each data point; This represents a two-level minimum difference, where the first level min represents the minimum difference for each... Find the minimum absolute difference among all comparison sequences. The second level min represents the minimum absolute difference among all comparison sequences. Find the global minimum difference in the middle; This represents the maximum difference between two levels. The first level `max` represents the maximum difference for each... Find the maximum absolute difference among all comparison sequences. The second-level max represents the maximum absolute difference among all comparison sequences. Find the global maximum difference; The resolving factor is used to adjust the sensitivity of the correlation coefficient to extreme values, and is set to 0.
5. Among them, retain Indicators greater than 0.
7.
5. The new energy output prediction method based on multidimensional index correlation analysis according to claim 4, characterized in that, In step S2, the indicators selected after linear and nonlinear relationship analysis are subjected to significance tests, and indicators with p-values less than 0.05 are retained.
6. The new energy output prediction method based on multidimensional index correlation analysis according to claim 1, characterized in that, In step S3, the convolutional neural network includes convolutional layers and pooling layers; wherein, the output calculation formula of the convolutional layer is: In the formula, For the first The output of each convolutional kernel, ; Indicates the number of convolution kernels; For the first Input of each channel; Indicates the number of input channels; The activation function for the convolutional layer; and The first The weights and biases of each convolutional kernel; Pooling layers are used to aggregate interaction information between multi-dimensional metrics. The output calculation formula for pooling layers is as follows: In the formula, To represent the first step after the pooling operation j The output feature values of each channel ; Indicates the total number of input channels; This indicates the number of data items in the pooling window.
7. The new energy output prediction method based on multidimensional index correlation analysis according to claim 6, characterized in that, In step S3, the construction process of the Long Short-Term Memory network includes: Constructing a forget gate to filter historical information, expressed by the following formula: In the formula, Indicates the forget gate at the current moment. t The output vector; For the Sigmoid function; This represents the input information at the current moment; Indicates input information The weight of the Gate of Oblivion; Represents the hidden state at the previous time step. The weight of the Gate of Oblivion; For the bias term of the forget gate; The input gate controls the input of new information, expressed by the following formula: In the formula, This represents the output vector of the input gate at the current moment; Indicates input information The weights to the input gate; Represents the hidden state at the previous time step. The weights to the input gate; This represents the bias term of the input gate; The candidate cell state is calculated using the following formula: In the formula, This represents the candidate cell state at the current moment; tanh represents the hyperbolic tangent function. This indicates that information is input into the input gate. The weights for candidate states; Represents the hidden state in the input gate The weights for candidate states; Bias terms representing candidate states; The cell state at the previous time step is updated using the candidate cell state, expressed by the following formula: In the formula, This indicates the current, updated cell state. It indicates the cell state at the previous moment; This represents element-wise multiplication; This indicates historical information that has been retained. This indicates newly added current information; The output information is controlled by the output gate, and the current hidden layer state is calculated. The formula is as follows: In the formula, This represents the output of the cell unit at the current moment. This represents the current hidden state at the time step. Indicates input information The weights of the output gates; Represents hidden state The weights of the output gates; This represents the bias term of the output gate.
8. The new energy output prediction method based on multidimensional index correlation analysis according to claim 1, characterized in that, In step S1, the historical renewable energy output data includes: total solar irradiance, direct irradiance, global horizontal irradiance, temperature, atmospheric pressure, relative humidity, and wind speed and direction.
9. A computer apparatus comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of a new energy output prediction method based on multidimensional index correlation analysis as described in any one of claims 1 to 8.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the steps of a new energy output prediction method based on multidimensional index correlation analysis as described in any one of claims 1 to 8.