Photovoltaic power prediction method and device based on historical data and weather

By constructing a localized irradiance feature library and real-time data correction, combined with machine learning models, the problem of insufficient adaptation of general weather forecast data in photovoltaic power prediction has been solved, improving the accuracy and adaptability of prediction.

CN122000877APending Publication Date: 2026-05-08BEIJING TRUTH WISDOM POWER TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING TRUTH WISDOM POWER TECH CO LTD
Filing Date
2026-01-19
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

In existing photovoltaic power prediction methods, the general numerical weather prediction data is not well adapted to the local climate characteristics of specific photovoltaic power plants, resulting in reduced prediction accuracy.

Method used

A localized irradiance feature library was constructed. Historical data was classified and processed to generate daily irradiance variation curves representing different weather patterns. These curves were then corrected using numerical weather prediction data. Real-time operational data and machine learning models were used for prediction. Finally, measured power data were used to correct the predicted curves.

Benefits of technology

It improves the accuracy and adaptability of photovoltaic power forecasting, reduces meteorological input deviation, and achieves efficient adaptation and dynamic optimization to local climate patterns.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122000877A_ABST
    Figure CN122000877A_ABST
Patent Text Reader

Abstract

The invention discloses a photovoltaic power prediction method and device based on historical data and weather. The method comprises the following steps: acquiring historical irradiation data, historical output data and historical weather data; the method comprises the following steps: performing classification processing on historical irradiation data to generate a plurality of daily irradiation change curves, and constructing a localized irradiation feature library; acquiring numerical weather forecast data of the target photovoltaic power station; searching a target daily irradiation change curve, and correcting the numerical weather forecast data to generate target meteorological data; acquiring real-time power station operation data, inputting the target meteorological data and the real-time power station operation data into a machine learning prediction model, and generating an initial power prediction curve; and acquiring latest actually measured power data, determining a prediction deviation, and correcting the initial power prediction curve to obtain a final power prediction curve. By implementing the technical scheme provided by the invention, the adaptability of the prediction curve to the local unique climate law is improved, so that the overall accuracy of photovoltaic power prediction is effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of photovoltaic power generation technology, specifically to a method and apparatus for predicting photovoltaic power based on historical data and weather conditions. Background Technology

[0002] Photovoltaic power forecasting is a crucial step in ensuring the safe and stable operation of the power grid. Current photovoltaic power forecasting methods typically rely on two main data sources: general numerical weather prediction data released by meteorological services and historical output data collected by the photovoltaic power plants themselves. In practice, these methods use these two types of data as core inputs, processing them through pre-built machine learning or statistical models to generate power forecast curves for specific future time periods, providing a basis for power system dispatch decisions.

[0003] However, the aforementioned existing technologies suffer from insufficient accuracy in practical applications. The root cause lies in the fact that the numerical weather prediction data used in these methods is typically the output of a wide-area, general model, which is neither suitable nor capable of being deeply adapted to the local climate characteristics of a specific photovoltaic power station location. Therefore, when actual weather changes exhibit unique, periodic evolution patterns due to specific geographical factors such as topography and water bodies, the deviation between the general weather forecast data and the actual local meteorological conditions will significantly increase. This deviation will be directly transmitted to the output of the prediction model, leading to a decrease in the accuracy of the final power prediction. Summary of the Invention

[0004] To address the aforementioned technical problems, this application provides a method and apparatus for predicting photovoltaic power based on historical data and weather conditions.

[0005] The first aspect of this application provides a photovoltaic power prediction method based on historical data and weather, employing the following technical solution: Acquire historical irradiance data, historical power output data, and historical weather data of the target photovoltaic power station; By classifying and processing the historical irradiance data, multiple daily irradiance variation curves representing different weather patterns are generated, and a localized irradiance feature library is constructed based on the daily irradiance variation curves, the historical power output data, and the historical weather data. Numerical weather forecast data for the target photovoltaic power station is obtained based on the preset prediction start time and prediction duration. In the localized irradiance feature library, a target daily irradiance variation curve that matches the numerical weather forecast data is retrieved, and the numerical weather forecast data is corrected using the target daily irradiance variation curve to generate target meteorological data. The real-time power plant operation data of the target photovoltaic power plant is obtained, and the target meteorological data and the real-time power plant operation data are input into a preset machine learning prediction model to generate an initial power prediction curve. The latest measured power data of the target photovoltaic power station is obtained. Based on the latest measured power data and the initial power prediction curve, the prediction deviation is determined. The initial power prediction curve is then corrected based on the prediction deviation to obtain the final power prediction curve.

[0006] By employing the aforementioned technical solution, a localized irradiance feature library is constructed. Numerical weather forecast data is matched and corrected with the historical irradiance characteristics of the target power plant, effectively mitigating the systematic deviation between general forecast data and actual local meteorological conditions. Based on this, preliminary predictions are made by combining real-time operational data with machine learning models, and a deviation correction mechanism based on the latest measured power is introduced to form a closed-loop optimization. This method improves the adaptability of the prediction curve to unique local climate patterns, thereby effectively improving the overall accuracy of photovoltaic power prediction.

[0007] Optionally, the step of classifying the historical irradiance data to generate multiple daily irradiance variation curves representing different weather patterns, and constructing a localized irradiance feature library based on the daily irradiance variation curves, the historical power output data, and the historical weather data, includes: The historical irradiance data is sequentially processed by dividing it into natural days, aligning it with the time axis, and normalizing the amplitude to obtain multiple daily irradiance data sequences; Based on the data distribution characteristics of the multiple solar irradiance data sequences, a preset unsupervised clustering algorithm is applied to divide the multiple solar irradiance data sequences into multiple data clusters; For each data cluster, the corresponding solar radiation variation curve is generated by calculating the centroid of all solar radiation data sequences within the data cluster; For each data cluster, extract the output record corresponding to the solar irradiance data sequence associated with the data cluster in time from the historical output data, and calculate the historical output feature vector based on the output record; For each data cluster, weather records corresponding to the solar radiation data sequence associated with the data cluster in time are extracted from the historical weather data, and historical weather feature vectors are calculated based on the weather records. Each of the daily irradiance variation curves, the corresponding historical power output feature vectors, and the corresponding historical weather feature vectors are associated and stored to construct the localized irradiance feature library.

[0008] By adopting the above technical solution, historical irradiance data is first standardized and automatically clustered, thereby objectively and efficiently summarizing massive historical data into several representative typical daily weather patterns and generating their corresponding standard irradiance variation curves. Furthermore, each weather pattern is associated with its historical power output characteristics and meteorological condition characteristics, jointly constructing a structured, multi-dimensional, and integrated localized feature knowledge base.

[0009] Optionally, the step of retrieving a target daily irradiance variation curve that matches the numerical weather forecast data from the localized irradiance feature library, and using the target daily irradiance variation curve to correct the numerical weather forecast data to generate target meteorological data, includes: Weather features are extracted from the numerical weather forecast data, and a weather feature vector to be matched is constructed based on the weather features. For each diurnal radiation variation curve in the localized radiation feature library, the Euclidean distance between the weather feature vector to be matched and the historical weather feature vector associated with the diurnal radiation variation curve is calculated to obtain the dissimilarity score corresponding to the diurnal radiation variation curve. Among all the anisotropy scores, the target anisotropy score with the smallest value is determined, and the solar radiation variation curve corresponding to the target anisotropy score is determined as the target solar radiation variation curve; The irradiance prediction value in the numerical weather forecast data is fused with the target day irradiance variation curve to obtain the corrected irradiance prediction value. The corrected irradiance prediction value is used to replace the irradiance prediction value in the numerical weather forecast data to generate the target meteorological data.

[0010] By employing the aforementioned technical solution, the most suitable future weather pattern is effectively identified. Then, the irradiance of the general numerical weather prediction is corrected using the matched historical actual irradiance variation curves, refining the wide-area, averaged forecast data into "target meteorological data" that better reflects the local climate patterns of the power plant. This step reduces meteorological bias at the input source, laying a reliable data foundation for subsequent high-precision power prediction.

[0011] Optionally, the step of fusing the predicted irradiance value from the numerical weather forecast data with the target day's irradiance variation curve to obtain a corrected predicted irradiance value includes: The predicted total daily irradiance is obtained by integrating the predicted irradiance value over the predicted duration. The total irradiance of the base day is obtained by integrating the target day's irradiance variation curve along the predicted duration. Calculate the amplitude adjustment coefficient based on the predicted daily total irradiance and the baseline daily total irradiance; The amplitude adjustment coefficient is multiplied by the target daily irradiance variation curve, and the result of the multiplication is determined as the corrected irradiance prediction value.

[0012] By employing the aforementioned technical solution, refined correction of general forecast data is achieved. This method, while fully preserving the historical daily irradiance curve morphology that best reflects local meteorological evolution, utilizes daily total irradiance prediction information provided by numerical weather prediction to perform overall amplitude scaling. This effectively incorporates quantitative forecast information for the forecast day while maximizing the inheritance of the temporal variation characteristics of localized historical models, thereby generating high-quality meteorological input data that combines the latest forecast trends with actual local patterns.

[0013] Optionally, the step of obtaining the latest measured power data of the target photovoltaic power station, determining the prediction deviation based on the latest measured power data and the initial power prediction curve, and correcting the initial power prediction curve based on the prediction deviation to obtain the final power prediction curve includes: The measured power sequence of the target photovoltaic power station within a preset backtracking time before the start time of the prediction period is taken as the latest measured power data; From the initial power prediction curve, extract the historical predicted power sequence corresponding to each time point within the preset backtracking time of the latest measured power data; Based on the latest measured power data and the historical predicted power sequence, the average power deviation within the preset backtracking time is calculated, and the average power deviation is determined as the prediction deviation. The prediction deviation is input into a preset deviation attenuation algorithm to calculate a correction sequence that attenuates along the prediction duration; The correction sequence is superimposed on the initial power prediction curve to generate the final power prediction curve.

[0014] By adopting the above technical solution, the dynamic adaptability of the forecast is effectively enhanced. This method uses the latest measured data to quickly identify and quantify systematic short-term forecast biases, and then uses a decay algorithm to reasonably allocate this correction amount to future forecast periods. This process can automatically compensate for errors caused by factors not captured by the model, such as equipment performance fluctuations, temporary obstructions, or ultra-short-term weather changes, thereby significantly improving the accuracy and overall robustness of the forecast curve in the initial stage.

[0015] Optionally, the method further includes: Within the time interval of the prediction duration, the measured power values ​​corresponding to the time points of the final power prediction curve are continuously acquired from the data acquisition system of the target photovoltaic power station to obtain subsequent measured power data. By calculating the prediction error between the subsequent measured power data and the final power prediction curve, it is determined whether the prediction error continues to exceed a preset performance degradation threshold within a preset evaluation time window. When it is determined that the prediction error continues to exceed the performance degradation threshold, an online model update instruction is triggered, and an incremental training sample set is jointly constructed based on the subsequent measured power data obtained within the preset evaluation time window and the target meteorological data corresponding to the time range of the subsequent measured power data.

[0016] By adopting the above technical solution, adaptive closed-loop optimization of the prediction model was achieved. This method can keenly identify the decline in prediction performance caused by changes in the power plant's operating status or long-term drift in the external environment, and automatically initiate an incremental learning process. The model is dynamically updated using the latest measured data, enabling it to continuously track and adapt to changes in power plant characteristics, thereby effectively extending the model's effective lifespan, ensuring long-term prediction accuracy, and reducing reliance on large-scale periodic retraining.

[0017] Optionally, the method further includes: In response to the online update instruction for the model, the machine learning prediction model is deconstructed into a shared feature layer and a task-specific layer; When updating the machine learning prediction model, the network parameters of the shared feature layer are configured to not update the gradient. Based on the incremental training sample set, the gradient calculation and weight update of the network parameters of the specific task layer are performed only on the backpropagation algorithm to obtain the updated specific task layer. The updated task-specific layer is combined with the shared feature layer, which keeps the network parameters unchanged, to generate an updated machine learning prediction model.

[0018] By adopting the above technical solution, a hierarchical parameter update strategy is employed when initiating model updates. This method helps maintain the stability of the general meteorological and power correlation features learned by the model by freezing the parameters of the shared feature layer; simultaneously, lightweight incremental training is performed only on the parameters of specific task layers, enabling the model to adapt to the latest changes in the power plant's operating status using new data. This strategy helps reduce the risk of interfering with existing knowledge while improving the model's adaptability to current data, and achieves continuous model optimization with high efficiency.

[0019] A second aspect of this application provides an electronic device including a processor, a memory, a user interface, and a network interface, wherein the memory is used to store instructions, the user interface and the network interface are both used to communicate with other devices, and the processor is used to execute the instructions stored in the memory to cause the electronic device to perform the method as described in any of the foregoing.

[0020] A third aspect of this application provides a computer-readable storage medium storing instructions that, when executed, perform the method described in any of the preceding descriptions.

[0021] A fourth aspect of this application provides a computer program product that, when run on an electronic device, causes the electronic device to perform the method as described in any of the preceding claims.

[0022] In summary, one or more technical solutions provided in the embodiments of this application have at least the following technical effects or advantages: By constructing a localized feature library to perform morphological correction on general weather forecast data, meteorological input bias is reduced from the data source. Subsequently, an initial forecast is generated by combining a machine learning model with real-time operational data, and the forecast curve is dynamically corrected using a closed-loop bias correction mechanism based on the latest measured power, thereby effectively improving the adaptability and short-term accuracy of the forecast. In addition, the solution also designs a performance monitoring-based triggering mechanism and a hierarchical incremental update strategy, which enables the model to continuously track changes in the power plant status at a low cost while protecting core feature knowledge, thus achieving long-term, robust, and optimized operation of the forecast system. Attached Figure Description

[0023] Figure 1 This is a schematic diagram of the system architecture of an embodiment of a photovoltaic power prediction method based on historical data and weather, which applies the present application; Figure 2 This is a flowchart illustrating a photovoltaic power prediction method based on historical data and weather disclosed in an embodiment of this application. Figure 3 This is a schematic diagram of the structure of an electronic device disclosed in an embodiment of this application.

[0024] Explanation of reference numerals in the attached figures: 100, System architecture; 101, First terminal device; 102, Second terminal device; 103, Third terminal device; 104, Network; 105, Server; 301, Processor; 302, Communication bus; 303, User interface; 304, Network interface; 305, Memory. Detailed Implementation

[0025] To enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments.

[0026] In the description of the embodiments of this application, the words "for example" or "for instance" are used to indicate examples, illustrations, or explanations. Any embodiment or design that is described as "for example" or "for instance" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design options. Rather, the use of the words "for example" or "for instance" is intended to present the relevant concepts in a specific manner.

[0027] In the description of the embodiments of this application, the term "multiple" means two or more. For example, multiple systems means two or more systems, and multiple screen terminals means two or more screen terminals. Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the indicated technical features. Thus, a feature defined with "first" or "second" may explicitly or implicitly include one or more of that feature. The terms "comprising," "including," "having," and variations thereof all mean "including but not limited to," unless otherwise specifically emphasized.

[0028] Figure 1 This is a schematic diagram of the system architecture of an embodiment of a photovoltaic power prediction method based on historical data and weather, which applies the present application.

[0029] like Figure 1 As shown, system architecture 100 may include terminal devices 101, 102, and 103, a network 104, and a server 105. Network 104 serves as the medium for providing communication links between terminal devices 101, 102, and 103 and server 105. Network 104 may include various connection types, such as wired or wireless communication links, or fiber optic cables, etc.

[0030] Users can use terminal devices 101, 102, and 103 to interact with server 105 via network 104 to receive or send messages, etc. Various communication client applications can be installed on terminal devices 101, 102, and 103, such as model training applications, video recognition applications, web browser applications, social platform software, etc.

[0031] Terminal devices 101, 102, and 103 can be either hardware or software. When terminal devices 101, 102, and 103 are hardware, they can be various electronic devices with displays, including but not limited to smartphones, tablets, e-book readers, MP3 (Moving Picture Experts Group Audio Layer III) players, MP4 (Moving Picture Experts Group Audio Layer IV) players, laptops, and desktop computers, etc. When terminal devices 101, 102, and 103 are software, they can be installed in the aforementioned electronic devices. They can be implemented as multiple software programs or software modules (e.g., multiple software programs or software modules used to provide distributed services) or as a single software program or software module. No specific limitations are imposed here.

[0032] This embodiment discloses a photovoltaic power prediction method based on historical data and weather conditions. Figure 2 This is a flowchart illustrating a photovoltaic power prediction method based on historical data and weather disclosed in an embodiment of this application. Figure 2 As shown, the method includes the following steps: S201. Obtain historical irradiance data, historical power output data, and historical weather data of the target photovoltaic power station; In this embodiment, this step is the data preparation stage before performing power prediction. Its core task is to collect and integrate the operation and environmental records of the target photovoltaic power plant over a period of time from one or more data sources, providing a foundation for subsequent construction of a localized feature library and training models. Specifically, the system needs to acquire three key types of historical data. The first type is historical irradiance data, which typically refers to the time series of solar radiation intensity measured by irradiance meters (e.g., total radiation meters or direct / scattered radiation meters) deployed within the photovoltaic power plant, usually in watts per square meter (W / m²), reflecting the most direct energy input reaching the surface of the photovoltaic modules. The second type is historical output data, which is the time series of actual power generation at the same AC-side grid-connected point recorded by the photovoltaic power plant through its Supervisory Control and Data Acquisition (SCADA) system, usually in kilowatts (kW) or megawatts (MW), representing the final output performance of the power plant under real operating conditions. The third category is historical weather data, which includes other relevant meteorological parameters besides irradiance, such as but not limited to ambient temperature, relative humidity, wind speed, wind direction, atmospheric pressure, and cloud cover. This data provides a richer environmental context for understanding and modeling photovoltaic power output.

[0033] To ensure the comprehensiveness and representativeness of the data, the three types of data acquired should have a sufficiently long time span to cover various weather types and seasonal variations. For example, it can be configured to acquire all historical data of the target power plant since its commissioning, or at least data from the most recent 1 to 5 years. In terms of specific technical implementation, the methods for acquiring this data can be flexible and diverse, including but not limited to the following: The first method is based on API (Application Programming Interface) or database retrieval. The server deploying the forecasting method acts as a client, actively calling the API provided by the photovoltaic power plant's SCADA system or historical database at a preset frequency (e.g., early morning every day) via a secure network connection (such as a VPN), or directly executing SQL queries to batch download data files within the required time range. The second method is based on data push. A data agent program is deployed on the photovoltaic power plant side. This program is responsible for collecting data from various systems within the plant and actively pushing data packets to the server's designated receiving endpoint according to agreed protocols (such as MQTT (Message Queuing Telemetry Transport)) and formats (such as JSON (JavaScript Object Notation)). After acquiring the raw data, the system will also perform necessary data preprocessing, the core of which is to perform timestamp alignment to ensure that the irradiance, power output and weather data can correspond one-to-one at each sampling point (for example, one point every 15 minutes), forming a unified formatted multidimensional time series record, laying the foundation for subsequent analysis.

[0034] S202. By classifying and processing the historical irradiance data, multiple daily irradiance variation curves representing different weather patterns are generated, and a localized irradiance feature library is constructed based on the daily irradiance variation curves, the historical power output data, and the historical weather data. In this embodiment, to extract typical weather patterns in the target power station's area from massive historical data, the system first performs standardized preprocessing on the acquired historical radiation data. This includes segmenting by natural day, standardizing time axis sampling, and normalizing amplitude, resulting in a series of daily radiation data sequences that only reflect the "morphology" of intraday light changes. Next, the system applies a pre-defined unsupervised clustering algorithm (such as K-means clustering or DBSCAN (Density-Based Spatial Clustering of Applications with Noise)) to automatically classify these morphological sequences, grouping sequences with similar shapes into the same data cluster. For each formed data cluster, by calculating the centroid (i.e., the mean curve) of all sequences within the cluster, a standard daily radiation variation curve representing that type of weather pattern is generated. Finally, to enrich the model's content, the system identifies all historical dates constituting the cluster and extracts the corresponding average power output characteristics (such as average daily power generation efficiency) and average meteorological conditions (such as average daily temperature and humidity) from the original historical power output data and historical weather data, quantifying them into historical power output feature vectors and historical weather feature vectors. By associating each daily irradiance variation curve with its corresponding historical power output and historical weather feature vectors and storing them in a structured manner, the localized irradiance feature library is constructed.

[0035] Optionally, the step of classifying the historical irradiance data to generate multiple daily irradiance variation curves representing different weather patterns, and constructing a localized irradiance feature library based on the daily irradiance variation curves, the historical power output data, and the historical weather data, includes: sequentially performing daily segmentation, time axis alignment, and amplitude normalization on the historical irradiance data to obtain multiple daily irradiance data sequences; based on the data distribution characteristics of the multiple daily irradiance data sequences, applying a preset unsupervised clustering algorithm to divide the multiple daily irradiance data sequences into multiple data clusters; and for each data cluster, calculating the centroid of all daily irradiance data sequences within the data cluster. The process involves generating corresponding solar radiation variation curves; for each data cluster, extracting time-corresponding solar radiation data sequences associated with the data cluster from the historical power output data, and calculating historical power output feature vectors based on these power output records; for each data cluster, extracting time-corresponding weather records from the historical weather data, and calculating historical weather feature vectors based on these weather records; and associating and storing each solar radiation variation curve, the corresponding historical power output feature vector, and the corresponding historical weather feature vector to construct the localized radiation feature library.

[0036] Specifically, a series of standardized preprocessing steps are performed on the acquired continuous historical irradiance data to obtain multiple daily irradiance data sequences suitable for pattern recognition. This process includes the following steps: (1) Segmentation by natural day: Based on the local time of the target photovoltaic power station, the long sequence data spanning multiple days is segmented according to 00:00 to 23:59 (or sunrise to sunset) of each day to form data segments in units of "days". (2) Time axis alignment: In order to make all daily sequences have the same dimension and comparability, the system resamples each data segment to a uniform time resolution. For example, if the sampling point is uniformly set to one sampling point every 15 minutes, then there are 96 data points for 24 hours in a day. For missing points in the original data, methods such as linear interpolation, polynomial interpolation or spline interpolation can be used to fill them. (3) Amplitude normalization: In order to eliminate the interference of differences in total solar energy caused by different seasons and dates on the curve shape analysis, the system performs normalization processing on each aligned daily irradiance sequence. A preferred implementation is to use maximum-minimum normalization, which divides each irradiance value in the sequence by the maximum irradiance value of that day, so that all processed sequence amplitudes fall within the interval, thereby highlighting their relative change trend.

[0037] Furthermore, based on the data distribution characteristics of the multiple solar irradiance data sequences obtained after preprocessing, a pre-defined unsupervised clustering algorithm is applied to automatically divide these sequences into multiple data clusters. The purpose of this step is to automatically group solar irradiance sequences with similar variation patterns together, with each cluster representing a typical local weather pattern (such as clear skies, foggy mornings turning sunny in the afternoon, or cloudy skies with fluctuations throughout the day). The pre-defined unsupervised clustering algorithm may include, but is not limited to: the first scheme is the K-means clustering algorithm, in which the number of clusters K needs to be pre-set (which can be determined by, for example, the elbow rule or silhouette coefficient), and the algorithm iteratively calculates to assign each sequence to the cluster represented by its nearest centroid. The second scheme is the DBSCAN algorithm, which does not require a pre-set number of clusters, can automatically discover clusters of arbitrary shapes based on the density of data points, and can effectively identify "noise" or abnormal days that do not belong to any typical pattern, offering better flexibility.

[0038] Furthermore, for each data cluster formed in the previous step, the system generates a representative solar radiation variation curve corresponding to that cluster by calculating the centroid of all solar radiation data sequences within that cluster. Specifically, calculating the centroid involves taking the arithmetic mean of all solar radiation data sequences within the cluster at each corresponding time point. For example, if the time resolution is 15 minutes, each sequence is a 96-dimensional vector, and the calculated centroid is also a 96-dimensional vector, where the value of its i-th dimension is equal to the average value of all sequences within the cluster at the i-th time point. This generated centroid curve is smooth and typical in shape, and can be regarded as a standard template or ideal prototype of this weather model.

[0039] Furthermore, to assign corresponding power generation performance characteristics to each weather pattern (i.e., each data cluster), the system extracts the temporally corresponding power output records of all daily irradiance data sequences associated with that cluster from the original historical power output data for each data cluster. Then, based on these extracted power output records, a historical power output feature vector for that cluster is calculated. This vector is a set of multiple statistical indicators, specifically including but not limited to: the average daily total power generation under that weather pattern, the average power generation efficiency (i.e., the ratio of daily total power generation to daily total irradiance), the average capacity factor (i.e., the ratio of actual power generation to theoretical maximum power generation), and the volatility of the power output curve, etc. This feature vector quantitatively describes the typical power generation performance of a photovoltaic power plant under that type of weather pattern.

[0040] Similarly, to characterize the environmental conditions under each weather pattern, the system extracts temporally corresponding weather records (such as temperature, humidity, wind speed, etc.) from the original historical weather data for each data cluster. Based on these weather records, a historical weather feature vector is calculated. Similar to the output feature vector, this vector is also a set of statistical values, which may include, but are not limited to: daily average temperature, daily maximum / minimum temperature, daily average relative humidity, prevailing wind direction, and average wind speed under this weather pattern. This feature vector depicts the most common combinations of external environmental parameters when such radiation variation patterns occur.

[0041] Furthermore, the system associates and stores the elements generated in the aforementioned steps to formally construct the localized irradiance feature library. Specifically, the storage method could be as follows: A table could be created in a relational database, where each row represents a weather pattern (a cluster), and the fields (columns) would store the cluster ID, the daily irradiance variation curve representing that cluster (which can be stored as a serialized string or a pointer to a file), the corresponding historical power output feature vector, and the corresponding historical weather feature vector. Alternatively, a NoSQL database could be used, storing all the information for each pattern (curve, power output feature, weather feature) as a complete JSON document. This constructed localized irradiance feature library is equivalent to a structured, multi-dimensional knowledge base, where each record fully defines a "local weather pattern - power plant response - environmental background" relationship, providing a solid data foundation for subsequent accurate forecasting.

[0042] S203. Based on the preset prediction start time and prediction duration, obtain the numerical weather forecast data of the target photovoltaic power station; In this embodiment, this step is the input stage of the forecasting process. Its purpose is to obtain key external conditions affecting the power generation capacity of the photovoltaic power station over a future period, namely, numerical weather prediction (NWP) data from professional meteorological agencies. The triggering condition for executing this step is a preset, user-configurable forecast start time and forecast duration within the system. For example, the system can be configured as a daily automatic task, starting at 6:00 AM each day to obtain forecast data for the current day and the next 48 hours. Here, 6:00 AM is the forecast start time, and 48 hours is the forecast duration. These configurations ensure that the forecasting task can obtain the latest weather forecast information on a regular and automatic basis.

[0043] To ensure forecast accuracy, the NWP data acquired by the system must be high-precision forecast data specific to the geographical coordinates (latitude and longitude) of the target photovoltaic power station. In terms of specific technical implementation, acquiring this data can be achieved through at least two methods: The first method is real-time invocation based on an application programming interface (API), which is also a preferred implementation. The forecasting system, acting as a client, initiates an HTTP request with authentication credentials to a professional commercial meteorological service provider (such as IBM The Weather Company, AccuWeather, etc.) or the National Meteorological Public Data Service Platform via the internet. This request precisely specifies the latitude and longitude coordinates of the target photovoltaic power station, the required meteorological elements, and the forecast time range. The second method is batch download based on file transfer protocols or HTTP. This method is commonly used to interface with publicly released model data from national meteorological centers. The system periodically logs into a designated server, downloads gridded forecast data files (e.g., GRIB or NetCDF format) covering the area where the power station is located, corresponding to the latest model's start time, and then extracts the grid point data closest to the power station's location using a local parsing program. Regardless of the approach used, the acquired NWP data must include irradiance forecasts crucial for photovoltaic output prediction, such as global horizontal irradiance, and preferably include direct normal irradiance and diffuse horizontal irradiance. In addition, other auxiliary meteorological elements should be acquired, such as, but not limited to, ambient temperature, relative humidity, wind speed, wind direction, and cloud cover. After acquiring these raw forecast data, the system will perform necessary preprocessing, such as interpolating the temporal resolution (e.g., processing hourly forecast data into 15-minute intervals using linear or spline interpolation) to match the time granularity required by subsequent prediction models.

[0044] S204. In the localized irradiance feature library, retrieve the target daily irradiance variation curve that matches the numerical weather forecast data, and use the target daily irradiance variation curve to correct the numerical weather forecast data to generate target meteorological data. Specifically, the system first calculates a predicted weather feature vector based on the acquired NWP data (including forecast values ​​such as temperature, humidity, and irradiance) for the future forecast period. The vector's dimension and physical meaning are consistent with historical weather feature vectors stored in the localized irradiance feature library. Then, the system compares this predicted weather feature vector with each historical weather feature vector stored in the feature library to retrieve the best-matching historical weather pattern. Various techniques can be used to calculate the similarity, such as, but not limited to: one approach is to calculate the Euclidean distance between the two vectors and select the weather pattern corresponding to the vector with the smallest distance; another approach is to calculate the cosine similarity and select the pattern with the highest similarity score. Once the best-matching historical weather pattern is determined, the system extracts the associated normalized diurnal irradiance variation curve as the target diurnal irradiance variation curve. Next, the system performs a correction operation, the core of which is a morphological reshaping method: the system first calculates the total irradiance of the original NWP data irradiance forecast sequence over the entire forecast day (i.e., integrating the forecast curve), and then uses this scalar value representing the total energy level to proportionally scale the target daily irradiance variation curve retrieved in the previous step, which only represents the shape. The new irradiance sequence generated by this method retains the macroscopic accuracy of the daily total energy level predicted by NWP, and is endowed with a more realistic intraday fluctuation pattern derived from local historical statistics. Finally, this corrected daily irradiance sequence is recombine with other unmodified auxiliary meteorological element forecast values ​​(such as temperature, humidity, etc.) in the NWP data to form a high-quality target meteorological data set as input for subsequent prediction models.

[0045] Optionally, the step of retrieving a target daily irradiance variation curve that matches the numerical weather forecast data from the localized irradiance feature library, and using the target daily irradiance variation curve to correct the numerical weather forecast data to generate target meteorological data, includes: extracting weather features from the numerical weather forecast data; constructing a weather feature vector to be matched based on the weather features; for each daily irradiance variation curve in the localized irradiance feature library, calculating the Euclidean distance between the weather feature vector to be matched and the historical weather feature vector associated with the daily irradiance variation curve to obtain the dissimilarity score corresponding to the daily irradiance variation curve; determining the target dissimilarity score with the smallest value among all the dissimilarity scores, and determining the daily irradiance variation curve corresponding to the target dissimilarity score as the target daily irradiance variation curve; fusing the irradiance prediction value in the numerical weather forecast data with the target daily irradiance variation curve to obtain a corrected irradiance prediction value; and replacing the irradiance prediction value in the numerical weather forecast data with the corrected irradiance prediction value to generate the target meteorological data.

[0046] Specifically, to achieve automated comparison later, the system extracts key meteorological indicators from numerical weather forecast data and constructs a standardized weather feature vector to be matched. For example, the system can calculate multiple statistics for the forecast day, such as the average temperature, total precipitation, average wind speed, average cloud cover, and predicted total daily radiation, and arrange these values ​​according to a preset dimensional order to form a multi-dimensional vector. This vector serves as a quantified weather identifier, enabling objective comparison of the forecast day's weather conditions with historical data within a unified framework.

[0047] Furthermore, the system quantifies the similarity between the predicted day and the historical day's weather patterns by calculating the Euclidean distance. Specifically, the system iterates through each record in the localized irradiance feature library, extracts its associated historical weather feature vector, and then calculates the straight-line distance between this historical vector and the weather feature vector to be matched generated in the previous step in the multidimensional feature space. The result of this distance calculation is defined as the dissimilarity score of the corresponding daily irradiance variation curve. Generally, a smaller score value means that the weather patterns represented by the two are more similar. Those skilled in the art can also choose other distance measurement methods, such as Manhattan distance or Mahalanobis distance, depending on the data characteristics; these alternatives are all included within the scope of the inventive concept of this application.

[0048] Furthermore, the system filters through all calculated dissimilarity scores to determine the globally optimal match. This process compares all scores and finds the minimum value, which is then determined as the target dissimilarity score. The historical daily irradiance variation curve corresponding to this target dissimilarity score is selected as the best reference template for revising the current forecast, because it originates from an actual day with weather conditions most similar to the predicted day.

[0049] Furthermore, the system performs a fusion calculation, combining the total energy prediction from numerical weather prediction data with the fine-grained shape of the target-day irradiance variation curve to obtain a corrected irradiance prediction. A specific fusion method is as follows: first, the original irradiance time series from the numerical weather prediction is summed or integrated over the entire prediction day to obtain a predicted total irradiance. Then, this total irradiance is used as a scaling factor and multiplied by the already normalized target-day irradiance variation curve. The resulting new time series is the corrected irradiance prediction, which maintains consistency with the original prediction in total energy but reproduces the true patterns of the most similar historical days in terms of intraday fluctuation details.

[0050] Furthermore, the system generates target meteorological data for subsequent steps through a data replacement operation. In this step, the system completely replaces the irradiance prediction portion of the original numerical weather prediction data with the corrected irradiance prediction value sequence generated in the previous process. Simultaneously, the forecast values ​​of other meteorological elements in the original data, such as temperature, humidity, air pressure, and wind speed, are fully preserved. The dataset formed after this combination is the target meteorological data, an optimized hybrid data that provides a more accurate and reliable input for subsequent power prediction models.

[0051] Optionally, the step of fusing the predicted irradiance value in the numerical weather forecast data with the target day's irradiance variation curve to obtain a corrected predicted irradiance value includes: integrating the predicted irradiance value along the forecast duration to obtain a predicted total daily irradiance value; integrating the target day's irradiance variation curve along the forecast duration to obtain a reference day's total irradiance; calculating an amplitude adjustment coefficient based on the predicted total daily irradiance value and the reference day's total irradiance; multiplying the amplitude adjustment coefficient by the target day's irradiance variation curve, and determining the multiplication result as the corrected predicted irradiance value.

[0052] Specifically, the system obtains the predicted daily total irradiance by integrating the irradiance forecasts provided by numerical weather prediction over their forecast duration. In essence, the irradiance forecasts provided by numerical weather prediction are typically a time series, for example, data points every 15 minutes or 1 hour within the forecast day. The system sums these discrete irradiance data points over their corresponding time intervals; the summation result is a macroscopic prediction of the total solar energy for the entire forecast day. This predicted total daily irradiance represents the numerical weather prediction model's assessment of the total energy output under the combined influence of factors such as cloud cover and atmospheric transparency on that day.

[0053] Furthermore, the system employs a similar processing method to the previous step, integrating the selected target day's irradiance variation curve over the prediction duration to obtain a baseline day's total irradiance. The target day's irradiance variation curve is also a time series recording the actual irradiance variation on a historical day. By summing all its data points, the actual total irradiance energy for that historical day can be obtained. This baseline day's total irradiance serves as the reference for subsequent amplitude adjustments, representing the original energy scale under the selected similar weather pattern.

[0054] Furthermore, the system calculates a key amplitude adjustment coefficient based on the predicted daily total irradiance calculated in the first two steps and the total irradiance of the reference day. This coefficient is preferably calculated by dividing the predicted daily total irradiance by the total irradiance of the reference day. This dimensionless coefficient quantifies the proportional relationship between the total energy of the predicted day and the total energy of the selected historical template day. For example, if the coefficient is greater than 1, it indicates that the total irradiance of the predicted day is higher than that of the historical day; conversely, it is lower. This coefficient is the core of achieving energy conservation correction.

[0055] Furthermore, the system multiplies the calculated amplitude adjustment coefficient with the target daily irradiance variation curve, and uses the result as the final corrected irradiance prediction. Specifically, this multiplication operation involves multiplying the single scalar value of the amplitude adjustment coefficient by each data point in the time series of the target daily irradiance variation curve. Through this operation, the overall shape of the historical curve—including intraday fluctuations, peak times, and the intensity of fluctuations—is fully preserved, while its overall amplitude is precisely scaled. This ensures that the total integral value of the scaled curve, i.e., the total energy, is exactly equal to the daily total irradiance prediction from numerical weather prediction. The resulting corrected irradiance prediction thus possesses both the details of the historical true shape and conforms to the macroscopic energy prediction of numerical weather prediction, achieving an effective fusion of the two.

[0056] S205. Obtain the real-time power plant operation data of the target photovoltaic power plant, and input the target meteorological data and the real-time power plant operation data into a preset machine learning prediction model to generate an initial power prediction curve. In this embodiment, the system first performs the operation of acquiring real-time power plant operation data of the target photovoltaic power plant. Real-time operation data here refers to various dynamic information reflecting the current and recent operating status of the power plant. A preferred approach is to establish a communication connection through a SCADA system or directly with key equipment such as combiner boxes and inverters within the power plant, collecting data at a preset time frequency, such as every 1 minute, 5 minutes, or 15 minutes. The collected data may include, but is not limited to: the actual output power of the power plant's grid connection point, the DC and AC side power of each inverter, the voltage and current of the photovoltaic array, the surface temperature of the modules, and local meteorological data measured by the power plant's internal environmental monitoring unit, such as backsheet temperature and actual irradiance. This real-time data provides the prediction model with an accurate benchmark regarding the power plant's current health status, operating efficiency, and initial output level.

[0057] Furthermore, the system uses the target meteorological data and the newly acquired real-time power plant operation data as input features, feeding them into a pre-defined machine learning prediction model to generate an initial power prediction curve. The target meteorological data provides an accurate prediction of external environmental conditions for a future period, such as the next 4 to 72 hours; while the real-time power plant operation data provides the model with the initial state of the system at the prediction start time. The combination of these two forms a complete time-series prediction input vector. The pre-defined machine learning prediction model can be any model suitable for handling time series and nonlinear mapping problems, and it has been trained using massive amounts of historical meteorological data and corresponding historical power data before deployment. For example, the model can be a recurrent neural network, such as a Long Short-Term Memory network or a gated recurrent unit, which excels at capturing temporal dependencies in the data. As another implementation, the model can also be an ensemble model based on gradient boosting decision trees, such as XGBoost or LightGBM, which perform well when processing tabular feature data. After receiving the input features, the model performs inference calculations and outputs a time series, i.e., the initial power prediction curve. Specifically, this curve represents a power prediction value given at a specific time granularity, such as every 15 minutes, within a future prediction period. This curve is called the initial prediction curve because it is the raw prediction result directly output by the machine learning model, without subsequent rule corrections.

[0058] S206. Obtain the latest measured power data of the target photovoltaic power station, determine the prediction deviation based on the latest measured power data and the initial power prediction curve, and correct the initial power prediction curve based on the prediction deviation to obtain the final power prediction curve.

[0059] In this embodiment, to further improve the real-time accuracy of the prediction, the system executes a correction process based on real-time error. First, the system obtains the latest measured power data up to the current prediction start time from the SCADA system of the photovoltaic power plant. Then, the system compares this latest measured power value with the predicted value at the corresponding time on the initial power prediction curve, calculating an immediate prediction deviation by subtracting the two. This deviation reflects the prediction performance of the machine learning model at the current time. Finally, the system performs overall correction on the initial power prediction curve based on this prediction deviation to obtain the final power prediction curve. One specific correction method is to use this prediction deviation as a constant correction amount, superimposed on each future time point of the initial power prediction curve; this is an error persistence correction strategy. As another preferred option, this correction amount can gradually decay over time, for example, by using an exponential decay function for weighting, because the impact of prediction errors typically weakens over time. In this way, the generated final power prediction curve retains the machine learning model's trend judgment of future weather changes while being anchored to the latest actual operating point of the power plant, thereby improving the reliability of ultra-short-term predictions.

[0060] Optionally, the step of obtaining the latest measured power data of the target photovoltaic power station, determining the prediction deviation based on the latest measured power data and the initial power prediction curve, and correcting the initial power prediction curve based on the prediction deviation to obtain the final power prediction curve includes: taking the measured power sequence of the target photovoltaic power station within a preset backtracking time before the start time of the prediction period as the latest measured power data; extracting historical prediction power sequences corresponding to the latest measured power data at each time point within the preset backtracking time from the initial power prediction curve; calculating the average power deviation within the preset backtracking time based on the latest measured power data and the historical prediction power sequences, and determining the average power deviation as the prediction deviation; inputting the prediction deviation into a preset deviation attenuation algorithm to calculate a correction amount sequence that attenuates along the prediction period; and superimposing the correction amount sequence with the initial power prediction curve to generate the final power prediction curve.

[0061] Specifically, the system acquires the latest measured power data of the target photovoltaic power station. Here, "latest" does not refer to a single moment's data point, but rather a sequence of measured power data within a preset backtracking period prior to the prediction start time. For example, if the current prediction task starts at 10:00 AM and the preset backtracking period is 1 hour, the system will collect and organize all measured power values ​​recorded at specific time intervals (e.g., every 5 minutes) between 9:00 AM and 10:00 AM, forming a time series containing multiple data points. The specific value of the preset backtracking period can be configured according to the scale of the power station and the rate of weather change; for example, it can be set to 30 minutes, 1 hour, or 2 hours to achieve a balance between response speed and stability.

[0062] Accordingly, to calculate the bias, the system needs to extract a historical predicted power sequence from the initial power prediction curve previously generated by the machine learning model, which corresponds exactly in time to the measured power sequence. In other words, since the initial power prediction curve was generated at an earlier time point, its prediction range will inevitably cover this "preset backtracking period". The system uses timestamp matching to accurately find all predicted values ​​on the initial power prediction curve within the time period from 9:00 AM to 10:00 AM, forming a historical predicted power sequence with the same length as the measured power sequence and corresponding to each time point.

[0063] Based on this, the system calculates the average power deviation within a preset backtracking period using the latest measured power data (i.e., the measured power sequence) and historical predicted power sequences, and determines this deviation as the prediction deviation for subsequent correction. Specifically, the system can calculate the difference between the measured value and the predicted value at each time point within the backtracking period, and then take an arithmetic average of all these differences. For example, it calculates the average value of [measured power (t) - predicted power (t)] between 9:00 and 10:00. The advantage of this is that averaging can smooth out measurement noise or short-term disturbances that may exist at a single time point, thus obtaining a more stable deviation value that better reflects the model's persistent, systematic overestimation or underestimation within that time period.

[0064] Furthermore, after obtaining the average power deviation, the system does not simply apply it directly to the entire future prediction curve. Instead, it inputs it into a preset deviation decay algorithm to calculate a correction sequence that gradually decreases along the prediction duration. This is based on a key technical insight: the prediction deviation at the current moment has a greater impact on the near future and a smaller impact on the distant future. The preset deviation decay algorithm can include, but is not limited to, an exponential decay model or a linear decay model. For example, a preferred implementation is to use an exponential decay model, where the generated correction is equal to the average power deviation at the start of the prediction and then gradually decreases to zero exponentially over time. The decay rate of this decay model is a configurable parameter that can be optimized based on historical prediction results.

[0065] Furthermore, the system overlays the generated correction sequence with the initial power prediction curve to generate the final power prediction curve. This overlay operation involves adding each correction value from the correction sequence to the predicted value at the corresponding future time point on the initial power prediction curve. Through this overlay operation, the final power prediction curve is effectively pulled back to a state closer to the actual current operation of the power plant in the early stages of prediction due to the larger correction amounts. In the later stages of prediction, it gradually reverts to the long-term trend judgment of the original machine learning model, thus achieving a dynamic, smooth, and physically consistent fine-grained correction of the initial prediction results.

[0066] The method further includes: continuously acquiring measured power values ​​corresponding to the time points of the final power prediction curve from the data acquisition system of the target photovoltaic power station within the time interval of the prediction duration to obtain subsequent measured power data; determining whether the prediction error continuously exceeds a preset performance degradation threshold within a preset evaluation time window by calculating the prediction error between the subsequent measured power data and the final power prediction curve; when it is determined that the prediction error continuously exceeds the performance degradation threshold, triggering an online model update instruction, and jointly constructing an incremental training sample set based on the subsequent measured power data acquired within the preset evaluation time window and the target meteorological data corresponding to the time range of the subsequent measured power data.

[0067] In some embodiments of this application, to achieve continuous monitoring and adaptive optimization of the prediction model's performance, the method also includes a closed-loop online update process. After generating and outputting the final power prediction curve, the system does not terminate its operation. Instead, it continuously acquires the actual grid-connected power value of the target photovoltaic power plant from its data acquisition system (such as a SCADA system) in real time, at the same time granularity as the final power prediction curve (e.g., every 15 minutes), throughout the entire prediction duration (e.g., the next 4 hours or 24 hours). These power data, which actually occur and are measured after the prediction task begins, collectively constitute the subsequent measured power data sequence.

[0068] Further, the system enters a performance evaluation phase. It calculates the prediction error by comparing the subsequent measured power data with the corresponding values ​​on the final power prediction curve in real time. A preferred error calculation method is to calculate the relative error, i.e., |measured power - predicted power| / power plant installed capacity, to eliminate the impact of capacity differences. More importantly, the system needs to determine whether the prediction error has continuously exceeded a preset performance degradation threshold within a preset evaluation time window. This preset evaluation time window is a sliding time window, such as the most recent 2 hours. "Continuously exceeding" is a key judgment criterion; it does not refer to a single error exceeding the threshold, but rather a trend of deterioration. Specifically, this can be achieved by calculating the average or root mean square error of all error points within the evaluation time window and checking if it exceeds the threshold; or, a more robust approach is to statistically analyze the proportion of points where the error value exceeds the threshold within the evaluation time window, and only when this proportion reaches a preset trigger ratio (e.g., 80%) is it considered a continuous exceedance. The performance degradation threshold and trigger ratio are parameters that can be flexibly configured based on the power grid's assessment standards or operational experience.

[0069] When the system determines, based on the aforementioned rules, that the prediction error has consistently exceeded the performance degradation threshold, it indicates a significant decline in the current model's predictive performance, rendering it unable to meet accuracy requirements. At this point, the system automatically triggers an online model update command. This command signifies a shift from passive performance monitoring to proactive model correction. Following this, the system begins constructing an incremental training sample set for model updates. Specifically, the system precisely compiles all subsequent measured power data acquired within the preset evaluation time window that triggered the update; this data serves as the label for the incremental sample set. Simultaneously, the system retrieves target meteorological data (i.e., the NWp data used for prediction at that time) from the historical database within the time range completely corresponding to this time window; this meteorological data constitutes the features of the incremental sample set. By pairing these features and labels one-to-one according to time points, a high-quality, high-value, and highly targeted incremental training sample set is constructed, providing precise input for subsequent model fine-tuning or incremental learning.

[0070] Optionally, the method further includes: in response to the online model update instruction, deconstructing the machine learning prediction model into a shared feature layer and a task-specific layer; when updating the machine learning prediction model, configuring the network parameters of the shared feature layer to a gradient-no-update state; based on the incremental training sample set, performing gradient calculation and weight update only on the network parameters of the task-specific layer using the backpropagation algorithm to obtain the updated task-specific layer; and combining the updated task-specific layer with the shared feature layer whose network parameters remain unchanged to generate the updated machine learning prediction model.

[0071] To achieve efficient and low-cost online model updates, this application also provides a preferred implementation scheme based on the idea of ​​transfer learning. When the system responds to an online model update command, it first logically deconstructs the currently used machine learning prediction model into two core parts: a shared feature layer and a task-specific layer. In this embodiment, the machine learning prediction model can be a deep neural network. The shared feature layer typically corresponds to the part of the network near the input, such as multiple convolutional or recurrent layers, and its function is to extract general features from the raw meteorological data. These layers, through large-scale offline training, have learned to recognize universal patterns such as trends in light intensity changes and cloud movement patterns. The task-specific layer corresponds to the part of the network near the output, such as a fully connected layer, and its function is to map the extracted high-level features to specific power prediction values. This structural deconstruction of the model lays the foundation for subsequent implementation of an update strategy called fine-tuning.

[0072] Furthermore, when preparing for a model update, the system performs a crucial operation: configuring all network parameters of the shared feature layer—namely, weights and biases—to a gradient-no-update state. This operation is commonly referred to as freezing in the deep learning field. Technically, this can be achieved by setting the trainable properties of these layers to dummy values ​​in mainstream deep learning frameworks (such as TensorFlow or PyTorch). The core purpose of this is to protect the valuable knowledge learned by the model during long-term offline training, which possesses high generalization ability. By freezing these low-level parameters, a phenomenon known as catastrophic forgetting—where the model compromises its understanding of long-term patterns to adapt to short-term data—is effectively prevented when training with small batches of incremental samples. This ensures the stability and robustness of model updates.

[0073] Furthermore, the system utilizes the incremental training sample set constructed in the aforementioned steps to perform gradient calculations and weight updates specifically for the network parameters of the task-specific layer using the backpropagation algorithm. Since the shared feature layer is frozen, in each training iteration, although data and error gradients flow through the entire network, only the parameters of the task-specific layer are adjusted based on the calculated gradients. This means the computational load of the update process is significantly reduced, as the number of parameters to be optimized represents only a small fraction of the entire model. This makes the online update process extremely fast and resource-efficient, fully meeting the technical requirements for rapid response and real-time iteration in production environments.

[0074] Furthermore, once the network parameters for a specific task layer are updated through incremental training, the update process is complete. The system logically combines this updated task-specific layer, which incorporates the latest learned knowledge, with the previously unchanged shared feature layer, to form a completely new and updated machine learning prediction model. This updated model retains its original powerful feature extraction capabilities while rapidly adapting to the latest changes in operating conditions that lead to recent performance degradation, such as dust accumulation on solar panels, component aging, or special weather patterns, through fine-tuning the top-level network. This updated model can be immediately deployed to the next prediction cycle, thus forming an intelligent prediction closed-loop system capable of continuous self-evolution and performance optimization.

[0075] It should be noted that the above embodiments of the apparatus are only illustrated by the division of the above functional modules. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the apparatus and method embodiments provided in the above embodiments belong to the same concept, and the specific implementation process can be found in the method embodiments, which will not be repeated here.

[0076] This embodiment also discloses an electronic device, as shown in the reference. Figure 3 The electronic device may include: at least one processor 301, at least one communication bus 302, user interface 303, network interface 304, and at least one memory 305.

[0077] The communication bus 302 is used to enable communication between these components.

[0078] The user interface 303 may include a display screen and a camera. Optionally, the user interface 303 may also include a standard wired interface and a wireless interface.

[0079] The network interface 304 may optionally include a standard wired interface or a wireless interface (such as a Wi-Fi interface).

[0080] The processor 301 may include one or more processing cores. The processor 301 connects to various parts of the server using various interfaces and lines, and performs various server functions and processes data by running or executing instructions, programs, code sets, or instruction sets stored in memory 305, and by calling data stored in memory 305. Optionally, the processor 301 may be implemented using at least one hardware form of Digital Signal Processing (DSP), Field-Programmable Gate Array (FPGA), or Programmable Logic Array (PLA). The processor 301 may integrate one or a combination of several of the following: Central Processing Unit (CPU), Graphics Processing Unit (GPU), and modem. The CPU primarily handles the operating system, user interface, and applications; the GPU is responsible for rendering and drawing the content required for display; and the modem handles wireless communication. It is understood that the modem may also not be integrated into the processor 301 and may be implemented as a separate chip.

[0081] The memory 305 may include random access memory (RAM) or read-only memory. Optionally, the memory 305 may include a non-transitory computer-readable storage medium. The memory 305 can be used to store instructions, programs, code, code sets, or instruction sets. The memory 305 may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for at least one function (such as touch function, sound playback function, image playback function, etc.), instructions for implementing the above-described method embodiments, etc.; the data storage area may store data involved in the above-described method embodiments, etc. Optionally, the memory 305 may also be at least one storage device located remotely from the aforementioned processor 301. Figure 3 As shown, the memory 305, which serves as a computer storage medium, may include an operating system, a network communication module, a user interface module, and an application program for photovoltaic power prediction based on historical data and weather.

[0082] exist Figure 3 In the electronic device shown, the user interface 303 is mainly used to provide an input interface for the user and to obtain the user input data; while the processor 301 can be used to call an application stored in the memory 305 that is a photovoltaic power prediction method based on historical data and weather. When executed by one or more processors 301, the electronic device executes one or more methods as described in the above embodiments.

[0083] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this application.

[0084] In the several embodiments provided in this application, it should be understood that the disclosed apparatus can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the shown or discussed mutual couplings or direct couplings or communication connections may be through some service interfaces; indirect couplings or communication connections between apparatuses or units may be electrical or other forms.

[0085] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0086] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage device (CMD). Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory 305 and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods in the various embodiments of this application. The aforementioned memory 305 includes various media capable of storing program code, such as a USB flash drive, external hard drive, magnetic disk, or optical disk.

[0087] The foregoing description is merely an exemplary embodiment of this disclosure and should not be construed as limiting the scope of this disclosure. Any equivalent changes and modifications made in accordance with the teachings of this disclosure shall still fall within the scope of this disclosure. Other embodiments of this disclosure will be readily apparent to those skilled in the art upon consideration of the disclosure in this specification. This application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not described in this disclosure. The specification and embodiments are to be considered exemplary only, and the scope and spirit of this disclosure are defined by the claims.

Claims

1. A photovoltaic power prediction method based on historical data and weather, characterized in that, Applied to a server, the method includes: Acquire historical irradiance data, historical power output data, and historical weather data of the target photovoltaic power station; By classifying and processing the historical irradiance data, multiple daily irradiance variation curves representing different weather patterns are generated, and a localized irradiance feature library is constructed based on the daily irradiance variation curves, the historical power output data, and the historical weather data. Numerical weather forecast data for the target photovoltaic power station is obtained based on the preset prediction start time and prediction duration. In the localized irradiance feature library, a target daily irradiance variation curve that matches the numerical weather forecast data is retrieved, and the numerical weather forecast data is corrected using the target daily irradiance variation curve to generate target meteorological data. The real-time power plant operation data of the target photovoltaic power plant is obtained, and the target meteorological data and the real-time power plant operation data are input into a preset machine learning prediction model to generate an initial power prediction curve. The latest measured power data of the target photovoltaic power station is obtained. Based on the latest measured power data and the initial power prediction curve, the prediction deviation is determined. The initial power prediction curve is then corrected based on the prediction deviation to obtain the final power prediction curve.

2. The method according to claim 1, characterized in that, The process involves classifying the historical irradiance data to generate multiple daily irradiance variation curves representing different weather patterns, and constructing a localized irradiance feature library based on the daily irradiance variation curves, the historical power output data, and the historical weather data, including: The historical irradiance data is sequentially processed by dividing it into natural days, aligning it with the time axis, and normalizing the amplitude to obtain multiple daily irradiance data sequences; Based on the data distribution characteristics of the multiple solar irradiance data sequences, a preset unsupervised clustering algorithm is applied to divide the multiple solar irradiance data sequences into multiple data clusters; For each data cluster, the corresponding solar radiation variation curve is generated by calculating the centroid of all solar radiation data sequences within the data cluster; For each data cluster, extract the output record corresponding to the solar irradiance data sequence associated with the data cluster in time from the historical output data, and calculate the historical output feature vector based on the output record; For each data cluster, weather records corresponding to the solar radiation data sequence associated with the data cluster in time are extracted from the historical weather data, and historical weather feature vectors are calculated based on the weather records. Each of the daily irradiance variation curves, the corresponding historical power output feature vectors, and the corresponding historical weather feature vectors are associated and stored to construct the localized irradiance feature library.

3. The method according to claim 1, characterized in that, The step of retrieving a target daily irradiance variation curve that matches the numerical weather forecast data from the localized irradiance feature library, and using the target daily irradiance variation curve to correct the numerical weather forecast data to generate target meteorological data includes: Weather features are extracted from the numerical weather forecast data, and a weather feature vector to be matched is constructed based on the weather features. For each diurnal radiation variation curve in the localized radiation feature library, the Euclidean distance between the weather feature vector to be matched and the historical weather feature vector associated with the diurnal radiation variation curve is calculated to obtain the dissimilarity score corresponding to the diurnal radiation variation curve. Among all the anisotropy scores, the target anisotropy score with the smallest value is determined, and the solar radiation variation curve corresponding to the target anisotropy score is determined as the target solar radiation variation curve; The irradiance prediction value in the numerical weather forecast data is fused with the target day irradiance variation curve to obtain the corrected irradiance prediction value. The corrected irradiance prediction value is used to replace the irradiance prediction value in the numerical weather forecast data to generate the target meteorological data.

4. The method according to claim 3, characterized in that, The step of fusing the predicted irradiance value from the numerical weather forecast data with the target day's irradiance variation curve to obtain a corrected predicted irradiance value includes: The predicted total daily irradiance is obtained by integrating the predicted irradiance value over the predicted duration. The total irradiance of the base day is obtained by integrating the target day's irradiance variation curve along the predicted duration. Calculate the amplitude adjustment coefficient based on the predicted daily total irradiance and the baseline daily total irradiance; The amplitude adjustment coefficient is multiplied by the target daily irradiance variation curve, and the result of the multiplication is determined as the corrected irradiance prediction value.

5. The method according to claim 1, characterized in that, The process of obtaining the latest measured power data of the target photovoltaic power station, determining the prediction deviation based on the latest measured power data and the initial power prediction curve, and correcting the initial power prediction curve based on the prediction deviation to obtain the final power prediction curve includes: The measured power sequence of the target photovoltaic power station within a preset backtracking time before the start time of the prediction period is taken as the latest measured power data; From the initial power prediction curve, extract the historical predicted power sequence corresponding to each time point within the preset backtracking time of the latest measured power data; Based on the latest measured power data and the historical predicted power sequence, the average power deviation within the preset backtracking time is calculated, and the average power deviation is determined as the prediction deviation. The prediction deviation is input into a preset deviation attenuation algorithm to calculate a correction sequence that attenuates along the prediction duration; The correction sequence is superimposed on the initial power prediction curve to generate the final power prediction curve.

6. The method according to claim 1, characterized in that, The method further includes: Within the time interval of the prediction duration, the measured power values ​​corresponding to the time points of the final power prediction curve are continuously acquired from the data acquisition system of the target photovoltaic power station to obtain subsequent measured power data. By calculating the prediction error between the subsequent measured power data and the final power prediction curve, it is determined whether the prediction error continues to exceed a preset performance degradation threshold within a preset evaluation time window. When it is determined that the prediction error continues to exceed the performance degradation threshold, an online model update instruction is triggered, and an incremental training sample set is jointly constructed based on the subsequent measured power data obtained within the preset evaluation time window and the target meteorological data corresponding to the time range of the subsequent measured power data.

7. The method according to claim 6, characterized in that, The method further includes: In response to the online update instruction for the model, the machine learning prediction model is deconstructed into a shared feature layer and a task-specific layer; When updating the machine learning prediction model, the network parameters of the shared feature layer are configured to not update the gradient. Based on the incremental training sample set, the gradient calculation and weight update of the network parameters of the specific task layer are performed only on the backpropagation algorithm to obtain the updated specific task layer. The updated task-specific layer is combined with the shared feature layer, which keeps the network parameters unchanged, to generate an updated machine learning prediction model.

8. An electronic device, characterized in that, The device includes a processor, a memory, a user interface, and a network interface. The memory is used to store instructions. The user interface and the network interface are both used to communicate with other devices. The processor is used to execute the instructions stored in the memory to cause the electronic device to perform the method as described in any one of claims 1-7.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores instructions that, when executed, perform the method as described in any one of claims 1-7.

10. A computer program product, characterized in that, When the computer program product is run on an electronic device, it causes the electronic device to perform the method as described in any one of claims 1-7.