Photovoltaic power generation few-sample transferable prediction method based on large language model
Through the method of migratory prediction of photovoltaic power generation based on large language model, the time series data is mapped to the natural language semantic space, and the problem of insufficient generalization in the data scarce scenario is solved, and efficient and migratory photovoltaic power generation prediction is achieved.
Patent Information
- Application Number
- CN202510529749.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-25
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-04-25
AI Technical Summary
The existing photovoltaic power generation prediction methods are insufficient in data scarce scenarios, and their feature expression capabilities are limited, making it difficult to adapt to the migration application scenarios of multiple regions and multiple stations.
The photovoltaic power generation with a small sample migration prediction method based on a large language model is used to map the time series data to the natural language semantic space through steps such as input segmentation, feature extraction, semantic alignment, prompt generation and output projection, and predict using the context understanding ability of the pre-trained large language model.
It significantly improves the migration capability and prediction accuracy of the photovoltaic power generation prediction model, reduces the dependence on the local data volume of the target site, and maintains high-fidelity prediction performance in extremely short-term data and cross-scenario migration scenarios.
Smart Images

Figure CN120045943A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to photovoltaic power generation prediction technology, and particularly to a few-shot transferable prediction method for photovoltaic power generation based on large language models. Background Art
[0002] As an important part of the new power system, photovoltaic power generation has advantages such as green and low-carbon, and rich resources. However, photovoltaic power generation has significant intermittency and randomness characteristics, and its output is easily affected by factors such as weather changes and environmental conditions, bringing significant challenges to power grid dispatching and system operation. In this context, in order to ensure the safe and stable operation of the power system, improving the prediction accuracy of photovoltaic output has become an important technical means to enhance the new energy consumption capacity and support flexible system regulation.
[0003] The core problem of photovoltaic power generation prediction lies in how to fully explore the deep relationship between multi-source heterogeneous data such as historical power data and meteorological information, so as to construct an accurate prediction model. Currently, the mainstream prediction methods can be divided into two categories: physical mechanism-based modeling methods and historical data-based statistical learning methods. With the rapid development of artificial intelligence technology, especially the advantages of deep learning in pattern recognition, feature extraction, etc. have gradually emerged, and data-driven deep learning methods have become an important research direction in the field of photovoltaic prediction. This type of method can efficiently process large-scale data, automatically extract complex non-linear features, significantly improve the prediction accuracy and model efficiency, and promote the development of prediction technology towards real-time and intelligent directions. Although deep learning models have been widely used in photovoltaic prediction, due to the wide geographical distribution of photovoltaic power stations and the large differences in data distribution among stations, especially in newly built or small power stations with limited data volume, the generalization ability and adaptability of traditional models still face great challenges.
[0004] To solve the problem of insufficient data and improve the model transfer performance, transfer learning technology has been gradually introduced into the field of photovoltaic power generation prediction, by leveraging the modeling experience of existing stations to support the prediction modeling of new stations. However, existing transfer methods still have certain limitations in dealing with source-target differences and improving transfer stability.
[0005] Specifically, existing photovoltaic power generation prediction methods often face the problem of decreased prediction accuracy or even failure when the historical data of the target power station is insufficient. Traditional models usually need to be retrained or tuned on each power station, with limited generalization ability and difficulty in adapting to the migration application scenarios of multiple regions and multiple power stations. Therefore, the most critical technical problem at present is how to construct a prediction framework with strong transferability, which can use the data of existing source power stations for generalization training and migrate to the target power station, still maintaining high prediction performance in the absence of local data, so as to improve the universality and practicality of photovoltaic power generation prediction. In addition, different photovoltaic power stations are limited by factors such as geographical location, weather conditions, and component types, and the data shows obvious heterogeneity. How to design a model structure with strong adaptability that can effectively integrate heterogeneous data from multiple source power stations and improve its feature expression ability is the key challenge for realizing generalization prediction. In addition, in the process of feature extraction of existing methods, they mostly rely on artificial construction or shallow neural networks, and it is difficult to fully capture spatio-temporal dependencies and complex non-linear relationships.
[0006] It should be noted that the information disclosed in the above background technology section is only used for understanding the background of the present application, and therefore may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention
[0007] The main object of the present invention is to overcome the defects existing in the above background technology and provide a few-shot transferable prediction method for photovoltaic power generation based on a large language model.
[0008] To achieve the above object, the present invention adopts the following technical solutions: A few-shot transferable prediction method for photovoltaic power generation based on a large language model, comprising the following steps: S1. Input segmentation: Normalize and block-cut the original photovoltaic time series data to generate an embedding module adapted to the input of the large language model; S2. Feature extraction: Dynamically allocate feature weights through high-order partial correlation analysis and gated residual network to extract non-linear correlation features of multi-dimensional data; S3. Semantic alignment: Map the time series features to the semantic space of the large language model corpus to complete natural language alignment; S4. Prompt generation: Generate a prompt prefix containing background, domain, instruction, and statistical information based on a preset prompt word template; S5. Natural language prediction: Input the data after semantic alignment and the prompt into the large language model with frozen parameters to output a natural language prediction result; S6. Output projection: Inversely parse the natural language response into a numerical power generation prediction value through linear projection.
[0009] Further, in step S1, the input segmentation specifically includes: Performing extreme value normalization on the original photovoltaic time series data; Chunking the normalized data at a preset stride to generate multiple embedding modules; Mapping the segmented data to a feature space matching the input dimension of the large language model through linear transformation.
[0010] Further, in step S2, the feature extraction specifically includes: Based on high-order partial correlation analysis, dynamically evaluating the association strength between variables and excluding the interference of multicollinearity; Adopting a gated residual network combined with a gating mechanism and residual connection to extract the non-linear features of multi-dimensional data; Through a two-layer feature selection mechanism, fusing the results of statistical analysis and deep learning to generate a weighted normalized feature vector.
[0011] Further, in step S3, the semantic alignment specifically includes: Reducing the dimensionality of the large language model corpus to reduce the complexity of matching calculations; Using a masked multi-head self-attention mechanism to match the time series data with the corpus tokens, ensuring that only historical information is focused on; Combining the token sequences of the power generation data and the external feature data to generate an input format adapted to the large language model.
[0012] Further, in step S4, the prompt template specifically includes: Data background information describing the geographical location, time range, and external features of the photovoltaic power station; Defining the typical time series patterns and domain knowledge of photovoltaic power generation; Determining the historical step and prediction step instructions for the prediction task; Providing the statistical features of the historical power generation data to enhance semantic expression.
[0013] Further, in step S6, the output projection specifically includes: Flattening the natural language sequence output by the large language model into a one-dimensional vector; Mapping the flattened vector to the target prediction dimension through linear transformation to restore the numerical result.
[0014] Further, during the prediction process, the large language model keeps the pre-trained parameters fixed and only optimizes the parameters of the input segmentation, feature extraction, semantic alignment, prompt generation, and output projection modules.
[0015] Further, the evaluation metrics for the prediction results include: Symmetric Mean Absolute Percentage Error, which measures the relative deviation between the predicted value and the true value; Normalized Root Mean Square Error, which quantifies the global accuracy of the prediction result.
[0016] Furthermore, the method is adapted to any one of the following data scenarios: Only using the very short-term historical data of the target photovoltaic power station for training; Performing cross-scenario transfer prediction based on data from similar power stations; Performing cross-scenario transfer prediction based on data from dissimilar power stations.
[0017] A computer program product includes a computer program, and when the computer program is executed by a processor, it implements the above-mentioned few-shot transferable prediction method for photovoltaic power generation based on a large language model.
[0018] The present invention has the following beneficial effects: The present invention proposes a few-shot transferable prediction method for photovoltaic power generation based on a large language model. By integrating the generalization ability of the large language model and transfer learning technology, it effectively solves the key problems of insufficient generalization and limited feature expression ability of traditional photovoltaic prediction models in data-scarce scenarios. The method of the present invention innovatively maps time series data to the natural language semantic space, utilizes the context understanding ability of the pre-trained large language model, combines the normalized block processing of data by the input segmentation layer, the high-order partial correlation analysis and gated residual network dynamic weight allocation of the feature extraction layer, the masked multi-head attention mechanism matching of the semantic alignment layer, the structured template guidance of the prompt generation layer, and the linear inverse mapping of the output projection layer to construct an end-to-end automated prediction framework. The present invention can achieve high-fidelity transfer prediction across power stations without fine-tuning the ontology parameters of the large language model, significantly reducing the dependence on the local data volume of the target power station. At the same time, through the feature modeling mechanism that combines statistics and deep learning, it fully excavates the spatio-temporal dependence and non-linear correlation features of multi-source heterogeneous data, and solves the limitations of traditional methods in artificial feature construction and shallow network expression. Experimental verification shows that the symmetric mean absolute percentage error (SMAPE) and normalized root mean square error (NRMSE) of this method in very short-term data, similar and dissimilar power station transfer scenarios are significantly better than mainstream comparison models, demonstrating its excellent prediction accuracy, stability and cross-scenario adaptation ability under few-shot conditions, providing an innovative solution for the intelligentization and generalization of photovoltaic power generation prediction technology.
[0019] Other beneficial effects in the embodiments of the present invention will be further described below. Description of the Drawings
[0020] Figure 1 The flowchart of a few-shot transferable prediction method for photovoltaic power generation based on a large language model according to an embodiment of the present invention. Detailed implementation manners
[0021] The following provides a detailed description of the implementation manners of the present invention. It should be emphasized that the following description is merely exemplary and not intended to limit the scope of the present invention and its applications.
[0022] In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the embodiments of the present invention, "a plurality of" means two or more unless otherwise specifically defined.
[0023] Refer to Figure 1 , the embodiments of the present invention provide a few-shot transferable prediction method for photovoltaic power generation based on a large language model, including the following steps: S1. Input segmentation: Normalize and block-cut the original photovoltaic time series data to generate an embedding module adapted to the input of the large language model.
[0024] In some embodiments, in step S1, the input segmentation specifically includes: performing extreme value normalization processing on the original photovoltaic time series data; block-cutting the normalized data at a preset stride to generate a plurality of embedding modules; and mapping the segmented data to a feature space matching the input dimension of the large language model through linear transformation.
[0025] S2. Feature extraction: Dynamically allocate feature weights through high-order partial correlation analysis and gated residual network to extract non-linear correlation features of multi-dimensional data.
[0026] In some embodiments, in step S2, the feature extraction specifically includes: dynamically evaluating the correlation strength between variables based on high-order partial correlation analysis to exclude the interference of multi-collinearity; using a gated residual network combined with a gating mechanism and residual connection to extract non-linear features of multi-dimensional data; and generating a weighted normalized feature vector by fusing the results of statistical analysis and deep learning through a two-layer feature selection mechanism.
[0027] S3. Semantic alignment: Map the time series features to the semantic space of the large language model corpus to complete natural language alignment.
[0028] In some embodiments, in step S3, the semantic alignment specifically includes: performing dimensionality reduction and compression on the large language model corpus to reduce the complexity of matching calculations; using a masked multi-head self-attention mechanism to match the time series data with the corpus Tokens to ensure that only historical information is concerned; and combining the Token sequences of the power generation data and external feature data to generate an input format adapted to the large language model.
[0029] S4. Prompt Generation: Generate a prompt prefix containing background, domain, instructions, and statistical information based on a preset prompt template.
[0030] In some embodiments, in step S4, the prompt template specifically includes: data background information describing the geographical location, time range, and external characteristics of the photovoltaic power station; defining typical temporal patterns and domain knowledge of photovoltaic power generation; determining the historical step and prediction step instructions for the prediction task; providing statistical features of historical power generation data to enhance semantic expression.
[0031] S5. Natural Language Prediction: Input the semantically aligned data and the prompt into a large language model with frozen parameters to output a natural language prediction result.
[0032] S6. Output Projection: Inversely parse the natural language response into a numerical power generation prediction value through linear projection.
[0033] In some embodiments, in step S6, the output projection specifically includes: flattening the natural language sequence output by the large language model into a one-dimensional vector; mapping the flattened vector to the target prediction dimension through linear transformation to restore the numerical result.
[0034] In some embodiments, the pre-trained parameters of the large language model are kept fixed during the prediction process, and only the parameters of the input segmentation, feature extraction, semantic alignment, prompt generation, and output projection modules are optimized.
[0035] In some embodiments, the evaluation metrics for the prediction results include: symmetric mean absolute percentage error, which measures the relative deviation between the predicted value and the true value; and root mean square error, which quantifies the global accuracy of the prediction results.
[0036] In some embodiments, the method is adapted to any one of the following data scenarios: training only using the very short-term historical data of the target photovoltaic power station; cross-scenario transfer prediction based on data from similar power stations; cross-scenario transfer prediction based on data from dissimilar power stations.
[0037] The following further describes specific embodiments of the present invention, its algorithm examples, and experimental verification.
[0038] A few-shot transferable prediction method for photovoltaic power generation based on large language models, which is mainly implemented through five functional modules: an input segmentation layer, a feature extraction layer, a semantic alignment layer, a prompt word generation layer, and an output projection layer; among them, the input segmentation layer divides the original photovoltaic power generation time series data into multiple embedding modules (Patch) to adapt to the processing requirements of large language models for structured language input and enhance the ability to express local data features; the feature extraction layer combines deep learning and statistical knowledge to extract key dynamic features of power generation data and external factors such as the environment, improving the model's ability to depict complex driving factors; the semantic alignment layer maps the structured time series features to the semantic space of the large language model's pre-trained corpus, realizing the high-fidelity conversion of data into natural language; the prompt word generation layer automatically generates context-related prompt words based on the original time series information to guide the large language model (LLM) to deeply understand and mine features of the target time series segment; the output projection layer then reversely parses the natural language response generated by the LLM, extracts the power generation prediction information contained therein, and converts it into a numerical form. This method breaks through the expression bottleneck of traditional time series modeling methods and realizes natural language-driven transfer prediction of photovoltaic power generation data. Specifically, the embodiments of the present invention include the following processing steps: (1) Processing by the input segmentation layer.
[0039] In order to enable it to understand time series data without fine-tuning or modifying the large model, a large amount of time series data is first segmented. This facilitates subsequent semantic alignment of each segmented Patch. The core of the input segmentation layer includes two steps: normalization and Patch splitting.
[0040] For a given time series data sample , first normalize it:
[0041] In the formula, and are the maximum and minimum values in the time series data respectively, is the normalized time series data.
[0042] Next, divide the normalized training samples into several small pieces, thereby grouping the original time series data into individual Patches to retain the original semantic information to the greatest extent. Set the length of each Patch to , then the number of Patches can be calculated by the following formula:
[0043] In the formula, is the source data 's length, is the step size of the horizontal sliding of the split window, is the floor function.
[0044] Through the above steps, a sample with the shape of is obtained . Finally, in order to match the input dimension with the column input size of the large language model , the present invention introduces a linear layer to map the split sample:
[0045] In the formula, has a dimension of , is the bias term, is the output of the final input split layer, and its dimension is .
[0046] (2) Feature extraction layer processing.
[0047] In the power generation prediction of photovoltaic, the power generation is affected by various factors such as solar radiation, cloud cover, and solar radiation angle, resulting in the source input sample usually being multi-dimensional data. When performing power generation prediction, the influence of each input variable on the prediction result is often unclear, and there may be significant differences in the contribution degrees of different features.
[0048] To solve this problem, the present invention proposes a double-layer feature enhancement model, which assigns weights to each variable through feature engineering. This model combines statistics and machine learning, introduces high-order partial correlation analysis as prior knowledge, and calculates the high-order partial correlation coefficients between variables before each round of machine learning training, so that the model can dynamically adjust the feature weights on the premise of excluding feature correlations. In addition, the present invention introduces a gated residual network (GRN) to measure the correlation between the external input and the target variable. GRN combines a gating mechanism and a residual connection, can flexibly capture non-linear correlations, and enhance the expression ability of the model.
[0049] (2-1) High-order partial correlation analysis.
[0050] First of all, the present invention introduces high-order partial correlation analysis. If the number of variables is , for any two variables with orders of and , the calculation formula for their high-order partial correlation coefficient is: , In the formula, represents the high-order partial correlation coefficient between variables and , and the right side of the equal sign is all The high-order partial correlation coefficient of the order. The judgment criteria for the correlation strength are as follows: At it is uncorrelated, at it is weakly correlated, at it is moderately correlated, and at it is strongly correlated.
[0051] (2-2) Gated Residual Network.
[0052] Considering that there are a large number of non-linear associations between photovoltaic power generation and its influencing factors, the present invention introduces a gated residual network model. The preprocessed data is first combined with the context vector and passes through to generate the intermediate layer : ,
[0053] In the formula, is the weight matrix, is the bias term. ELU acts as a feature recognition function at and acts as an activation function to produce a constant output at , thus exhibiting linear layer behavior.
[0054] Then, a weight transformation is applied to and the bias term is added to obtain the intermediate layer . It is input into the gated linear unit (GLU), combined with the initial feature and normalized through the regularization layer to obtain the output of the GRN: ,
[0055] In the formula, is an index representing weight sharing. The calculation formula of the gating mechanism of GLU is as follows: , In the formula, is the sigmoid activation function, is the Hadamard product.
[0056] GLU enables the prediction model to dynamically control the contribution degree of GRN to the initial feature and has the ability of adaptive adjustment.
[0057] When GRN finds that no additional transformation is required in dealing with non - linear contributions, GLU can approximately output zero, enabling the model to directly skip this layer. Additionally, for instances without context vectors, GRN will treat as zero.
[0058] (2 - 3) First - layer feature selection.
[0059] The high - order partial correlation coefficients obtained through high - order partial correlation analysis and GRN and the features after non - linear transformation by GRN are input into Softmax for normalization: , , where is the weight matrix after the first - layer feature selection mechanism.
[0060] (2 - 4) Second - layer feature selection.
[0061] At each time step, each original feature is sent into the corresponding GRN for non - linear feature extraction: , Then, according to the weight matrix calculated in the first layer, the feature matrix is weighted and summed to obtain the final normalized feature vector :
[0062] where has a dimension of , has a dimension of , and finally has a dimension of . This feature vector converts the external features at each time step into a unified influencing feature through weighted summation.
[0063] (2 - 5) Feature mapping and output.
[0064] Finally, in order to match the feature vector with the input dimension of the large - language model, segmentation and linear transformation are performed. Through formula (x), a sample with a shape of is obtained and mapped to the final output through linear transformation:
[0065] where has a dimension of , is the bias term, and the final output has a dimension of .
[0066] (3) Semantic alignment layer processing.
[0067] (3-1) Corpus dimensionality reduction.
[0068] The corpus of the LLM usually contains hundreds of millions of tokens. If directly matching time series data, it will cause extremely large memory and computational overhead. Therefore, the corpus is dimensionally reduced, and then each data Patch is matched.
[0069] Let the original corpus of the LLM be , where there are tokens. A random boolean matrix (with a dimension of ) is used to reduce its dimension, obtaining a new corpus :
[0070] In the formula, only contains tokens, and , effectively reducing the computational complexity (3-2) Token matching mechanism.
[0071] The present invention proposes to use an improved multi-head self-attention mechanism to match each Patch with the tokens in the corpus . The common attention mechanism scales the value( ) through the relationship between the key( ) and the query( ). In Token matching, is the time series data Patch to be matched, is the corpus after dimensionality reduction , is the finally matched Token sequence . The matching process is completed through the self-attention mechanism, and its core calculation method is as follows: , In the formula, is the normalization function. The traditional attention mechanism may see future data when calculating the matching of the current Patch, resulting in "distraction". Therefore, the present invention introduces a mask matrix to ensure that only historical information is focused on at the current moment: , Wherein, is the dimension of the key for scaling calculation to prevent the gradient from vanishing due to an overly large inner product value. When it means not paying attention to ; when it means paying attention to .
[0072] To reduce the time and memory consumption during the matching process, the self-attention mechanism is extended to the multi-head attention mechanism. Specifically, the value weight matrix is shared among all attention heads, and additive aggregation instead of concatenation operation is adopted to reduce the computational complexity: , , Wherein, is the attention output after aggregation, is the shared value weight matrix. The calculation method of the improved attention weight is as follows: , Wherein, is the number of attention heads, and are the weight matrices of the query and key corresponding to the th attention head.
[0073] Finally, based on the improved self-attention mechanism, the data Patch can be matched to the corpus Token to form a Token sequence: . (3-3) Token combination.
[0074] Through the above matching process, the Token sequences of the power generation data and the Token sequence of the external feature data are obtained respectively. Finally, the sequences are combined to obtain the input that can be converted from the power generation data into Tokens and embedded into the LLM model:
[0075] This process completes the conversion of time series data into natural language, enabling the LLM to directly understand and use it for prediction.
[0076] (4) Prompt generation layer processing.
[0077] A prompt is a simple and efficient method for activating an LLM. However, in the photovoltaic power generation prediction task, converting time series data into a Token "black box" causes the LLM to limit its semantic understanding ability without being informed, hindering its process of analyzing Tokens and predicting outputs. Therefore, before inputting the converted natural language data, a set of prompt words is designed as an input prefix to guide the LLM to correctly understand the data and embed the corresponding Tokens into a suitable structure.
[0078] Therefore, a prompt word template in the following format is designed: [Description] This dataset is a photovoltaic power generation dataset located in the xx region, with a time range from xx to xx and a sampling interval of xx. The dataset includes xx external features such as power generation values and xx, xx.
[0079] [Domain] The power generation data in this scenario usually starts at 6 am, stops at 7 pm, and reaches its peak at noon.
[0080] [Instruction] Based on historical <x>Relevant data of the steps to predict the future <x>Power generation value of the step.
[0081] [Statistics] The minimum value of the historical power generation data is xx, the maximum value is xx, and the median is xx. The overall trend is xx, with strong time series characteristics.
[0082] Embed the above prompt as a prefix before the Patch input. [Description] and [Domain] provide background information on the input time series for the LLM, helping it understand the data source and characteristics. [Instruction] clarifies the role of the LLM in the prediction task, guiding it to reasonably transform the Patch. [Statistics] enhances the semantic expression of the time series data, providing key pattern recognition and reasoning clues for the LLM.
[0083] (5) Output projection layer processing.
[0084] After obtaining the Token combination from the remapping of the prompt and the photovoltaic power generation data, input it into the frozen LLM for forward propagation to obtain the prediction output represented in natural language. . Subsequently, flatten it and obtain the prediction result finally restored to time series data through linear projection:
[0085] In the formula, is the sample corresponding photovoltaic power generation prediction value, is the flattening operation, which converts to a format suitable for the fully connected layer, is the weight matrix of the linear projection, is the bias term.
[0086] This layer regresses the natural language output by the LLM to time series data, thus completing the photovoltaic power generation prediction task.
[0087] (6) Establish an evaluation index system for the prediction results.
[0088] This method uses the symmetric mean absolute percentage error (SMAPE) and the normalized root mean square error (NRMSE) to evaluate the accuracy of the prediction results. The definitions and calculations of the evaluation indexes are as follows: , , Among them, is the predicted value, is the true value, is the mean of the true values.
[0089] To evaluate the performance of the proposed method, the present invention conducted experimental tests under three different data scarcity scenarios: (F1) training using only the historical data of the target photovoltaic power station in the recent 7 days; (F2) training using the data of a photovoltaic power station similar to the target power station in the recent 30 days; (F3) training using the data of a photovoltaic power station not similar to the target power station in the recent 30 days. All models were finally verified on the 5-day test data of the unified target photovoltaic power station.
[0090] In the method of the present invention, the pre-trained large language model Qwen-7B is introduced as a plug-in and embedded into the prediction framework. Qwen-7B receives inputs from the semantic alignment layer and the prompt word generation layer, and maps the natural language results generated by it to the predicted photovoltaic power generation value through the output projection layer. Qwen-7B is built based on an improved Transformer architecture, contains 7 billion parameters, the vocabulary size is 153K, and the maximum input length is 4096. It should be particularly noted that during all experiments, the parameters of Qwen-7B remain frozen, and only the parameters of the remaining modules in the proposed method are trained.
[0091] To comprehensively verify the effectiveness of the proposed method, the inventors selected four mainstream prediction models for comparative experiments, namely ConvTrans, TFT, Seq2Seq, and STGCN. The experimental results are shown in Table 1. In the three data scarcity scenarios of F1, F2, and F3, the proposed large language model-based photovoltaic power generation transferable prediction method achieved the best performance, fully demonstrating its accuracy and generalization ability under small sample conditions.
[0092] Table 1 Evaluation indicators of the method of the present invention and comparative models under different scenarios
[0093] In summary, the present invention proposes a few-shot transferable prediction method for photovoltaic power generation based on a large language model. The pre-trained large language model is embedded as a plug-in into photovoltaic power generation prediction, making full use of its powerful generalization ability and language modeling ability to maximize the extraction of potential features of photovoltaic power generation data under data scarcity conditions, thereby significantly improving the transfer ability of the prediction model. The proposed semantic alignment layer can remap structured time series data into natural language expression forms to achieve alignment with the corpus space of the large language model, so that the prediction task can be efficiently completed without adjusting the parameters of the large language model itself. In addition, the proposed feature extraction layer fuses the deep learning mechanism on the basis of fully isolating the correlation between different features, realizes end-to-end automatic feature extraction, and constructs an efficient and scalable feature modeling framework. The present invention constructs a new type of photovoltaic prediction framework with low data adaptability, strong semantic understanding ability, and high transferability, thus promoting the photovoltaic prediction technology to a higher level of intelligence and generalization.
[0094] An embodiment of the present invention further provides a storage medium for storing a computer program, which, when executed, at least executes the method described above.
[0095] An embodiment of the present invention further provides a control device, including a processor and a storage medium for storing a computer program; wherein, the processor is configured to at least execute the method described above when executing the computer program.
[0096] An embodiment of the present invention further provides a processor, which executes a computer program and at least executes the method described above.
[0097] The storage medium can be implemented by any type of non-volatile storage device, or a combination thereof. Among them, the non-volatile memory can be a read-only memory (ROM, Read Only Memory), a programmable read-only memory (PROM, Programmable Read-Only Memory), an erasable programmable read-only memory (EPROM, Erasable Programmable Read-Only Memory), an electrically erasable programmable read-only memory (EEPROM, Electrically Erasable Programmable Read-Only Memory), a ferromagnetic random access memory (FRAM, Ferromagnetic Random Access Memory), a flash memory (Flash Memory), a magnetic surface memory, an optical disc, or a compact disc read-only memory (CD-ROM, Compact Disc Read-Only Memory); the magnetic surface memory can be a disk memory or a tape memory. The storage medium described in the embodiments of the present invention is intended to include, but not limited to, these and any other suitable types of memories.
[0098] In several embodiments provided by the present invention, it should be understood that the disclosed systems and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined, or can be integrated into another system, or some features can be ignored, or not executed. In addition, the coupling, direct coupling, or communication connection between the components shown or discussed with each other can be through some interfaces, and the indirect coupling or communication connection of the devices or units can be electrical, mechanical, or other forms.
[0099] The units described above as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0100] In addition, in each embodiment of the present invention, all the functional units may be integrated into one processing unit, or each unit may be a separate unit alone, or two or more units may be integrated into one unit. The above integrated units may be implemented in the form of hardware, or in the form of a combination of hardware and software functional units.
[0101] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps including the above method embodiments. The foregoing storage medium includes: various media such as removable storage devices, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.
[0102] Alternatively, if the above integrated units of the present invention are implemented in the form of software functional modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the embodiments of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the methods described in the various embodiments of the present invention. The foregoing storage medium includes: various media such as removable storage devices, ROM, RAM, magnetic disks, or optical discs that can store program codes.
[0103] The methods disclosed in several method embodiments provided by the present invention can be combined arbitrarily without conflict to obtain new method embodiments.
[0104] The features disclosed in several product embodiments provided by the present invention can be combined arbitrarily without conflict to obtain new product embodiments.
[0105] The features disclosed in several method or device embodiments provided by the present invention can be combined arbitrarily without conflict to obtain new method embodiments or device embodiments.
[0106] The above content is a further detailed description of the present invention in combination with specific preferred embodiments. It cannot be determined that the specific implementation of the present invention is only limited to these descriptions. For those skilled in the art to which the present invention pertains, without departing from the concept of the present invention, several equivalent substitutions or obvious modifications can be made, and as long as the performance or use is the same, they should all be regarded as falling within the protection scope of the present invention.< / x> < / x>
Claims
1. A small sample transferable prediction method for photovoltaic power generation based on a large language model, characterized in that: The following steps are involved: S1. Input segmentation: normalize and segment the original photovoltaic time series data to generate an embedding module that is suitable for the input of the large language model; S2. Feature extraction: Dynamically assign feature weights through high-order partial correlation analysis and gated residual networks to extract nonlinear correlation features of multidimensional data; S3, semantic alignment: Map the temporal features to the semantic space of the large language model corpus to complete natural language alignment; S4, prompt generation: generating a prompt word prefix containing background, field, instruction and statistical information based on a preset prompt word template; S5, natural language prediction: input the semantically aligned data and the prompt words into the large language model with frozen parameters, and output the natural language prediction results; S6. Output projection: The natural language response is deconstructed into a numerical power generation prediction value through linear projection.
2. The method according to claim 1, characterized in that In step S1, the input segmentation specifically includes: Perform extreme value normalization on the original photovoltaic time series data; The normalized data is divided into blocks according to the preset step size to generate multiple embedding modules; The segmented data is mapped to a feature space that matches the input dimension of the large language model through a linear transformation.
3. The method according to claim 1, characterized in that In step S2, the feature extraction specifically includes: Dynamically evaluate the strength of association between variables based on high-order partial correlation analysis to eliminate multicollinearity interference; The gated residual network is used to combine the gating mechanism and residual connection to extract the nonlinear features of multidimensional data; The statistical analysis and deep learning results are integrated through a two-layer feature selection mechanism to generate a weighted normalized feature vector.
4. The method according to claim 1, characterized in that: In step S3, the semantic alignment specifically includes: Perform dimensionality reduction compression on large language model corpora to reduce the complexity of matching calculations; A masked multi-head self-attention mechanism is used to match time series data with corpus tokens to ensure that only historical information is focused on; The token sequence of power generation data and external feature data is combined to generate an input format suitable for the large language model.
5. The method according to claim 1, characterized in that In step S4, the prompt word template specifically includes: Data background information describing the geographical location, time range and external characteristics of the PV plant; Define typical timing patterns and domain knowledge of photovoltaic power generation; Determine the historical step size and prediction step size instructions for the prediction task; Provide statistical features of historical power generation data to enhance semantic expression.
6. The method according to claim 1, characterized in that In step S6, the output projection specifically includes: Flatten the natural language sequence output by the large language model into a one-dimensional vector; The flattened vector is mapped to the target prediction dimension through linear transformation to restore the numerical result.
7. The method according to claim 1, characterized in that The large language model keeps the pre-trained parameters fixed during the prediction process and only optimizes the parameters of the input segmentation, feature extraction, semantic alignment, prompt generation and output projection modules.
8. The method according to claim 1, characterized in that The evaluation indicators of the prediction results include: Symmetric mean absolute percentage error, which measures the relative deviation between the predicted value and the true value; Normalized root mean square error quantifies the global accuracy of the prediction results.
9. The method according to claim 1, characterized in that: The method is applicable to any of the following data scenarios: Only very short-term historical data of the target PV plant is used for training; Cross-scenario migration prediction based on similar power station data; Cross-scenario migration prediction based on non-similar power station data.
Citation Information
Patent Citations
Electric vehicle charging load prediction method and device
CN117526307A
Airport short-term load prediction method based on large language model migration
CN118966422A
System and Method for Resource Efficient Natural Language Processing
US20220292266A1
Cited By
Language model-based time sequence prediction method and prediction device
CN120688643A
Photovoltaic power generation power cross-substation prediction method based on self-attention mechanism
CN121072883A
Photovoltaic power generation power prediction method and system based on large language model
CN121172753A
Regional fertility census statistical prediction system based on big data
CN121301808A