High-precision medium-and-long-term photovoltaic power generation prediction method based on multi-modal data fusion

By collecting and fusing multimodal data from photovoltaic devices and utilizing a photovoltaic power generation prediction model with multi-head attention and gating mechanisms, the problems of insufficient accuracy and robustness in photovoltaic power generation prediction in existing technologies are solved, and high-precision medium- and long-term predictions are achieved.

CN120671072APending Publication Date: 2025-09-19JIANGSU LINYANG ZHIWEI TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510737374.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-04
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

Existing photovoltaic power generation prediction methods have deficiencies in accuracy and robustness, especially in medium- and long-term predictions and under small sample conditions, where it is difficult to achieve high precision and high requirements are placed on data volume and quality.

Method used

Multimodal data of photovoltaic equipment, including time series, image and text data, is collected. Through the multimodal data fusion framework and the improved photovoltaic power generation prediction model, multimodal features are extracted using the multi-head attention mechanism and gating mechanism to perform medium- and long-term power generation forecasts.

Benefits of technology

The accuracy and efficiency of photovoltaic power generation prediction are improved, high-precision medium- and long-term predictions can be achieved under small sample conditions, and the stability and robustness of the model are enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120671072A_ABST
    Figure CN120671072A_ABST
Patent Text Reader

Abstract

The invention discloses a high-precision medium-and-long-term photovoltaic power generation prediction method based on multi-modal data fusion, and relates to the related field of photovoltaic power generation prediction.The method comprises the steps that power generation data of photovoltaic equipment within a period of time is collected to serve as photovoltaic power generation historical data; preprocessing the collected photovoltaic power generation historical data, and converting the data into stable data; extracting multi-modal data from the preprocessed photovoltaic power generation historical data by applying a learner in the photovoltaic power generation prediction model, and obtaining multi-modal advanced feature representation; a gating mechanism is introduced into the photovoltaic power generation prediction model to fuse multi-modal features, and photovoltaic power generation data in a future medium-and-long-term time period is predicted through multi-modal interaction. The problems that an existing photovoltaic power generation prediction method is poor in stability and robustness and too high in requirement for data quality are solved, the complementary advantage of multi-modal data is fully utilized, the dependence of model performance on the data quality is reduced, and high-precision medium-and-long-term photovoltaic power generation prediction is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of photovoltaic power generation prediction, and in particular to a high-precision medium- and long-term photovoltaic power generation prediction method based on multimodal data fusion. Background Art

[0002] As a renewable resource, photovoltaic power generation continues to grow its share of the global energy mix. However, the intermittent and uncertain nature of photovoltaic power generation poses significant challenges to grid scheduling and energy management. Accurate photovoltaic power generation forecasts can help grid operators optimize scheduling, reduce reserve capacity requirements, and improve grid operation efficiency and economics.

[0003] Traditional physical models are based on the physical properties of photovoltaic cell modules and calculate theoretical output power through mathematical formulas. However, this method is sensitive to meteorological conditions and has large prediction errors. Statistical methods are used to make predictions using historical data trends and seasonal information on photovoltaic power generation, mining data patterns, and predicting short-term power generation data. This can effectively reduce prediction errors, but this method has limited ability to capture nonlinear relationships.

[0004] Existing methods widely apply machine learning and deep learning technologies. From the perspective of big data analysis, predictions are achieved by constructing a mapping relationship between environmental data and photovoltaic power generation data. However, these methods rely on a large amount of labeled data and have high requirements on the quantity and quality of data. In addition, the method of using environmental data to predict power generation data relies on the learning ability of the neural network model, and its accuracy and robustness are limited. It can only be applied to short-term power generation predictions. There is an urgent need to improve the feature types, model structure and training methods to further improve the accuracy and efficiency of photovoltaic power generation predictions. Summary of the Invention

[0005] In response to the technical problems of the above-mentioned prior art, the present invention applies to provide a high-precision medium- and long-term photovoltaic power generation prediction method based on multimodal data fusion, which collects power generation data of photovoltaic equipment and extracts multimodal data from it, including time series, image and text data; further applies the multimodal data fusion framework, introduces an improved photovoltaic power generation prediction model, realizes medium- and long-term power generation prediction of photovoltaic equipment, and improves the accuracy and efficiency of photovoltaic power generation prediction.

[0006] This application provides a high-precision medium- and long-term photovoltaic power generation prediction method based on multimodal data fusion, including:

[0007] (1) Collecting the power generation data of photovoltaic equipment over a period of time as photovoltaic power generation history data;

[0008] (2) Preprocess the collected historical photovoltaic power generation data and convert the data into stable data;

[0009] (3) Apply the learner in the photovoltaic power generation prediction model to extract multimodal data from the preprocessed photovoltaic power generation historical data and obtain high-level feature representation of multimodality;

[0010] (4) The photovoltaic power generation prediction model introduces a gating mechanism to integrate multimodal features and uses multimodal interaction to predict photovoltaic power generation data in the future medium and long term.

[0011] Furthermore, during the data collection process, a smart meter is installed at the output end of the photovoltaic equipment. The collected photovoltaic power generation data includes the generated power and generated amount, and the real-time generated power and cumulative generated amount of the photovoltaic equipment at each time point are recorded;

[0012] The collected data is transmitted to a computer terminal in real time for subsequent processing. The data is stored in a relational database in the form of a table, where each row represents a data record and each column represents a data field;

[0013] The power generation data over a period of time is organized into historical data of photovoltaic power generation, which serves as training samples for the photovoltaic power generation prediction model.

[0014] Furthermore, the main steps of preprocessing the collected photovoltaic power generation historical data include data cleaning, data smoothing and differential processing:

[0015] The data cleaning step includes removing outliers and filling missing values. Outliers refer to data points that significantly deviate from the normal range, such as negative power generation or power that exceeds the rated power of the equipment by several times. Missing values ​​refer to values ​​that are not recorded at certain time points in the data.

[0016] Outliers were detected using boxplots, which divided the data into quartiles. Data points outside the interquartile range (IQR) by 1.5 were identified as outliers and replaced with the median. Missing values ​​in the data were handled using linear interpolation, which imputed missing values ​​by calculating the mean of two adjacent data points.

[0017] The purpose of data smoothing is to reduce random fluctuations in the data and facilitate subsequent analysis. The moving average method is used to smooth the data and calculate the average value of each time point and several time points before and after it. The calculation formula is as follows: t , its n-period moving average MA t Expressed as:

[0018]

[0019] Among them, n represents the selected time point range; X t-n+1 and X t-n+2 Represents the data at the t-n+1th and t-n+2th time points.

[0020] The photovoltaic power generation history data is converted into stable data through differential processing. The difference between the data of adjacent time points is calculated by calculating the first-order difference method. t , its first-order difference ΔX t Expressed as:

[0021] ΔX t =X t -X t-1

[0022] Among them, X t-1 Represents data from the previous point in time; stabilized data is suitable for analyzing long-term data of photovoltaic power generation and improving prediction accuracy.

[0023] Furthermore, for the preprocessed photovoltaic power generation historical data, the photovoltaic power generation prediction model extracts time series, image and text data through the learner to obtain high-level temporal features, visual features and text features:

[0024] The cued-enhanced learner of the photovoltaic power generation prediction model segments historical photovoltaic power generation data into multiple time series data blocks and extracts temporal features from them, preserving the original memory information of the data. The data block embedding interacts through multi-head self-attention and pooling mechanisms to further retain rich temporal feature information, enhance the modeling ability of long-range dependencies, and improve prediction robustness.

[0025] The image enhancement learner in the photovoltaic power generation prediction model converts time series into information-rich three-channel images through multi-scale convolution, frequency and periodicity encoding. The images are processed by a pre-trained VLM (Visual Language Model) visual encoder to extract hierarchical visual features that capture fine-grained details and high-level temporal patterns.

[0026] The text-enhanced learner in the photovoltaic power generation prediction model generates contextual text prompts for the input time series, including statistical features (e.g., mean, variance, trend), domain-specific contextual information, and image descriptions as text data of photovoltaic power generation history data; these prompt data are encoded through the pre-trained VLM text encoder to generate text features.

[0027] The visual and textual features extracted by the VLM are combined with temporal features through a gated fusion mechanism to capture complementary information; these rich multimodal features are then processed by a fine-tuned predictor to output predicted values ​​of photovoltaic power generation data in the future medium and long term.

[0028] Furthermore, the photovoltaic power generation prediction model uses a multimodal fusion network to aggregate temporal, visual, and textual features, leveraging their complementary advantages to enhance prediction accuracy:

[0029] To address the distribution shift between temporal and multimodal features, all modalities are projected into a shared space of the same dimension. Temporal memory embedding F from retrieval-enhanced learner tem The high-level temporal patterns of photovoltaic power generation data are encoded and used as the query vector Q for the cross-modal multi-head attention mechanism; the multimodal embedding F from the two visual language model encoders mm Being used as the key vector K and the value vector V, the cross-modal multi-head attention CM-MHA is defined as:

[0030] CM-MHA(Q,K,V)=Cat(head1,...,head h )W O

[0031]

[0032] Among them, Cat represents splicing according to channel dimension, Q = F tem W Q , K=F mm W K , V=F mm W V ;head i is the i-th self-attention, W O , W Q , W K , W V , W i Q , W i K and W i V is the learnable mapping matrix, d k represents the normalization parameter, which is determined by the feature dimension and the number of heads of the attention mechanism; softmax represents the softmax activation function, which ensures that the weight value is within a specific range.

[0033] The attention mechanism distributes and aggregates temporal and multimodal features, capturing fine-grained patterns and high-level contextual feature representations. Secondly, residuals and layer normalization (LayerNorm) are introduced to stabilize the training process:

[0034] F attn =LayerNorm(F tem +CM_MHA(Q,K,V)

[0035] The gating mechanism further obtains the output feature F by dynamically weighting each modality fused :

[0036] G=σ(W g [F tem ; F mm ]+bg )

[0037] F fused =G⊙F attn +(1-G)⊙F mm

[0038] Among them, G is the gating weight matrix, W g and b g is a learnable parameter, σ(·) is the sigmoid activation function, [;] represents feature concatenation, and ⊙ represents element-wise multiplication. This gating mechanism adaptively balances temporal and multimodal features to achieve robust prediction.

[0039] Furthermore, the photovoltaic power generation prediction model is trained end-to-end using the mean square error loss, assuming that the photovoltaic power generation historical data is Where N represents the dimension of photovoltaic power generation data, T represents the time step; the training goal is to predict the power generation data for the next H time steps. in The optimization objective of the model is expressed as:

[0040]

[0041] in, and Y h are the predicted and true values ​​at the hth time point.

[0042] After training is complete, the model parameters that achieved optimal performance are retained as the standard model for medium- and long-term photovoltaic power generation forecasting, meeting the needs of photovoltaic power generation forecasting. The model's hyperparameters are adjusted and updated based on forecast feedback. Furthermore, the photovoltaic power generation forecasting model proposed in this invention can be used for few- and zero-sample training, achieving high-precision power generation forecasts even with insufficient data.

[0043] The present invention discloses the following technical effects:

[0044] This paper proposes a high-precision medium- and long-term photovoltaic power generation forecasting method based on multimodal data fusion. This method extracts multimodal information from collected historical photovoltaic power generation data and uses the complementarity between modalities to predict power generation data. While photovoltaic power generation time series data provides information about photovoltaic power generation at each moment, the information available is limited, limiting the accuracy of photovoltaic power generation forecasts. Therefore, this method analyzes the photovoltaic power generation time series data and extracts image and text data to supplement the power generation information, leveraging the complementary nature of multimodal data to improve forecast accuracy. Based on the principles of big data analysis, this paper constructs a photovoltaic power generation forecasting model. It uses a cue-enhanced learner, an image-enhanced learner, and a text-enhanced learner to process and analyze photovoltaic power generation time series, image, and text data, respectively. This model comprehensively captures temporal and multimodal features. Furthermore, it extracts rich information for power generation forecasting from the fused features, enabling long-range temporal context modeling and enabling the model to predict medium- and long-term photovoltaic power generation data. This model utilizes a pre-trained encoder, an attention mechanism, and a gating mechanism to enhance feature extraction capabilities, improve the stability and robustness of model training, and achieve high-precision medium- and long-term photovoltaic power generation forecasts even with limited training samples. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings of the embodiments of the present invention are briefly introduced below. Flowcharts are used in this application to illustrate the operations performed by the system according to the embodiments of the present application. It should be understood that the preceding or following operations are not necessarily performed in precise order. Instead, various steps may be processed in reverse order or simultaneously as needed. Furthermore, other operations may be added to these processes, or one or more operations may be removed from these processes.

[0046] Figure 1 A flowchart of a high-precision medium- and long-term photovoltaic power generation prediction method based on multimodal data fusion provided in an embodiment of the present application.

[0047] Figure 2 A schematic diagram of the structure of a photovoltaic power generation prediction model based on multimodal data fusion provided in an embodiment of the present application. DETAILED DESCRIPTION

[0048] The above description is only an overview of the technical solution of the present application. In order to more clearly understand the technical means of the present application, it can be implemented in accordance with the contents of the specification. In order to make the above and other purposes, features and advantages of the present application more obvious and easy to understand, the specific implementation methods of the present application are listed below.

[0049] In order to make the purpose, technical solutions and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limiting this application. All other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.

[0050] In the following description, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or server that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or modules that are not explicitly listed or are inherent to these processes, methods, products, or devices. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application belongs. The terms used herein are for the purpose of describing the embodiments of this application only.

[0051] Example 1: This application embodiment provides a high-precision medium- and long-term photovoltaic power generation prediction method based on multimodal data fusion, such as Figure 1 As shown, the method includes:

[0052] Step S10: collecting power generation data of the photovoltaic equipment over a period of time as photovoltaic power generation history data.

[0053] During the data collection process, a smart meter is installed at the output end of the photovoltaic equipment. The collected photovoltaic power generation data includes power generation and power generation, and records the real-time power generation and cumulative power generation of the photovoltaic equipment at each time point;

[0054] The collected data is transmitted to a computer terminal in real time for subsequent processing. The data is stored in a relational database in the form of a table, where each row represents a data record and each column represents a data field;

[0055] The power generation data over a period of time is organized into historical data of photovoltaic power generation, which serves as training samples for the photovoltaic power generation prediction model.

[0056] Step S20 , pre-processing the collected photovoltaic power generation historical data to convert the data into stable data.

[0057] The main steps for processing the collected historical photovoltaic power generation data include data cleaning, data smoothing and differential processing:

[0058] The data cleaning step includes removing outliers and filling missing values:

[0059] In this example, outliers were detected using a boxplot, which divides the data into quartiles. Data points outside the interquartile range (IQR) by 1.5 times were identified as outliers and replaced with the median. Missing values ​​in the data were handled using linear interpolation, which filled the gaps by calculating the mean of two adjacent data points.

[0060] In this embodiment, the moving average method is used to smooth the data and calculate the average value of each time point and several time points before and after it. The calculation formula is as follows: t , its n-period moving average MA t Expressed as:

[0061]

[0062] Among them, n represents the selected time point range; X t-n+1 and X t-n+2 Represents the data at the t-n+1th and t-n+2th time points.

[0063] This embodiment converts photovoltaic power generation history data into stable data through differential processing. The method of calculating first-order difference is used to calculate the difference between the data of adjacent time points. t , its first-order difference ΔX t Expressed as:

[0064] ΔX t =x t -X t-1

[0065] Among them, X t-1 Represents data from the previous point in time; stabilized data is suitable for analyzing long-term data of photovoltaic power generation and improving prediction accuracy.

[0066] Step S30 , applying the learner in the photovoltaic power generation prediction model to extract multimodal data from the preprocessed photovoltaic power generation historical data, and obtaining a high-level feature representation of the multimodal data.

[0067] In this embodiment, for the pre-processed photovoltaic power generation historical data, the photovoltaic power generation prediction model obtains multimodal data in a specific manner and further extracts multimodal features:

[0068] The cue-enhanced learner in the photovoltaic power generation prediction model divides the photovoltaic power generation historical data into multiple time series data blocks and extracts temporal features from them, preserving the original memory information of the data. The data block embedding interacts through multi-head self-attention and pooling mechanisms to further retain rich temporal feature information, enhance the modeling ability of long-range dependencies, and improve prediction robustness.

[0069] The image reinforcement learner in the photovoltaic power generation prediction model converts time series into information-rich three-channel images through multi-scale convolution, frequency, and periodicity encoding. The images are processed by a pre-trained VLM visual encoder to extract hierarchical visual features that capture fine-grained details and high-level temporal patterns. This example uses ViLT as the pre-trained VLM visual encoder.

[0070] The text-enhanced learner in the photovoltaic power generation prediction model generates contextual text prompts for the input time series, including statistical features (e.g., mean, variance, trend), domain-specific contextual information and image descriptions, as text data of the photovoltaic power generation history data; these prompt data are encoded by a pre-trained VLM text encoder to generate text features. The pre-trained VLM text encoder in this embodiment is CLIP.

[0071] Step S40 : The photovoltaic power generation prediction model introduces a gating mechanism to fuse multimodal features, and uses multimodal interaction to predict photovoltaic power generation data in the future medium and long term.

[0072] In this embodiment, the photovoltaic power generation prediction model uses a multimodal fusion network to aggregate temporal, visual, and textual features, leveraging their complementary advantages to enhance prediction accuracy:

[0073] To address the distribution shift between temporal and multimodal features, multiple modalities are projected into a shared space with the same dimension. Temporal memory embedding F from retrieval-enhanced learner tem The high-level temporal patterns of photovoltaic power generation data are encoded and used as the query vector Q of the cross-modal multi-head attention mechanism; the multimodal embedding f from the two visual language model encoders mm Being used as the key vector K and the value vector V, the cross-modal multi-head attention CM-MHA is defined as:

[0074] CM-MHA(Q,K,V)=Cat(head1,...,head h )W O

[0075]

[0076] Among them, Cat represents splicing according to channel dimension, Q = F tem W Q , K=F mm W K , V=F mm W V ;head i is the i-th self-attention, W O , W Q , W K , W V , Wi Q , W i K and W i V is the learnable mapping matrix, d k represents the normalization parameter, which is determined by the feature dimension and the number of heads of the attention mechanism; softmax represents the softmax activation function, which ensures that the weight value is within a specific range.

[0077] The attention mechanism distributes and aggregates temporal and multimodal features, capturing fine-grained patterns and high-level contextual feature representations. Secondly, residuals and layer normalization (LayerNorm) are introduced to stabilize the training process:

[0078] F attn =LayerNorm(F tem +CM_MHA(Q,K,V)

[0079] The gating mechanism further obtains the output feature F by dynamically weighting each modality fused :

[0080] G=σ(W g [F tem ; F mm ]+b g )

[0081] F fused =G⊙F attn +(1-G)⊙F mm

[0082] Among them, G is the gating weight matrix, W g and b g is a learnable parameter, σ(·) is the sigmoid activation function, [;] represents feature concatenation, and ⊙ represents element-wise multiplication. This gating mechanism adaptively balances temporal and multimodal features to achieve robust prediction.

[0083] Finally, the fused multimodal features are fed into the fine-tuned predictor for processing, and the predicted values ​​of photovoltaic power generation data in the future medium and long term are output.

[0084] In this embodiment, the photovoltaic power generation prediction model is trained end-to-end using the mean square error loss. Assuming that the photovoltaic power generation history data is Where N represents the dimension of photovoltaic power generation data, T represents the time step; the training goal is to predict the power generation data for the next H time steps. in The optimization objective of the model is expressed as:

[0085]

[0086] in, and Y h are the predicted and true values ​​at the hth time point.

[0087] After training is complete, the model parameters that achieved optimal performance are retained as the standard model for medium- and long-term photovoltaic power generation forecasting, meeting the needs of photovoltaic power generation forecasting. The model's hyperparameters are adjusted and updated based on forecast feedback. Furthermore, the photovoltaic power generation forecasting model proposed in this invention can be used for few- and zero-sample training, achieving high-precision power generation forecasts even with insufficient data.

[0088] Example 2: This application embodiment provides a high-precision medium- and long-term photovoltaic power generation prediction method based on multimodal data fusion. The detailed process of generating multimodal features by the photovoltaic power generation prediction model is as follows: Figure 2 As shown:

[0089] The retrieval reinforcement learner in the photovoltaic power generation prediction model, the detailed structure is as follows Figure 2 -(a). This module extracts high-level temporal features through data block segmentation and memory-enhanced attention. It dynamically retrieves and aggregates temporal features to adapt them to complex time series structures to enhance prediction results. It mainly includes two key stages:

[0090] The first key stage is to extract data block embeddings. For the input time series Where B is the batch size, L is the sequence length, and D represents the variable dimension. In this embodiment, D = 2, which includes the photovoltaic power and power generation at each time step.

[0091] The sequence is first split into multiple overlapping data blocks of length patch_len, and each data block is mapped to d model dimensional feature space, adding position embedding to it to preserve the time order. Finally, the data block embedding is obtained Contains local time domain pattern information, where N oatches is the number of data blocks.

[0092] The second key stage is to compute memory-enhanced attention, introducing a series of learnable query matrices Used to implement data block embedding interaction through multi-head attention. The data block embedding is mapped to the key matrix K o Sum matrix V o , used to calculate memory attention:

[0093]

[0094] To better capture complex temporal dynamics, a hierarchical memory structure is introduced. Local memory captures fine-grained patterns within each data block to capture short-range dependencies, while global memory aggregates information across data blocks to capture long-range dependencies and high-level temporal trends. The two types of memory are fused through a gating mechanism:

[0095] M fused =α·M local +(1-α)·M global

[0096] Among them, M fused represents fusion memory, M local and M global denote local memory and global memory respectively, and α is a learnable gating parameter that adaptively balances the contribution of local memory and global memory.

[0097] The size of the time domain feature finally output by the retrieval reinforcement learner is B×N queries ×d query This feature preserves high-level temporal patterns and enables dynamic retrieval and integration of other modalities. Each memory fragment captures semantic and temporal relationships, adaptively retrieving relevant temporal patterns based on the input context. This hierarchical memory mechanism effectively overcomes the challenges of capturing long-range dependencies and complex temporal dynamics.

[0098] Image reinforcement learning in photovoltaic power generation prediction model, the detailed structure is as follows Figure 2 -(b). This module adaptively converts time series into images, preserving fine-grained and high-level temporal patterns through the following steps:

[0099] In order to obtain the spectrum and time domain dependence, this module uses two complementary encoding techniques to add frequency and time domain information to time series data. Fast Fourier Transform (FFT) is applied for frequency encoding to extract frequency components:

[0100]

[0101] Where k is the frequency index and t is the time index. The obtained frequency features and the input time series are concatenated to obtain the shape Tensor of .

[0102] Second, sine and cosine functions are applied to encode temporal dependencies at each time step:

[0103]

[0104] Where encoding represents the encoding operation and P is a periodic hyperparameter. The generated encoding is concatenated with the input time series to obtain a shape of Tensor of .

[0105] The concatenated tensor is processed through multiple convolutional layers to obtain hierarchical temporal patterns. The one-dimensional convolutional layer obtains local dependencies and converts the input into And take the average value along the D dimension to get Then, a two-dimensional convolution is applied to halve the channel dimension and map the features into C output channels, generating output features containing local and global time domain information.

[0106] Finally, bilinear interpolation is used to transform the output features into the desired image dimensions (H, W). For the target pixel (x, y), the interpolated value I(x, y) is calculated by the following formula:

[0107]

[0108] Among them, (x i ,y i ) are the four nearest coordinates, w ij The pixel values ​​are further scaled to [0, 255] by minimum-maximum normalization to generate the normalized image I norm :

[0109]

[0110] Among them, I raw is the original image generated after interpolation. To prevent the denominator from being 0, ∈=10 -5 The size of the normalized image is B×C×H×W, which is consistent with the input distribution of the visual VLM encoder to achieve effective image feature extraction. The normalized image is processed by the pre-trained visual VLM encoder to obtain the visual features of photovoltaic power generation.

[0111] The text-enhanced learner in the photovoltaic power generation prediction model, the detailed structure is as follows Figure 2 -(c). This module provides contextual text representations containing pre-defined or dynamically generated information, providing flexibility for different PV power generation forecasting needs. For dynamically generated prompts, the text reinforcement learning phase extracts key statistical properties from the input time series, including:

[0112] Range: maximum and minimum values;

[0113] Central tendency: central value;

[0114] Trend: Determine the upward or downward trend based on the first-order difference;

[0115] Periodicity: main frequency components;

[0116] Task context: prediction horizon and historical data window size;

[0117] These features are presented in a structured text format. Figure 2 (c) shows the form of the generated textual prompts. When specific knowledge is available, the textual reinforcement learner incorporates predefined textual descriptions. These descriptions are combined with the dynamically generated prompts to enhance contextual understanding. The final textual representation is processed by the VLM text encoder to generate textual features that complement the visual and temporal features.

[0118] By supporting both predefined and dynamically generated text representations, the Text Enhanced Learner provides a flexible and adaptable mechanism for leveraging textual information in time series forecasting. This adaptability ensures that the model can effectively address a wide range of PV power generation forecasting needs, regardless of the forecast horizon.

[0119] The time domain features, visual features and text features generated through the above steps are fused and then fine-tuned to achieve high-precision medium- and long-term photovoltaic power generation prediction.

[0120] The above specific embodiments do not constitute a limitation to the scope of protection of this application. It should be understood by those skilled in the art that various modifications, combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent replacements and improvements made within the spirit and principles of this application should be included in the scope of protection of this application. In some cases, the actions or steps recorded in this application can be performed in an order different from that in the embodiments and can still achieve the desired results. In addition, the processes depicted in the accompanying drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

Claims

1. A high-precision medium- and long-term photovoltaic power generation prediction method based on multimodal data fusion, characterized by: The method comprises: (1) Collecting the power generation data of photovoltaic equipment over a period of time as photovoltaic power generation history data; (2) Preprocess the collected historical photovoltaic power generation data and convert the data into stable data; (3) Apply the learner in the photovoltaic power generation prediction model to extract multimodal data from the preprocessed photovoltaic power generation historical data and obtain high-level feature representation of multimodality; (4) The photovoltaic power generation prediction model introduces a gating mechanism to integrate multimodal features and uses multimodal interaction to predict photovoltaic power generation data in the future medium and long term.

2. The high-precision medium- and long-term photovoltaic power generation prediction method based on multimodal data fusion according to claim 1, characterized in that: During the data collection process, a smart meter is installed at the output end of the photovoltaic equipment. The collected photovoltaic power generation data includes power generation and power generation, and records the real-time power generation and cumulative power generation of the photovoltaic equipment at each time point; The collected data is transmitted to a computer terminal in real time for subsequent processing. The data is stored in a relational database in the form of a table, where each row represents a data record and each column represents a data field; The power generation data over a period of time is organized into historical data of photovoltaic power generation, which serves as training samples for the photovoltaic power generation prediction model.

3. The high-precision medium- and long-term photovoltaic power generation prediction method based on multimodal data fusion according to claim 1, characterized in that: The main steps of preprocessing the collected photovoltaic power generation historical data include data cleaning, data smoothing and differential processing.

4. The high-precision medium- and long-term photovoltaic power generation prediction method based on multimodal data fusion according to claim 3, characterized in that: The data cleaning step includes removing outliers and filling missing values. First, outliers are detected by box plots, which divide the data into quartiles. Data points that exceed 1.5 times the interquartile range are identified as outliers and replaced with the median. Second, linear interpolation is used to process missing values ​​in the data, and the missing values ​​are filled by calculating the mean of two adjacent data points.

5. The high-precision medium- and long-term photovoltaic power generation prediction method based on multimodal data fusion according to claim 3, characterized in that: The moving average method is used to smooth the data and calculate the average value of each time point and several time points before and after it. The calculation formula is as follows: t , its n-period moving average MA t Expressed as: Among them, n represents the selected time point range; X t-n+1 and X t-n+2 Represents the data at time points t-n+1 and t-n+2; The photovoltaic power generation history data is converted into stable data through differential processing; the difference between the data of adjacent time points is calculated by calculating the first-order difference method. t , its first-order difference ΔX t Expressed as: ΔX t =X t -X t-1 Among them, X t-1 Represents data at the previous point in time.

6. The high-precision medium- and long-term photovoltaic power generation prediction method based on multimodal data fusion according to claim 1, characterized in that: For the preprocessed photovoltaic power generation historical data, the photovoltaic power generation prediction model extracts time series, image and text data through the learner to obtain advanced time domain features, visual features and text features.

7. The high-precision medium- and long-term photovoltaic power generation prediction method based on multimodal data fusion according to claim 6, characterized in that: The cue-enhanced learner in the photovoltaic power generation prediction model divides the historical photovoltaic power generation data into multiple time series data blocks and extracts time domain features from them, maintaining the original memory information of the data; data block embedding interacts through multi-head self-attention and pooling mechanisms to model long-distance dependencies and retain time domain feature information.

8. The high-precision medium- and long-term photovoltaic power generation prediction method based on multimodal data fusion according to claim 6, characterized in that: The image enhancement learner in the photovoltaic power generation prediction model converts time series into information-rich three-channel images through multi-scale convolution, frequency and periodicity encoding; the images are processed by the visual encoder in the pre-trained visual language model to extract hierarchical visual features and capture fine-grained details and high-level temporal patterns.

9. The high-precision medium- and long-term photovoltaic power generation prediction method based on multimodal data fusion according to claim 6, characterized in that: The text-enhanced learner in the photovoltaic power generation prediction model generates contextual text prompts for the input time series, including statistical features, domain-specific contextual information and image descriptions, as text data of the photovoltaic power generation history data; these prompt data are encoded through the text encoder in the pre-trained visual language model to generate text features.

10. The high-precision medium- and long-term photovoltaic power generation prediction method based on multimodal data fusion according to claim 6, characterized in that: The photovoltaic power generation prediction model uses a multimodal fusion network to aggregate temporal, visual, and textual features, leveraging their complementary advantages to achieve high-precision photovoltaic power generation data prediction: Temporal memory embedding F from retrieval-enhanced learner tem The high-level temporal patterns of photovoltaic power generation data are encoded and used as the query vector Q for the cross-modal multi-head attention mechanism; the multimodal embedding F from the two visual language model encoders mm Being used as the key vector K and the value vector V, the cross-modal multi-head attention is defined as: CM-MHA(Q,K,V)=Cat(head1,...,head h )W O Among them, Cat represents splicing according to channel dimension, Q = F tem W Q , K=F mm W K , V=F mm W V ;head i is the i-th self-attention, W O , W Q , W K , W V , W i Q , W i K and W i V is a learnable mapping matrix; d k represents the normalization parameter, which is determined by the feature dimension and the number of heads of the attention mechanism; softmax represents the softmax activation function, which ensures that the weight value is within a specific range; The attention mechanism distributes and aggregates temporal and multimodal features, capturing fine-grained patterns and advanced contextual feature representations. Secondly, residuals and layer normalization LayerNorm are introduced to stabilize the training process: F attn =LayerNorm(F tem +CM_MHA(Q,K,V)) The gating mechanism further obtains the output feature F by dynamically weighting each modality fused : G=σ(W g [F tem ;F mm ]+b g ) F fused =G⊙F attn +(1-G)⊙F mm Among them, G is the gating weight matrix, W g and b g is a learnable parameter, σ(·) is the sigmoid activation function, [;] represents feature concatenation; ⊙ represents element-wise multiplication; These multimodal features are then processed by a fine-tuned predictor to output predicted values ​​of PV power generation data in the future medium- to long-term time periods.