Source load prediction method and device based on multi-modal data, equipment and medium
By performing feature fusion and position coding on multimodal data, multiple training sub-data sets are generated to train the Informer model, which solves the problems of large memory usage and low computational efficiency of source load prediction methods in the prior art, and achieves more efficient source load prediction.
Patent Information
- Application Number
- CN202510601187.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-12
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-05-12
AI Technical Summary
The existing source load prediction method occupies a large amount of memory and has low computational efficiency, making it difficult to effectively process multimodal data.
The preprocessing method based on multimodal data is adopted to feature fusion of text data and cloud image video data, and position encoding is performed through the Informer model, and finally multiple training sub-data sets are generated to train the Informer model to form a source load prediction model.
It improves the generalization ability and robustness of the model, reduces memory usage and computing costs, and improves computing efficiency.
Smart Images

Figure CN120124816A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of source load prediction technology, and in particular to a source load prediction method, device, equipment and medium based on multimodal data. Background Art
[0002] Photovoltaic (PV) power generation is a technology that uses the photovoltaic effect of semiconductor materials to directly convert solar energy into electrical energy. However, its volatility, grid adaptability and sustainable development issues still need to be solved through technological innovation and system optimization. Photovoltaic power generation curve and load curve prediction are key technologies for new energy power systems, and their accuracy directly affects the economy and safety of the power grid. With the advancement of AI technology and the fusion of multi-source data, the prediction methods of photovoltaic power generation curves and load curves are developing from traditional statistical models to deep learning and physical information fusion. In the future, photovoltaic power generation and load forecasting will be more intelligent and automated, supporting the stable operation of power grids with a high proportion of renewable energy.
[0003] Time Series Networks are neural network structures specifically used to process time series data in deep learning. They can effectively capture the temporal dependencies and dynamic changes in photovoltaic power generation data. From the initial exploration of RNN (Recurrent Neural Network), to the optimization of LSTM (Long Short-Term Memory) / GRU (Gated Recurrent Unit), to the innovation of Transformer and diffusion models, time series neural networks have always revolved around the two core issues of modeling and dynamic adaptability. Transformer is a deep learning model based on the self-attention mechanism. It was originally used for natural language processing and has performed well in time series prediction in recent years. However, the model also has some limitations, such as large consumption of computing resources, high complexity of self-attention calculations, low computing efficiency, high training costs, and memory usage. Summary of the invention
[0004] The embodiments of the present invention provide a source load prediction method, device, equipment and medium based on multimodal data, aiming to solve the problems of large memory usage and low computational efficiency of existing source load prediction methods.
[0005] In a first aspect, an embodiment of the present invention provides a source load prediction method based on multimodal data, which includes: Collecting multimodal initial data, and preprocessing the multimodal data to obtain multimodal target data, wherein the multimodal target data includes text data and cloud image video data; Performing feature fusion on the text data and the cloud image video data to obtain fused feature data, and performing position encoding on the fused feature through an Informer model to obtain position encoding data; The position coding data and the fused feature data are used together as a training data set, and the training data set is grouped to obtain a plurality of training sub-data sets; Using multiple groups of the training sub-data sets to train the Informer model to obtain a source load prediction model; The acquired multimodal data to be predicted is input into the source load prediction model for prediction to obtain a prediction result.
[0006] In a second aspect, an embodiment of the present invention further provides a source load prediction device based on multimodal data, which includes: An acquisition and processing unit, used for acquiring multimodal initial data, and preprocessing the multimodal data to obtain multimodal target data, wherein the multimodal target data includes text data and cloud image video data; A fusion encoding unit, used for performing feature fusion on the text data and the cloud image video data to obtain fused feature data, and performing position encoding on the fused feature through an informer model to obtain position encoding data; A grouping unit, used to use the position coding data and the fused feature data together as a training data set, and group the training data set to obtain a plurality of training sub-data sets; A training unit, used for training the Informer model using multiple sets of training sub-data sets to obtain a source load prediction model; The prediction unit is used to input the acquired multimodal data to be predicted into the source-load prediction model for prediction to obtain a prediction result.
[0007] In a third aspect, an embodiment of the present invention further provides a computer device, which includes a memory and a processor, wherein a computer program is stored in the memory, and the processor implements the above method when executing the computer program.
[0008] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium, wherein the storage medium stores a computer program, and the computer program implements the above method when executed by a processor.
[0009] The embodiment of the present invention provides a source load prediction method, device, equipment and medium based on multimodal data. The method includes: collecting multimodal initial data, and preprocessing the multimodal data to obtain multimodal target data, wherein the multimodal target data includes text data and cloud image video data; fusing the text data and the cloud image video data to obtain fused feature data, and position encoding the fused feature through the informer model to obtain position encoding data; using the position encoding data and the fused feature data as a training data set, and grouping the training data set to obtain multiple groups of training sub-data sets; using the multiple groups of the training sub-data sets to train the informer model to obtain a source load prediction model; inputting the acquired multimodal data to be predicted into the source load prediction model for prediction to obtain a prediction result. The technical solution of the embodiment of the present invention first performs feature fusion on text data and cloud image video data in multimodal target data to obtain fused feature data, and position encodes the fused feature data through an Informer model to obtain position encoded data; then, multiple sets of training sub-data sets are generated according to the position encoded data and the fused feature data, and the Informer model is trained with the multiple sets of training sub-data sets to obtain a source load prediction model; finally, the acquired multimodal data to be predicted is input into the source load prediction model for prediction to obtain a prediction result. Since the source load prediction model is a model obtained by training the Informer model with the multiple sets of training sub-data sets generated from the multimodal target data, the generalization ability and robustness of the model are improved. When the source load prediction model is used to perform source load prediction, less memory is occupied and the computational efficiency is high. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying any creative work.
[0011] Figure 1 A schematic flow chart of a source load prediction method based on multimodal data provided by an embodiment of the present invention; Figure 2 A schematic diagram of a sub-process of a source load prediction method based on multimodal data provided by an embodiment of the present invention; Figure 3 A schematic diagram of another sub-process of a source load prediction method based on multimodal data provided by an embodiment of the present invention; Figure 4 A schematic diagram of another sub-process of a source load prediction method based on multimodal data provided by an embodiment of the present invention; Figure 5 It is an overall schematic diagram of the source load prediction model training provided by an embodiment of the present invention; Figure 6 A schematic block diagram of a source load prediction device based on multimodal data provided by an embodiment of the present invention; Figure 7 A schematic block diagram of a computer device provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0012] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0013] It should be understood that when used in this specification and the appended claims, the terms "include" and "comprises" indicate the presence of described features, integers, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or combinations thereof.
[0014] It should also be understood that the terms used in this specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. As used in the specification of the present invention and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an" and "the" are intended to include plural forms.
[0015] It should be further understood that the term "and / or" used in the present description and the appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.
[0016] As used in this specification and the appended claims, the term "if" may be interpreted as "when" or "upon" or "in response to determining" or "in response to detecting," depending on the context. Similarly, the phrases "if it is determined" or "if [described condition or event] is detected" may be interpreted as meaning "upon determination" or "in response to determining" or "upon detection of [described condition or event]" or "in response to detecting [described condition or event]," depending on the context.
[0017] See also Figure 1 , Figure 1 FIG. 1 is a flow chart of a source load prediction method based on multimodal data provided by an embodiment of the present invention. Figure 1As shown, the method includes the following steps S110-S150.
[0018] S110, collecting multimodal initial data, and preprocessing the multimodal data to obtain multimodal target data, wherein the multimodal target data includes text data and cloud image video data.
[0019] In an embodiment of the present invention, multimodal initial data is collected, wherein the multimodal initial data includes initial text data and initial cloud image video data, and the multimodal data is preprocessed to obtain multimodal target data, wherein the multimodal target data includes text data and cloud image video data. It can be understood that the initial text data is preprocessed to obtain the text data, and the initial cloud image video data is preprocessed to obtain the cloud image video data. It should be noted that the preprocessing includes deduplication, denoising, data anomaly, data missing and normalization. The initial text data includes user energy portrait data, date attribute information, historical source load curve, equipment ledger and externally called weather forecast data. The text data is the user energy portrait data, date attribute information, historical source load curve, equipment ledger and externally called weather forecast data after preprocessing. It should also be noted that when collecting the initial text data and the initial cloud image video data, the collection is carried out in units of days. The initial cloud image video data is cloud image video data about weather.
[0020] S120, performing feature fusion on the text data and the cloud image video data to obtain fused feature data, and performing position encoding on the fused feature through an Informer model to obtain position encoding data.
[0021] In an embodiment of the present invention, after obtaining the text data and the cloud image video data, the text data and the cloud image video data are feature fused to obtain fused feature data, and the fused feature data is positionally encoded through the Informer model to obtain positionally encoded data. Specifically, the fused feature data is positionally encoded through the PositionalEncoding algorithm in the Informer model to obtain positionally encoded data. It should be noted that Informer is a deep learning model based on the improved Transformer architecture, which is specially designed for the Long Sequence Time-Series Forecasting (LSTF) task.
[0022] In one embodiment, if Figure 2 As shown, step S120 may specifically include steps S121-S122: S121, extracting features of the cloud image video data through a CNN network to obtain cloud image video features; S122: Fusing the cloud image video features with the text data to obtain the fused feature data.
[0023] In an embodiment of the present invention, the CNN network is a convolutional neural network, which is a deep learning model specifically used to process grid-shaped data (such as images, videos, audio spectrograms, etc.). The CNN network includes a convolution layer and a pooling layer, and the convolution layer includes a first convolution layer, a second convolution layer, a third convolution layer, and a fourth convolution layer. The feature of the cloud image video data is extracted through the CNN network to obtain the cloud image video feature, including: sequentially performing dimensionality reduction processing on the cloud image video data through the first convolution layer, the second convolution layer, the third convolution layer, and the fourth convolution layer to obtain the reduced dimensional cloud image video data; performing maximum pooling processing on the reduced dimensional cloud image video data through the pooling layer to obtain the cloud image video feature. It should be noted that the first convolution layer is a 3*48 convolution layer, the second convolution layer is a 48*128 convolution layer, the third convolution layer is a 128*256 convolution layer, the fourth convolution layer is a 256*512 convolution layer, and the pooling layer is a 2*2 MaxPool pooling layer. It should also be noted that the fusion of the cloud image video features and the text data is specifically to splice the cloud image video features and the text data.
[0024] S130, taking the position coding data and the fused feature data together as a training data set, and grouping the training data set to obtain a plurality of training sub-data sets.
[0025] In an embodiment of the present invention, after obtaining the position coding data, the position coding data and the fusion feature data are used together as a training data set, and the training data set is grouped to obtain multiple groups of training sub-data sets. It should be noted that in this embodiment, the training data set is grouped into groups of 8 data to obtain multiple groups of training sub-data sets, and the training sub-data sets include source-load training data and label data. In the training sub-data set, the first 7 data are the source-load training data, and the last data, that is, the 8th data, is the label data. It can be understood that the last data can be predicted by the first 7 data. It should also be noted that in other embodiments, the training data set is also grouped into groups of 10 data to obtain multiple groups of training sub-data sets. In the training sub-data set, the first 9 data are the source-load training data, and the last data, that is, the 10th data, is the label data, that is, the number of data in the training sub-data set is set according to actual needs and is not specifically limited.
[0026] S140: Using multiple groups of the training sub-data sets to train the Informer model to obtain a source load prediction model.
[0027] In an embodiment of the present invention, the training sub-dataset includes source-load training data and label data, and both the source-load training data and the label data include the fusion feature data and the position coding data. Figure 3 As shown, step S140 may specifically include steps S141-S142: S141, for each group of the training sub-datasets, splicing the fusion feature data and the position coding data in the source load training data to obtain the target source load training data; S142, using the target source load training data and the label data to train the Informer model to obtain the source load prediction model. It can be understood that when the Informer model is trained, the data set is transmitted in groups, that is, each time a group of the training sub-datasets is transmitted, the attention mechanism in the Informer model will calculate the similarity between each data vector in the training sub-dataset and other vectors respectively, and then obtain the final prediction result based on the similarity, and the prediction result and the label data in the training sub-dataset are used to calculate the source load loss, and the parameters in the Informer model are updated through back propagation according to the source load loss, so that the Informer model effectively fits the training data, so that the prediction result is more in line with the expected result. It should be noted that in this embodiment, the Informer model includes a self-attention mechanism module, a sparse attention module, a multi-head attention module, and a cross attention module.
[0028] In one embodiment, if Figure 4 As shown, step S142 may specifically include steps S1421-S1426: S1421, inputting the target source load training data into the self-attention mechanism module to perform data fusion to obtain a dictionary set; S1422, inputting the dictionary set into the sparse attention module to perform dimensionality reduction to obtain a reduced dimensionality dictionary set; S1423, inputting the dimension reduction dictionary set into the multi-head attention module to perform information aggregation to obtain an aggregated dictionary set; S1424, inputting the aggregate dictionary set into the cross attention module to calculate the similarity between the aggregate dictionary set and the target source-load training data to obtain source-load similarity; S1425, calculating a source load loss value according to the source load similarity, the target source load training data and the label data; S1426. Update the parameters of the Informer model through back propagation according to the source load loss value to obtain the source load prediction model.
[0029] In an embodiment of the present invention, for each group of the training sub-data sets, the number of training epochs is set to 20 times. During the training process, the model with the smallest source load loss is saved. The initial learning rate is defined as 0.00001. After the model training starts, the fused feature data and the position encoding data in the source load training data are first spliced to obtain the target source load training data and transmitted to the self-attention mechanism module. The self-attention mechanism module performs data aggregation to obtain a dictionary set, and the dictionary set is transmitted to the sparse attention module for data dimension reduction to obtain a reduced dimension dictionary set, and the reduced dimension dictionary set is transmitted to the multi-head attention module. The multi-head attention module performs information aggregation on the reduced dimension dictionary set to obtain an aggregated dictionary set, and the aggregated dictionary set is transmitted to the cross attention module to calculate the similarity between the aggregated dictionary set and the target source load training data to obtain the source load similarity, and determine the source load prediction result according to the source load similarity and the target source load training data; calculate the error loss between the source load prediction result and the labeled result in the label data to obtain the source load loss value, and based on the source load loss value, update the parameters of the Informer model by back propagation to obtain the source load prediction model.
[0030] S150, inputting the acquired multimodal data to be predicted into the source load prediction model for prediction to obtain a prediction result.
[0031] In an embodiment of the present invention, after the source-load prediction model is trained, the source-load prediction model is saved, a python Restful service is written to load the source-load prediction model, and the external system calls the Restful service interface to pass in the multimodal data to be predicted to obtain the prediction result. It should be noted that in this embodiment, the multimodal data to be predicted includes text data to be predicted and cloud image video data to be predicted. The prediction results include photovoltaic power generation curves and load curves.
[0032] See also Figure 5 , Figure 5 Schematic diagram of the overall source load prediction model training provided by the embodiment of the present invention. Figure 5 As shown, the features of cloud atlas video data are extracted through a CNN network to obtain cloud atlas video features, the cloud atlas video features and text data are fused to obtain fused feature data, the fused feature data are position-encoded to obtain position-encoded data, the position-encoded data and the fused feature data are used together as a training data set, and the training data set is grouped to obtain multiple groups of training sub-data sets, and the Informer model is trained using the multiple groups of training sub-data sets to obtain a source load prediction model.
[0033] To summarize, in this embodiment, since the source load prediction model is a model obtained by training the Informer model based on multiple sets of training sub-data sets generated according to multimodal target data, the generalization ability and robustness of the model are improved. When the source load prediction model is used for source load prediction, less memory is occupied and the computational efficiency is high.
[0034] Figure 6 is a schematic block diagram of a source load prediction device 200 based on multimodal data provided by an embodiment of the present invention. Figure 6 As shown, corresponding to the above source load prediction method based on multimodal data, the present invention also provides a source load prediction device 200 based on multimodal data. The source load prediction device 200 based on multimodal data includes a unit for executing the above source load prediction method based on multimodal data, and the device can be configured in a computer device. Specifically, please refer to Figure 6 The source load prediction device 200 based on multimodal data includes a collection and processing unit 201, a fusion coding unit 202, a grouping unit 203, a training unit 204 and a prediction unit 205.
[0035] Among them, the acquisition and processing unit 201 is used to acquire multimodal initial data, and pre-process the multimodal data to obtain multimodal target data, wherein the multimodal target data includes text data and cloud image video data; the fusion coding unit 202 is used to perform feature fusion on the text data and the cloud image video data to obtain fused feature data, and position encode the fused feature through the Informer model to obtain position coded data; the grouping unit 203 is used to use the position coded data and the fused feature data together as a training data set, and group the training data set to obtain multiple groups of training sub-data sets; the training unit 204 is used to train the Informer model using multiple groups of the training sub-data sets to obtain a source-load prediction model; the prediction unit 205 is used to input the acquired multimodal data to be predicted into the source-load prediction model for prediction to obtain a prediction result.
[0036] In some embodiments, such as this embodiment, the training unit 204 includes a splicing unit and a training sub-unit.
[0037] Among them, the splicing unit is used to splice the fused feature data and the position coding data in the source load training data for each group of the training sub-data sets to obtain the target source load training data; the training sub-unit is used to use the target source load training data and the label data to train the Informer model to obtain the source load prediction model.
[0038] In some embodiments, such as the present embodiment, the training sub-units include a data fusion unit, a dimensionality reduction unit, an information aggregation unit, a first calculation unit, a second calculation unit, and an update unit.
[0039] Among them, the data fusion unit is used to input the target source load training data into the self-attention mechanism module for data fusion to obtain a dictionary set; the dimension reduction unit is used to input the dictionary set into the sparse attention module for dimension reduction to obtain a reduced dimension dictionary set; the information aggregation unit is used to input the reduced dimension dictionary set into the multi-head attention module for information aggregation to obtain an aggregated dictionary set; the first calculation unit is used to input the aggregated dictionary set into the cross attention module to calculate the similarity between the aggregated dictionary set and the target source load training data to obtain the source load similarity; the second calculation unit is used to calculate the source load loss value according to the source load similarity, the target source load training data and the label data; the update unit is used to update the parameters of the Informer model through back propagation according to the source load loss value to obtain the source load prediction model.
[0040] In some embodiments, such as this embodiment, the second calculation unit includes a determination unit and a third calculation unit.
[0041] Among them, the determination unit is used to determine the source load prediction result according to the source load similarity and the target source load training data; the third calculation unit is used to calculate the error loss between the source load prediction result and the annotation result in the label data to obtain the source load loss value.
[0042] In some embodiments, such as this embodiment, the fusion encoding unit 202 includes a feature extraction unit and a fusion unit.
[0043] Among them, the feature extraction unit is used to extract the features of the cloud image video data through the CNN network to obtain the cloud image video features; the fusion unit is used to fuse the cloud image video features with the text data to obtain the fused feature data.
[0044] In some embodiments, such as the present embodiment, the feature extraction unit includes a dimensionality reduction processing unit and a pooling processing unit.
[0045] Among them, the dimensionality reduction processing unit is used to perform dimensionality reduction processing on the cloud image video data through the first convolution layer, the second convolution layer, the third convolution layer and the fourth convolution layer in sequence to obtain reduced dimensionality cloud image video data; the pooling processing unit is used to perform maximum pooling processing on the reduced dimensionality cloud image video data through the pooling layer to obtain the cloud image video features.
[0046] The specific implementation of the source load prediction device 200 based on multimodal data in the embodiment of the present invention corresponds to the above-mentioned source load prediction method based on multimodal data, which will not be repeated here.
[0047] The above-mentioned source load prediction device based on multimodal data can be implemented in the form of a computer program. The computer program can be used in Figure 7 Runs on the computer device shown.
[0048] See also Figure 7 , Figure 7 300 is a schematic block diagram of a computer device provided in an embodiment of the present application. The computer device 300 is a device having a source load prediction function based on multimodal data.
[0049] See also Figure 7 The computer device 300 includes a processor 302 , a memory and a network interface 305 connected via a system bus 301 , wherein the memory may include a storage medium 303 and an internal memory 304 .
[0050] The storage medium 303 may store an operating system 3031 and a computer program 3032. When the computer program 3032 is executed, the processor 302 may execute a source load prediction method based on multi-modal data.
[0051] The processor 302 is used to provide computing and control capabilities to support the operation of the entire computer device 300 .
[0052] The internal memory 304 provides an environment for the operation of the computer program 3032 in the storage medium 303. When the computer program 3032 is executed by the processor 302, the processor 302 can execute a source load prediction method based on multi-modal data.
[0053] The network interface 305 is used to communicate with other devices over the network. Figure 7 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device 300 to which the solution of the present application is applied. The specific computer device 300 may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.
[0054] The processor 302 is used to run a computer program 3032 stored in the memory to implement any embodiment of the above-mentioned source load prediction method based on multimodal data.
[0055] It should be understood that in the embodiment of the present application, the processor 302 may be a central processing unit (CPU), and the processor 302 may also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0056] It is understood by those skilled in the art that all or part of the processes in the method for implementing the above embodiment can be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a storage medium, which is a computer-readable storage medium. The computer program is executed by at least one processor in the computer system to implement the process steps of the embodiment of the above method.
[0057] Therefore, the present invention also provides a storage medium. The storage medium may be a computer-readable storage medium. The storage medium stores a computer program. When the computer program is executed by a processor, the processor executes any embodiment of the source load prediction method based on multimodal data.
[0058] The storage medium may be a USB flash drive, a mobile hard disk, a read-only memory (ROM), a magnetic disk, or an optical disk, etc., which are computer-readable storage media that can store program codes.
[0059] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in terms of function in the above description. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.
[0060] In the several embodiments provided by the present invention, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of each unit is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed.
[0061] The steps in the method of the embodiment of the present invention can be adjusted in order, combined and deleted according to actual needs. The units in the device of the embodiment of the present invention can be combined, divided and deleted according to actual needs. In addition, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0062] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a storage medium. Based on this understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for a computer device (which can be a personal computer, terminal, or network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention.
[0063] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0064] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalents, the present invention is also intended to include these modifications and variations.
[0065] The above is only a specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any technician familiar with the technical field can easily think of various equivalent modifications or replacements within the technical scope disclosed by the present invention, and these modifications or replacements should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention shall be based on the protection scope of the claims.
Claims
1. A source load prediction method based on multimodal data, characterized in that: include: Collecting multimodal initial data, and preprocessing the multimodal data to obtain multimodal target data, wherein the multimodal target data includes text data and cloud image video data; Performing feature fusion on the text data and the cloud image video data to obtain fused feature data, and performing position encoding on the fused feature data through an Informer model to obtain position encoding data; The position coding data and the fused feature data are used together as a training data set, and the training data set is grouped to obtain a plurality of training sub-data sets; Using multiple groups of the training sub-data sets to train the Informer model to obtain a source load prediction model; The acquired multimodal data to be predicted is input into the source load prediction model for prediction to obtain a prediction result.
2. The method according to claim 1, characterized in that The training sub-dataset includes source-load training data and label data, and both the source-load training data and the label data include the fusion feature data and the position coding data. The using multiple sets of the training sub-datasets to train the Informer model to obtain a source-load prediction model includes: For each group of the training sub-data sets, the fusion feature data and the position coding data in the source load training data are spliced to obtain target source load training data; The target source load training data and the label data are used to train the Informer model to obtain the source load prediction model.
3. The method according to claim 2, characterized in that The Informer model includes a self-attention mechanism module, a sparse attention module, a multi-head attention module, and a cross-attention module. The target source load training data and the label data are used to train the Informer model to obtain the source load prediction model, including: Inputting the target source load training data into the self-attention mechanism module to perform data fusion to obtain a dictionary set; Inputting the dictionary set into the sparse attention module to perform dimensionality reduction to obtain a reduced dimensionality dictionary set; Inputting the dimension reduction dictionary set into the multi-head attention module to aggregate information to obtain an aggregated dictionary set; Inputting the aggregate dictionary set into the cross attention module to calculate the similarity between the aggregate dictionary set and the target source-load training data to obtain source-load similarity; Calculate a source load loss value according to the source load similarity, the target source load training data and the label data; The source load prediction model is obtained by updating the parameters of the Informer model through back propagation according to the source load loss value.
4. The method according to claim 3, characterized in that The calculating the loss value according to the source load similarity, the target source load training data and the label data includes: Determine a source load prediction result according to the source load similarity and the target source load training data; The error loss between the source load prediction result and the labeling result in the label data is calculated to obtain the source load loss value.
5. The method according to claim 1, characterized in that The step of fusing the text data and the cloud image video data to obtain fused feature data includes: Extracting features of the cloud image video data through a CNN network to obtain cloud image video features; The cloud image video feature is fused with the text data to obtain the fused feature data.
6. The method according to claim 5, characterized in that The CNN network includes a convolution layer and a pooling layer, the convolution layer includes a first convolution layer, a second convolution layer, a third convolution layer and a fourth convolution layer, and the cloud image video features are obtained by extracting the features of the cloud image video data through the CNN network, including: Performing dimensionality reduction processing on the cloud image video data through the first convolution layer, the second convolution layer, the third convolution layer, and the fourth convolution layer in sequence to obtain reduced-dimensional cloud image video data; The cloud image video features are obtained by performing maximum pooling processing on the dimension-reduced cloud image video data through the pooling layer.
7. The method according to claim 1, characterized in that The preprocessing includes deduplication, denoising, data anomaly, data missing and normalization.
8. A source load prediction device based on multimodal data, characterized in that: include: An acquisition and processing unit, used for acquiring multimodal initial data, and preprocessing the multimodal data to obtain multimodal target data, wherein the multimodal target data includes text data and cloud image video data; A fusion encoding unit, used for performing feature fusion on the text data and the cloud image video data to obtain fused feature data, and performing position encoding on the fused feature through an Informer model to obtain position encoding data; A grouping unit, used to use the position coding data and the fused feature data together as a training data set, and group the training data set to obtain a plurality of training sub-data sets; A training unit, used for training the Informer model using multiple sets of training sub-data sets to obtain a source-load prediction model; The prediction unit is used to input the acquired multimodal data to be predicted into the source-load prediction model for prediction to obtain a prediction result.
9. A computer device, characterized in that: The computer device includes a memory and a processor, the memory stores a computer program, and the processor implements the method according to any one of claims 1 to 7 when executing the computer program.
10. A computer-readable storage medium, characterized in that: The storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Commercial building load prediction method and system based on Informer network
CN116979503A