Multi-modal power price analysis method, system and device and storage medium
Through the multimodal data processing method, data of different modes are combined, and pre-trained and iteratively optimized using ViT model and self-supervised comparison learning, solving the problem of poor prediction effect of nonlinear fluctuations of power prices in the existing technology, and achieving higher prediction accuracy and accuracy.
Patent Information
- Application Number
- CN202510374665.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-27
- Publication Date
- 2025-06-27
AI Technical Summary
The existing power market transaction price analysis method has poor prediction effect when dealing with nonlinear fluctuations in power prices, with large deviations in the results and low prediction accuracy.
Multimodal data processing method is adopted to combine data of different modes such as time series, text, images and sound, and pre-training and iterative optimization through ViT model and self-supervised comparison learning to build the optimal power price prediction model.
It significantly improves the accuracy of power price prediction, reduces model complexity, can effectively capture cross-modal correlation, improves prediction accuracy, and provides reliable decision support for power market participants.
Smart Images

Figure CN120218980A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of power market analysis, and particularly relates to a multi-modal power price analysis method, system, device, and storage medium. Background Art
[0002] In recent years, with the rapid development and utilization of clean energy such as wind energy and solar energy, the installed capacity of wind power and photovoltaic has increased sharply, and thus large-scale new energy is gradually incorporated into the power grid. However, due to the low power generation cost of new energy such as wind power and photovoltaic, the trading price in the power market has shown a downward trend; moreover, electric energy cannot be stored in large quantities, and to always meet the supply-demand balance, the randomness and instability of new energy such as wind and light make the trading price in the power market fluctuate strongly, showing non-linear characteristics. And the trading price in the power market is a complex and variable index, which is affected by many factors. These factors are not only diverse but also often intertwined and jointly act on the fluctuation of power prices. Among these factors, time series data, text data, image data, and sound data have the most prominent influence on the trading price in the power market.
[0003] At present, the trading price analysis methods in the power market can be divided into statistical analysis methods. Statistical analysis methods include generalized autoregressive conditional heteroskedasticity, time series, etc., which realize electricity price prediction by learning the recursive relationship of trading prices at different times. However, these methods do not have good prediction effects for non-linear sequences, and when dealing with the non-linear fluctuation of power prices, the analysis results have large deviations and the prediction accuracy is low. Summary of the Invention
[0004] In order to solve the problem that the existing traditional prediction methods have large deviations in the analysis results when dealing with the non-linear fluctuation of power prices, resulting in low prediction accuracy, the present invention provides a multi-modal power price analysis method, system, device, and storage medium. By combining data of different modalities (time series, text, image, sound, etc.), and through a multi-modal processing model and feature extraction with a shared encoder-decoder parameter, the accuracy of power price prediction is enhanced.
[0005] To achieve the above object, the present invention provides the following technical solutions: The present invention proposes a multi-modal power price analysis method, including the following steps: Divide the data of different modalities obtained into a training data set and a test data set; Construct an initial ViT model, first train it on the LAION-2B data set based on the initial ViT model, and then train it on the training data set, and perform contrastive learning using self-supervised learning methods for pre-training to obtain a multi-modal data processing model; Preprocess the test data set to obtain multimodal data, and iteratively train the multimodal data processing model based on the multimodal data to obtain an optimal prediction model; Input the currently obtained different-modal data into the optimal prediction model to predict electricity price data.
[0006] Preferably, the different-modal data includes audio data, image data, and text data, and both the training data set and the test data set contain audio data, the image data, and the text data.
[0007] Preferably, constructing the training data set and the test data set from the obtained different-modal data includes: Obtain different-modal data, sort the different-modal data in chronological order to obtain a time series data set, and process the time series data set through a sliding window technique to be segmented into a training data set and a test data set.
[0008] Preferably, constructing the initial ViT model, and after training the initial ViT model based on the LAION-2B data set first, then training it through the training data set, and performing contrastive learning using self-supervised learning methods for pre-training to obtain a multimodal data processing model, includes: Perform normalization processing on the training data set to obtain training multimodal data; Construct an initial ViT model based on the ViT backbone network; Input the obtained LAION-2B data set into the initial ViT model for training; Batch select data from the training multimodal data and input it into the initial ViT model after training with the LAION-2B data set; The initial ViT model generates enhanced views for each patch in the training multimodal data; Obtain the embedding representation of each enhanced view, calculate the contrastive loss based on the embedding representation, according to the contrastive loss, use the backpropagation algorithm to calculate the gradient data of the model parameters, and use the optimizer to update the parameters of the initial ViT model according to the calculated gradient data to obtain a prediction model, and freeze the parameters of the backbone network in the prediction model.
[0009] Preferably, performing normalization processing on the training data set to obtain training multimodal data includes: Convert the image data in the training data set to a unified size and unified format to obtain a normalized training data set; Segment the image data in the normalized training data set into patches of a defined size, and associate the text data and audio data corresponding to the patches; Convert the image patches and their associated text data and audio data into embedding vectors respectively to obtain training multi-modal data.
[0010] Preferably, preprocess the test data set to obtain multi-modal data, including: Use natural language processing technology to segment the text data into multiple text segments, and then convert each text segment into corresponding text multi-modal data; Cut the image data into image patches, and convert each image patch into a high-dimensional vector to obtain image multi-modal data; Convert the audio data into a logarithmic Mel spectrogram to obtain audio multi-modal data; Segment the time series data into a series of time windows, define the data within each window as a data point, and convert all data points into time series multi-modal data through normalization; Integrate the time series multi-modal data, the audio multi-modal data, the image multi-modal data, and the text multi-modal data to obtain multi-modal data.
[0011] Preferably, iteratively train a multi-modal data processing model based on the multi-modal data to obtain an optimal prediction model, including: Input the multi-modal data into the prediction model to capture different features in the multi-modal data and obtain first feature data analysis intermediate data; Perform residual connection and normalization processing on the first feature data analysis intermediate data to obtain first multi-modal associated feature data analysis intermediate data; Perform a high-dimensional space non-linear transformation on the first multi-modal associated feature data analysis intermediate data to obtain first fusion feature data; Perform residual connection and normalization processing on the first fusion feature data and the first multi-modal associated feature data to obtain second fusion feature vector data; Pass the second fusion feature vector data through the prediction model to mask the information of subsequent moments for each state as the current moment to obtain second multi-modal associated feature data; Process the second multi-modal associated feature data through the prediction model to obtain network structure parameters; Perform a high-dimensional space non-linear transformation on the network structure parameters to obtain third multi-modal associated feature data; Process the third multi-modal associated feature data through a fully connected layer and output to obtain electricity price prediction data; Compare the electricity price prediction data with the actual electricity price data to obtain an error value and iteratively optimize the network parameters multiple times; Adjust the parameters of the prediction model based on the optimized network parameters to obtain an optimal prediction model.
[0012] The present invention proposes a multimodal electricity price analysis system, and applies a multimodal electricity price analysis method, including: An acquisition module, configured to: Obtain data of different modalities and current data of different modalities; A set construction module, configured to: Split the data of different modalities obtained into a training data set and a test data set; A model construction module, configured to: Construct an initial ViT model, first train it on the LAION-2B data set based on the initial ViT model, and then train it on the training data set, and perform contrastive learning using self-supervised learning methods for pre-training to obtain a multimodal data processing model; An optimization module, configured to: Preprocess the test data set to obtain multimodal data, and iteratively train the multimodal data processing model based on the multimodal data to obtain an optimal prediction model; An output head module, configured to: Input the currently obtained data of different modalities into the optimal prediction model to predict electricity price data.
[0013] The present invention proposes an electronic device, including a memory, a processor, and a computer program stored in the memory and executable in the processor. When the processor executes the computer program, the steps of the above-mentioned multimodal electricity price analysis method are implemented.
[0014] The present invention proposes a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the steps of the above-mentioned multimodal electricity price analysis method are implemented.
[0015] Compared with the prior art, the present invention has the following beneficial technical effects: The present invention proposes a multi-modal electricity price analysis method. This method uses the large-scale publicly available dataset LAION-2B for pre-training, enabling the Vision Transformer (ViT) model to fully capture the general feature representations between image and text modalities, laying a solid foundation for subsequent task-specific training. Then, a self-supervised contrastive learning paradigm is adopted to conduct secondary training on multi-modal task data, effectively solving the problem of scarce labeled data, while enhancing the model's ability to understand the internal correlations of different modal data. Through an iterative optimization process, the model can dynamically adjust parameters to adapt to the complex and changing electricity market dynamics. Finally, an optimal model with strong interpretability and prediction ability is constructed, breaking through the expression limitations of single-modal data. Through the complementary integration of multi-source heterogeneous information, the prediction accuracy is significantly improved; its two-stage training strategy also achieves a balance between knowledge transfer and task specialization, enabling the model to not only have the ability to extract general features but also accurately capture the specific laws in the field of electricity prices. The end-to-end training-prediction process ensures the coherence and efficiency of data processing, providing a decision-making support tool with both timeliness and reliability for electricity market participants.
[0016] Furthermore, the present invention analyzes data of different modalities through a prediction model, fully utilizing the rich information contained in multi-source heterogeneous data, providing a more comprehensive basis for prediction, enhancing the ability to capture position information in time series data and structured data, and helping to better mine the temporal and structural dependencies in the data. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 It is a schematic flowchart of a multi-modal electricity price analysis method proposed by the present invention; Figure 2 It is a schematic diagram of a computer device provided by an embodiment of the present invention; Figure 3 It is a block diagram of a chip provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0018] In the following, only some exemplary embodiments are briefly described. As those skilled in the art can recognize, the described embodiments can be modified in various different ways without departing from the spirit or scope of the present invention. Therefore, the drawings and the description are considered to be exemplary in nature rather than restrictive.
[0019] In the description of the present invention, it should be understood that the orientation or positional relationship indicated by the terms "center", "longitudinal", "transverse", "length", "width", "thickness", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", "clockwise", "counterclockwise", "axial", "radial", "circumferential", etc. is based on the orientation or positional relationship shown in the drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation on the present invention.
[0020] In addition, the terms "first" and "second" are only used for descriptive purposes and should not be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the present invention, "a plurality of" means two or more unless otherwise specifically defined.
[0021] In the present invention, unless otherwise clearly specified and limited, the terms "mounted", "connected", "coupled", "fixed", etc. should be construed in a broad sense. For example, it may be a fixed connection, a detachable connection, or integrated; it may be a mechanical connection, an electrical connection, or a communication connection; it may be directly connected, or indirectly connected through an intermediate medium, and may be the internal communication of two elements or the interaction relationship between two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.
[0022] In the present invention, unless otherwise clearly specified and limited, the first feature being "above" or "below" the second feature may include the direct contact between the first and second features, or may include the situation where the first and second features are not in direct contact but in contact through other features therebetween. Moreover, the first feature being "above", "over" and "on" the second feature includes that the first feature is directly above or obliquely above the second feature, or merely means that the horizontal height of the first feature is higher than that of the second feature. The first feature being "under", "beneath" and "under" the second feature includes that the first feature is directly below or obliquely below the second feature, or merely means that the horizontal height of the first feature is lower than that of the second feature.
[0023] The embodiments of the present invention will be described in detail below with reference to the accompanying drawings.
[0024] In recent years, with the rapid development and utilization of clean energy such as wind energy and solar energy, the installed capacity of wind power and photovoltaic power has increased sharply. As a result, a large amount of new energy is gradually incorporated into the power grid. However, due to the low power generation cost of new energy such as wind power and photovoltaic power, the trading price in the power market shows a downward trend. Moreover, electricity cannot be stored in large quantities, and the supply and demand balance needs to be maintained at all times. The randomness and instability of new energy such as wind and light make the trading price in the power market fluctuate strongly, showing non-linear characteristics. The trading price in the power market is a complex and variable indicator, which is affected by many factors. These factors are not only diverse but also often intertwined, jointly affecting the fluctuation of electricity prices. Among these factors, time series data, text data, image data, and sound data have the most prominent impact on the trading price in the power market.
[0025] At present, the analysis methods for the trading price in the power market can be divided into statistical analysis methods. Statistical analysis methods include generalized autoregressive conditional heteroskedasticity, time series, etc. By learning the recursive relationship of trading prices at different times, electricity price prediction is realized. However, these methods do not perform well in predicting non-linear sequences. When dealing with the non-linear fluctuation of electricity prices, the analysis results have large deviations and low prediction accuracy.
[0026] In view of the above problems, the present invention proposes a multi-modal electricity price analysis method. This method unifies data of different modalities and inputs them into a prediction model for feature extraction, and predicts the electricity price, which enhances the accuracy of electricity price prediction. It not only reduces the complexity of the model but also can effectively capture cross-modal correlations, thereby improving the accuracy of electricity price prediction. The specific process is as follows:
[0027]
[0028] Among them, , , respectively represent the data processing processes in the prediction model.
[0029] The present invention proposes a multi-modal electricity price analysis method. As Figure 1 shown, it includes the following steps: Obtain data of different modalities, and construct the data of different modalities into a training data set and a test data set; that is, obtain audio data, image data, and text data, sort the obtained audio data, image data, and text data in chronological order to obtain a time series data set, and process the time series data set through a sliding window technique to be segmented into a training data set and a test data set. Among them, in the sliding window technique, the time series data set is segmented into a series of time windows, the data in each window is defined as a data point, and all data points are converted into time series multimodal data after normalization processing; then, the time series multimodal data is constructed into a training data set and a test data set. Among them, both the training data set and the test data set contain audio data, image data, and text data; Preprocess the training data set, construct an initial ViT (Vision Transformer) model, first train it with the LAION-2B data set, and then input the preprocessed training data set into the initial ViT model trained with the LAION-2B data set, and use self-supervised learning methods for contrastive learning for pre-training to obtain a multimodal data processing model; Exemplarily, standardize the image data in the training data set so that the image data in the training data set is converted to a unified size and format to obtain a standardized training data set, and segment the image data in the standardized training data set into tiles of a defined size, and associate the text data and audio data corresponding to the tiles. Convert the tiles and their associated text data and audio data into embedding vectors respectively to obtain training multimodal data; Construct an initial ViT model based on the ViT backbone network, that is, use ViT as the backbone network, and create a Transformer encoding module, a Transformer decoding module, and a position encoding module; The Transformer encoding module includes multiple encoding processing sub-modules, and the multiple encoding processing sub-modules are connected in series to process the data in sequence to form a first iterative data channel. Each encoding processing sub-module includes a first encoding residual block and a second encoding residual block. The data first passes through the first encoding residual block and then is input into the second encoding residual block for processing. The first encoding residual block includes a multi-head attention mechanism layer, and the multi-head attention mechanism layer uses residual connection and then performs layer normalization. The second encoding residual block includes a feed-forward neural network layer, and the feed-forward neural network layer includes two linear fully connected layers. The two linear fully connected layers use residual connection and are associated with the GELU non-linear activation function, and then perform layer normalization; The Transformer decoding module includes multiple decoding processing sub-modules connected in series to process data sequentially, forming a second iterative data channel. Each decoding processing sub-module includes a first decoding residual block and a second decoding residual block. The data first passes through the first decoding residual block and then is input into the second decoding residual block for processing. The first decoding residual block includes a multi-head attention mechanism layer and an information masking multi-head attention layer. The multi-head attention mechanism layer and the information masking multi-head attention layer use residual connections and then perform layer normalization. The second decoding residual block includes a feed-forward neural network layer, which includes two linear fully connected layers. The two linear fully connected layers use residual connections and are associated with the GELU non-linear activation function, and then perform layer normalization. Among them, the information masking multi-head attention layer masks future positions in the attention matrix, masking the information of subsequent moments for each state as the current moment, simulating the situation of segmenting the known in the past and the unknown in the future in actual situations. The Transformer decoding module simulates the data at each moment as the current moment. Exemplarily, assuming the time is t, the analysis content of the decoder should only depend on the information before the t moment and its own output, rather than the information after that. The masked MSA masks future positions in the attention matrix by using the sequence masking method, hiding the information after t so that it cannot be used at the t moment.
[0030] The position encoding module often uses sine and cosine functions. The position encoding module includes a first position encoding module and a second position encoding module. The first position encoding module and the second position encoding module respectively perform position encoding on the data to obtain encoded multi-modal data. The formula is as follows:
[0031]
[0032] Among them, pos represents the position of the data, i represents the feature index, d is the feature dimension, and 2i and 2i + 1 distinguish whether the position is the odd or even feature.
[0033] The ViT initial model is constructed based on the Transformer encoding module, the Transformer decoding module, and the position encoding module. In the ViT initial model, the input channel of the second decoding residual block of the Transformer decoding module is connected to the output channel of the second encoding residual block of the Transformer encoding module. The input channel of the first decoding residual block of the Transformer decoding module is connected to the output channel of the second position encoding module. The input channel of the first encoding residual block of the Transformer encoding module is connected to the output channel of the first position encoding module. Input the obtained LAION-2B dataset into the initial ViT model for training. Batch select data from the training multi-modal data and input it into the initial ViT model trained with the LAION-2B dataset. The initial ViT model generates two augmented views for each patch, obtains the embedding representation of each augmented view, calculates the contrastive loss through the embedding representation, calculates the gradient data of the model parameters according to the contrastive loss using the backpropagation algorithm, updates the parameters of the initial ViT model using the optimizer according to the calculated gradient data to obtain a prediction model, and freezes the parameters of the backbone network in the prediction model.
[0034] Preprocess the test dataset to obtain multi-modal data; Exemplarily, process audio data, image data, and text data respectively; that is, use natural language processing techniques, such as WordPiece, to split the text data into multiple text segments, and then convert each text segment into corresponding text multi-modal data; Among them, the text data includes policy information and market reports, and the policy information and market reports can provide important information about market dynamics and regulatory changes.
[0035] Slice the image data into image patches, and convert each image patch into a high-dimensional vector to obtain image multi-modal data; The image data includes satellite maps, circuit grid structure diagrams, and infrared detection image data; Among them, the satellite map can capture the geographical environment of the power station, the circuit grid structure diagram can perform power flow analysis to obtain the blocking situation, and the infrared detection image data can detect abnormal heating of equipment.
[0036] The formulaic expression of the image processing process is:
[0037] where H and W are the height and width of the image respectively, is the number of channels, is the size of each patch, is the number of flat patches.
[0038] Convert the audio data into a logarithmic Mel spectrogram to obtain audio multi-modal data. The audio data includes vibration data and electromagnetic wave signal data of a power generation device. Specifically, the process includes: segmenting the audio data to obtain segmented audio data, performing Fourier transform on the segmented audio data to obtain spectral data; using a Mel-scale filter bank to convert the spectral data, mapping the frequency characteristics to the Mel scale to obtain a Mel spectrogram, making the spectrogram more in line with human auditory perception; taking the logarithm of the Mel spectrogram to amplify the differences in low-intensity frequencies to obtain a logarithmic Mel spectrogram. In the logarithmic Mel spectrogram, one dimension represents time and the other dimension represents frequency; finally, using time windowing and frequency segmentation methods, divide the logarithmic Mel spectrogram into small blocks in the time and frequency dimensions to obtain audio multi-modal data.
[0039] Aggregate the audio multi-modal data, text multi-modal data, and image multi-modal data to obtain multi-modal data.
[0040] Input the multi-modal data into the prediction model. The prediction model first processes the multi-modal data through a multi-head attention mechanism, uses ViT features as the input benchmark for the Key / Value matrix, performs cross-region relationship modeling on the ViT output tokens, then calculates the weight distribution to capture and encode different features in the multi-modal data to obtain the first feature data analysis intermediate data; then perform residual connection and normalization processing on the first feature data analysis intermediate data, scale the data with different features to the same numerical range, and remove the differential data, enabling the model to more easily learn the correlation between features during training to obtain the first multi-modal associated feature data analysis intermediate data; the first multi-modal associated feature data passes through the prediction model for high-dimensional space non-linear transformation to obtain the first fusion feature data, perform residual connection and normalization processing on the first fusion feature data and the first multi-modal associated feature data to obtain the second fusion feature vector data, process the second fusion feature vector data through the masked MSA layer, mask the information of subsequent moments for each moment as the current state, simulate the actual situation of dividing the known past and the unknown future, obtain the third feature data, then scale the third feature data to the same numerical range and remove the differential data to obtain the second multi-modal associated feature data, pass the second multi-modal associated feature data through the prediction model to obtain the second fusion feature vector data, then input it into the multi-head attention mechanism layer, residual connection and normalization layer to obtain the network structure parameters for analyzing past data to obtain future predictions with the knowledge background of the full amount of data, then perform high-dimensional space non-linear transformation on the network structure parameters to optimize the complete network structure to obtain the third multi-modal associated feature data, and finally process it through the fully connected layer to output the electricity price prediction data. Compare the electricity price prediction data with the actual electricity price data to obtain the error value and iteratively optimize the network parameters multiple times.
[0041] Based on the optimized network parameters, adjust the parameters of the prediction model to obtain the optimal prediction model. Exemplarily, unfreeze the frozen parameters in the prediction model, adopt an adaptive adjustment strategy for the optimized network parameters, and determine the adjustment parameters of the prediction model. Among the determined adjustment parameters, if the adjustment parameter is negative, it indicates that the value of the adjustment parameter needs to be reduced based on the frozen parameters; if the adjustment parameter is positive, it indicates that the value of the adjustment parameter needs to be increased based on the frozen parameters, so as to obtain the optimal prediction model.
[0042] When performing real-time electricity price analysis, obtain current data of different modalities, and input the data of different modalities into the optimal prediction model to predict electricity price data.
[0043] In another embodiment of the present invention, a multi-modal electricity price analysis system is provided for implementing the above-mentioned multi-modal electricity price analysis method, including an acquisition module, a set construction module, a model construction module, an optimization module, and an output head module; Among them, the acquisition module is configured to: Obtain data of different modalities and current data of different modalities; The set construction module is configured to: Divide the obtained data of different modalities into a training data set and a test data set; The model construction module is configured to: Construct an initial ViT model, first train it on the LAION-2B data set based on the initial ViT model, and then train it on the training data set, and perform contrast learning using self-supervised learning methods for pre-training to obtain a multi-modal data processing model; The optimization module is configured to: Preprocess the test data set to obtain multi-modal data, and perform iterative training on the multi-modal data processing model based on the multi-modal data to obtain the optimal prediction model; The output head module is configured to: Input the obtained current data of different modalities into the optimal prediction model to predict electricity price data.
[0044] In another embodiment of the present invention, a terminal device is provided. The terminal device includes a processor and a memory. The memory is used to store a computer program. The computer program includes program instructions. The processor is used to execute the program instructions stored in the computer storage medium. The processor may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. It is the computing core and control core of the terminal, and is suitable for implementing one or more instructions. Specifically, it is suitable for loading and executing one or more instructions to implement the corresponding method flow or corresponding function. The processor described in the embodiment of the present invention can be used for the operation of a multi-modal electricity price analysis method, including: Dividing the data obtained in different modalities into a training data set and a test data set; constructing an initial ViT model, training it on the LAION-2B data set based on the initial ViT model first, and then training it on the training data set, and performing contrastive learning using self-supervised learning methods for pre-training to obtain a multi-modal data processing model; preprocessing the test data set to obtain multi-modal data, and iteratively training the multi-modal data processing model based on the multi-modal data to obtain an optimal prediction model; inputting the currently obtained data in different modalities into the optimal prediction model to predict electricity price data.
[0045] In another embodiment of the present invention, the present invention also provides a storage medium, specifically a computer-readable storage medium (Memory). The computer-readable storage medium is a memory device in the terminal device, used to store programs and data. It can be understood that the computer-readable storage medium here can include both the built-in storage medium in the terminal device and, of course, the extended storage medium supported by the terminal device. The computer-readable storage medium provides a storage space, and this storage space stores the operating system of the terminal. And, in this storage space, one or more instructions suitable for being loaded and executed by the processor are also stored. These instructions can be one or more computer programs (including program codes). It should be noted that the computer-readable storage medium here can be a high-speed RAM memory, or a non-volatile memory, such as at least one disk memory.
[0046] One or more instructions stored in a computer-readable storage medium can be loaded and executed by a processor to implement the corresponding steps of the multi-modal electricity price analysis method in the above embodiments; one or more instructions in the computer-readable storage medium are loaded and executed by the processor to perform the following steps: Split the data obtained in different modalities into a training data set and a test data set; construct an initial ViT model, train it on the LAION-2B data set based on the initial ViT model first, and then train it on the training data set. Use self-supervised learning methods for contrastive learning for pre-training to obtain a multi-modal data processing model; preprocess the test data set to obtain multi-modal data, and iteratively train the multi-modal data processing model based on the multi-modal data to obtain an optimal prediction model; input the currently obtained different-modal data into the optimal prediction model to predict electricity price data.
[0047] Please refer to Figure 2 The terminal device is a computer device. The computer device 60 in this embodiment includes: a processor 61, a memory 62, and a computer program 63 stored in the memory 62 and executable on the processor 61. When the computer program 63 is executed by the processor 61, it implements the reservoir stimulation wellbore fluid composition calculation method in the embodiment. To avoid repetition, it will not be elaborated here one by one. Alternatively, when the computer program 63 is executed by the processor 61, it implements the functions of each model / unit in the reservoir stimulation wellbore fluid composition calculation system in the embodiment. To avoid repetition, it will not be elaborated here one by one.
[0048] The computer device 60 can be a computing device such as a desktop computer, a notebook, a palm computer, and a cloud server. The computer device 60 may include, but is not limited to, a processor 61 and a memory 62. Those skilled in the art can understand that Figure 2 This is only an example of the computer device 60 and does not constitute a limitation on the computer device 60. It may include more or fewer components than shown in the figure, or combine certain components, or different components. For example, the computer device may also include input / output devices, network access devices, buses, etc.
[0049] The so-called processor 61 may be a Central Processing Unit (CPU), or it may also be other general-purpose processors, central processors, graphics processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, quantum computing-based data processing logic units, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0050] The memory 62 may be an internal storage unit of the computer device 60, such as the hard disk or memory of the computer device 60. The memory 62 may also be an external storage device of the computer device 60, such as a plug-in hard disk equipped on the computer device 60, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc.
[0051] Furthermore, the memory 62 may also include both an internal storage unit of the computer device 60 and an external storage device. The memory 62 is used to store computer programs and other programs and data required by the computer device. The memory 62 may also be used to temporarily store data that has been output or is to be output.
[0052] In each of the embodiments provided in this application, any reference to a memory, database, or other medium may include at least one of non-volatile and volatile memories. Non-volatile memory may include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory may include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.
[0053] In each of the embodiments provided in this application, the database involved may include at least one of a relational database and a non-relational database. The non-relational database may include a distributed database based on blockchain, etc., without limitation. In each of the embodiments provided in this application, the processor involved may be a general-purpose processor, a central processing unit, a graphics processing unit, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, etc., without limitation.
[0054] Please refer to Figure 3 , the terminal device is a chip. The chip 600 of this embodiment includes a processor 622, the number of which can be one or more, and a memory 632 for storing computer programs executable by the processor 622. The computer programs stored in the memory 632 may include one or more modules each corresponding to a set of instructions. In addition, the processor 622 may be configured to execute the computer program to perform the above-mentioned generalizable general monocular absolute depth map estimation method.
[0055] In addition, the chip 600 may further include a power supply component 626 and a communication component 650. The power supply component 626 may be configured to perform power management of the chip 600, and the communication component 650 may be configured to implement communication of the chip 600, for example, wired or wireless communication. In addition, the chip 600 may further include an input / output interface 658. The chip 600 may operate based on an operating system stored in the memory 632.
[0056] The foregoing has shown and described the basic principles, main features and advantages of the present invention. For a person skilled in the art, it is obvious that the present invention is not limited to the details of the above-described exemplary embodiments, and without departing from the spirit or basic features of the present invention, the present invention can be implemented in other specific forms. Therefore, in any aspect, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be embraced within the present invention. Any reference signs in the claims should not be construed as limiting the claims involved.
[0057] In addition, it should be understood that although this specification is described in terms of embodiments, not every embodiment only contains an independent technical solution. This narrative manner of the specification is only for clarity. A person skilled in the art should regard the specification as a whole. The technical solutions in the various embodiments can also be appropriately combined to form other embodiments that can be understood by a person skilled in the art. The above content is only to illustrate the technical idea of the present invention and cannot be used to limit the protection scope of the present invention. Any modification made on the basis of the technical solution according to the technical idea proposed by the present invention falls within the protection scope of the claims of the present invention.
Claims
1. A multimodal electricity price analysis method, characterized in that: The following steps are involved: The data obtained from different modalities are divided into training data sets and test data sets; Constructing a ViT initial model, first training the ViT initial model with the LAION-2B data set, then training it with the training data set, using a self-supervised learning method for comparative learning and pre-training to obtain a multimodal data processing model; Preprocessing the test data set to obtain multimodal data, iteratively training a multimodal data processing model based on the multimodal data to obtain an optimal prediction model; The current different modal data are obtained and input into the optimal prediction model to predict the electricity price data.
2. A multi-modal electricity price analysis method according to claim 1, characterized in that: The data of different modalities include audio data, image data and text data, and the training data set and the test data set both include audio data, image data and text data.
3. A multi-modal electricity price analysis method according to claim 2, characterized in that: The data of different modalities are obtained to construct training data sets and test data sets, including: Data of different modes are obtained and sorted in chronological order to obtain a time series data set. The time series data set is processed by sliding window technology and divided into a training data set and a test data set.
4. A multi-modal electricity price analysis method according to claim 3, characterized in that: The ViT initial model is constructed by first training the ViT initial model with the LAION-2B data set and then training with the training data set, and pre-training is performed by comparative learning using a self-supervised learning method to obtain a multimodal data processing model, including: Standardizing the training data set to obtain training multimodal data; Build the ViT initial model based on the ViT backbone network; Input the acquired LAION-2B data set into the ViT initial model for training; Batch selecting data from the training multimodal data, and inputting the data into the ViT initial model trained with the LAION-2B dataset; The ViT initial model generates an enhanced view for each tile in the training multimodal data; Obtain an embedded representation of each of the enhanced views, calculate the contrast loss based on the embedded representation, calculate the gradient data of the model parameters using a back propagation algorithm based on the contrast loss, use an optimizer to update the ViT initial model parameters according to the calculated gradient data, obtain a prediction model, and freeze the parameters of the backbone network in the prediction model.
5. A multi-modal electricity price analysis method according to claim 1, characterized in that: The training data set is standardized to obtain training multimodal data, including: Converting the image data in the training data set into a uniform size and a uniform format to obtain a standardized training data set; Divide the image data in the standardized training data set into tiles of defined size, and associate the text data and audio data corresponding to the tiles; The image tiles and their associated text data and audio data are converted into embedding vectors to obtain training multimodal data.
6. A multi-modal electricity price analysis method according to claim 5, characterized in that: The test data set is preprocessed to obtain multimodal data, including: Using natural language processing technology to segment the text data into multiple text segments, and then converting each text segment into corresponding text multimodal data; Dividing the image data into image blocks, converting each image block into a high-dimensional vector to obtain image multimodal data; Convert the audio data into a logarithmic Mel spectrum to obtain audio multimodal data; Split the time series data into a series of time windows, define the data in each window as a data point, and convert all data points into time series multimodal data after normalization; The time series multimodal data, the audio multimodal data, the image multimodal data and the text multimodal data are integrated to obtain multimodal data.
7. A multi-modal electricity price analysis method according to claim 1, characterized in that: The multimodal data processing model is iteratively trained based on the multimodal data to obtain the optimal prediction model, including: Inputting the multimodal data into the prediction model, capturing different features in the multimodal data, and obtaining first feature data analysis intermediate data; Performing residual connection and normalization processing on the first feature data analysis intermediate data to obtain first multimodal association feature data analysis intermediate data; The first multimodal association feature data analysis intermediate data performs a high-dimensional space nonlinear transformation to obtain first fused feature data; Performing residual connection and normalization processing on the first fused feature data and the first multimodal associated feature data to obtain second fused feature vector data; Passing the second fused feature vector data through a prediction model, masking the information of subsequent moments for each state at the current moment, and obtaining second multimodal associated feature data; The second multimodal association feature data is processed by a prediction model to obtain network structure parameters; Performing a high-dimensional spatial nonlinear transformation on the network structure parameters to obtain third multimodal association feature data; The third multimodal associated feature data is processed by a fully connected layer, and output to obtain power price prediction data; Compare the electricity price forecast data with the actual electricity price data to obtain an error value and iterate and optimize network parameters multiple times; The parameters of the prediction model are adjusted based on the optimized network parameters to obtain the optimal prediction model.
8. A multimodal electricity price analysis system, applying a multimodal electricity price analysis method according to any one of claims 1 to 7, characterized in that: include: Get module, configured as: Used to obtain data of different modalities and current data of different modalities; The collection building block is configured as: Used to split the data acquired from different modalities into training data sets and test data sets; The model building module is configured to: For constructing a ViT initial model, the ViT initial model is first trained with the LAION-2B data set, then trained with the training data set, and pre-trained by comparative learning using a self-supervised learning method to obtain a multimodal data processing model; The optimization module is configured as follows: Used to preprocess the test data set to obtain multimodal data, and iteratively train the multimodal data processing model based on the multimodal data to obtain an optimal prediction model; The output header module is configured as: It is used to obtain the current different modal data and input them into the optimal prediction model to predict the electricity price data.
9. An electronic device, characterized in that: The invention comprises a memory, a processor and a computer program stored in the memory and executable in the processor, wherein the processor implements the steps of a multimodal electricity price analysis method as claimed in any one of claims 1 to 7 when executing the computer program.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of a multimodal electricity price analysis method according to any one of claims 1 to 7 are implemented.
Citation Information
Cited By
Power generation strategy adjustment method and device based on multi-modal electricity price prediction result, equipment and medium
CN121120306A