Power time series data prediction method and system based on collaboration of large and small models

Through the collaborative power time series data prediction method of large and small models, combined with long short-term memory networks and Transformer architecture, and using knowledge distillation and dynamic triggering strategies, the problems of insufficient efficiency and accuracy of traditional methods are solved, and efficient and accurate power time series data prediction is achieved. It is suitable for scenarios such as power grid load forecasting, weather forecasting, and equipment operation status monitoring.

CN120278556BActive Publication Date: 2025-09-19UESTC (SHENZHEN) ADVANCED RES INST +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510712832.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-30
Publication Date
2025-09-19
Estimated Expiration
2045-05-30

AI Technical Summary

Technical Problem

Traditional power time series data prediction methods have shortcomings in efficiency and accuracy. It is difficult to take into account both long-term global patterns and short-term local characteristics at the same time, and they require high computing resources, which affects the safe operation of the power grid and the efficiency of resource allocation.

Method used

A power time series data prediction method that combines large and small models is adopted. By constructing a small model of long short-term memory network and a large model of Transformer architecture, combined with knowledge distillation and dynamic triggering strategy, supervised learning and real-time prediction of the small model are realized, and global analysis and correction are performed using the large model.

Benefits of technology

It improves prediction efficiency and accuracy, takes into account both real-time and accuracy, and is suitable for a variety of power time series data prediction tasks, optimizing resource allocation and grid operation stability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120278556B_ABST
    Figure CN120278556B_ABST
Patent Text Reader

Abstract

An embodiment of the present invention discloses a method and system for predicting electric power time series data based on the collaboration of large and small models. The method includes: step S1: obtaining an electric power time series data set for training and preprocessing it to unify the data format; step S2: constructing a small time series model and a large time series model for pre-training; step S3: extracting features from the large time series model through knowledge distillation, and performing supervised learning on the small model based on the features; step S4: collecting electric power time series data in real time, and predicting it through the small time series model. If the prediction result deviates from the historical pattern or exceeds a preset threshold, the large time series model is called to perform a global analysis to obtain and output the corrected prediction result. The small model distillation of the present invention optimizes short-term predictions, and triggers collaborative correction of the large model when the prediction deviates from the historical pattern, ultimately generating an accurate time series prediction result that takes into account both real-time performance and global laws. The present invention addresses the shortcomings of existing models in terms of accuracy, coverage, and practicality.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of smart grid technology, and in particular to a method and system for predicting power time series data based on collaboration between large and small models. Background Art

[0002] Power time-series data forecasting is a core task in the power grid sector, widely used in key scenarios such as grid load forecasting, power dispatch optimization, and equipment operating status monitoring. Accurate power time-series forecasting not only improves the stability and reliability of grid operations but also provides important support for energy transformation and intelligent development, with significant economic and social benefits.

[0003] However, traditional power time-series data forecasting methods have significant shortcomings in terms of efficiency and accuracy. These methods typically rely on a single model for forecasting, making it difficult to simultaneously account for long-term global patterns (such as seasonal load variations) and short-term local characteristics (such as sudden load fluctuations). Furthermore, power time-series data is large in scale, high in dimensionality, and highly time-dependent, making processing this data demanding on computational resources. While lightweight models (such as traditional shallow neural networks or statistical models) offer certain advantages in computational efficiency, their ability to capture complex global dependencies is limited, resulting in forecast accuracy that fails to meet practical needs.

[0004] Especially in power systems, the accuracy of forecast results directly impacts the safe operation of the power grid and the efficiency of resource allocation. For example, errors in grid load forecasting can lead to irrational power generation plans, resulting in energy waste or power shortages. Inaccurate equipment status monitoring can delay fault warnings and increase operational risks. Therefore, improving forecast accuracy while ensuring real-time performance has become a key issue that needs to be addressed in power time-series data forecasting. Summary of the Invention

[0005] The technical problem to be solved by the embodiments of the present invention is to provide a method and system for predicting power time series data based on the collaboration of large and small models to improve prediction efficiency and accuracy.

[0006] To solve the above technical problems, an embodiment of the present invention proposes a method for predicting power time series data based on collaboration between large and small models, including:

[0007] Step S1: Acquire multi-source power time series data and perform unified preprocessing to construct a standardized power time series dataset for training;

[0008] Step S2: Build a small time series model using a long short-term memory network, build a large time series model using a timer architecture, and pre-train the small and large time series models using data from the power time series dataset;

[0009] Step S3: Migrating the global features of the large time series model to the small model through knowledge distillation, performing supervised learning on the small model, and adjusting the parameters of the small model;

[0010] Step S4: Collect power time series data in real time, and output short-term prediction results through the adjusted small model first. If the prediction results deviate from the historical pattern or exceed the preset threshold, trigger the time series large model to perform a global analysis, obtain and output the revised prediction results; if the prediction results do not deviate from the historical pattern or exceed the preset threshold, directly output the prediction results.

[0011] Accordingly, an embodiment of the present invention further provides a power time series data prediction system based on collaboration between large and small models, comprising:

[0012] Data preprocessing module: acquires multi-source power time series data and performs unified preprocessing to construct a standardized power time series dataset for training;

[0013] Model construction module: Build a small time series model using a long short-term memory network, build a large time series model using a timer architecture, and pre-train the small and large time series models using data from the power time series dataset;

[0014] Knowledge distillation module: transfers the global features of the large time series model to the small model through knowledge distillation, performs supervised learning on the small model, and adjusts the parameters of the small model;

[0015] Dynamic collaborative prediction module: collects power time series data in real time, and outputs short-term prediction results through the adjusted small model first. If the prediction result deviates from the historical pattern or exceeds the preset threshold, the time series large model is triggered to perform a global analysis to obtain and output the revised prediction result; if the prediction result does not deviate from the historical pattern or exceed the preset threshold, the prediction result is directly output.

[0016] The beneficial effects of the present invention are:

[0017] 1) This invention introduces an advanced model collaboration strategy to achieve an efficient combination of large and small models.

[0018] 2) This invention introduces the design and refined training of small models driven by large models, effectively improving the prediction accuracy of small models.

[0019] 3) This invention significantly improves the generalization ability of the model by combining self-supervised learning with transfer learning.

[0020] 4) The present invention achieves efficient model switching through a dynamic triggering strategy, taking into account both real-time performance and accuracy.

[0021] 5) The present invention is applicable to a variety of power time series data prediction tasks and has broad application prospects. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Figure 1 This is a system framework diagram of a power time series data prediction method based on collaboration of large and small models according to an embodiment of the present invention. DETAILED DESCRIPTION

[0023] It should be noted that, unless there is a conflict, the embodiments in this application and the features in the embodiments can be combined with each other. The present invention is further described in detail below with reference to the drawings and specific embodiments.

[0024] In the embodiments of the present invention, if there are directional indications (such as up, down, left, right, front, back, etc.), they are only used to explain the relative position relationship and movement status of the various components under a specific posture (as shown in the accompanying drawings). If the specific posture changes, the directional indication will also change accordingly.

[0025] In addition, the terms "first," "second," and so on, used in this disclosure are for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features being referred to. Therefore, features specified as "first" or "second" may explicitly or implicitly include at least one of these features.

[0026] Please refer to Figure 1 The power time series data prediction method based on collaboration of large and small models according to an embodiment of the present invention includes steps S1 to S4. Figure 1 This is a system framework diagram of an electric power time series data prediction method based on collaboration of large and small models according to an embodiment of the present invention, in which the data flow and module collaboration mechanism are demonstrated by taking power grid load prediction as an example.

[0027] Step S1: Obtain multi-source power time series data and perform unified preprocessing to construct a standardized power time series dataset for training. Training data is extracted from real-time power grid monitoring data, using a sliding window approach to create short-term time series samples. The dataset is divided into training, validation, and test sets.

[0028] The time series data includes time series data from the past few years, such as meteorology, power generation, power consumption, equipment operating status, voltage data, circuit data, etc. Time series data is data that is arranged in chronological order and relies on time dimension analysis, satisfying time continuity and dependency. Structured data (such as sensor readings with timestamps, historical grid load data, etc.), text data (such as weather forecast text "The temperature will rise to 35°C tomorrow, which may trigger a peak in electricity consumption") and image data (such as video frames or continuous remote sensing images) are preprocessed, including normalization, missing value filling and noise filtering. Then, the data representation of the three modalities is obtained (that is, structured features, text features, and image features are processed separately), the data format is unified (that is, unified multimodal input features are obtained based on structured features and text features), noise and dimensional differences are eliminated, and standardized input is provided for model training. Text data includes labels, content, etc., where labels provide general information about "what type" or "what topic" the text is about. For example, Figure 1 The label in is "grid load"; "content" provides detailed information on "what is the specific content" of the text, for example, Figure 1 The "content" in this section refers to the textual descriptions obtained from the image data using the AIGC model. The textual data source is raw, unstructured textual information collected before training, such as logs, reports, maintenance plans, and notices written by power dispatchers and operations personnel. These documents have titles and may be manually labeled (e.g., "load forecast related," "equipment failure," "maintenance plan," etc.). The textual descriptions obtained from the image data using the AIGC model are supplementary.

[0029] AIGC model parameter description: Figure 1The AIGC model shown in the figure is typically a pre-trained large-scale vision-language model. For example, advanced image captioning or visual question answering models such as the GPT series (e.g., GPT-4V), CLIP, BLIP, or Visual Transformer, which are capable of image understanding and text generation, can be used. Its internal parameters are trained on a large number of general image-text pair datasets (e.g., LAION-5B, CC3M, etc.), and it possesses powerful general image understanding and text generation capabilities. In the present invention, the AIGC model can be used directly as a pre-trained component (e.g., via an API call), in which case the model parameters are fixed. Alternatively, if sufficient image-text pairing datasets specific to the power industry are available, the open-source AIGC model can be fine-tuned to improve its semantic extraction accuracy and relevance for specific power industry images. The parameters of the AIGC model are designed to ensure accurate and efficient extraction of key textual semantic information from images for downstream prediction tasks.

[0030] The processing flow of time series data is divided into three main stages: preprocessing, model training, and collaborative prediction.

[0031] Specifically, step S1 includes the following steps:

[0032] Step S11: data preprocessing.

[0033] Detect missing values ​​in the data (such as null values ​​or outliers) and fill them using linear interpolation or time series interpolation. Detect and remove noise in the data (such as abnormal fluctuations or outliers) and use a sliding average method to smooth the data and reduce random fluctuations.

[0034] Step S12: Unify the data format.

[0035] Unify data from different sources to the same time frequency (e.g., hourly). Ensure that timestamps are aligned for ease of subsequent processing. Standardize all data by subtracting the mean and dividing by the variance to minimize the impact of different dimensions on model training. Split the power time series dataset into a training set (70%), a validation set (15%), and a test set (15%).

[0036] Step S2: Build a small time series model using a long short-term memory network and a large time series model using a timer architecture. Pre-train the small and large time series models using data from the power time series dataset. The large model primarily performs image-text structured feature comparison learning, image-text structured feature matching, and masked text structured feature prediction modeling.

[0037] For specific tasks, such as grid load forecasting: predicting the power load demand within a certain period of time in the future (the next 1 hour or 24 hours) to provide a basis for grid scheduling and energy allocation. Long short-term memory networks are used to train lightweight time series models to capture short-term time series dependencies (such as the grid load in the next 1 hour). Appropriate model parameters are selected for specific data and tasks, and the corresponding functions are called to fit the model. This model is then used to perform short-term real-time forecasting, that is, the time series model is responsible for processing short-term forecasting tasks in real time, and its forecast results are used as preliminary outputs. Details of time series model training: The input dimension is determined by the data characteristics (for example, the input dimension for temperature, humidity, and wind speed is 3); a sliding window is used to generate samples; the mean square error (MSE) loss function and the Adam optimizer are used to prevent overfitting through the validation set.

[0038] This paper selects the Long Short-Term Memory (LSTM) network as a lightweight time series model because it is good at capturing short-term time series dependencies. Specifically, the training process of the small model is as follows:

[0039] Step S21: Model parameter selection.

[0040] Specifically involved are the following:

[0041] 1. Input dimension: Determine the input dimension based on data characteristics For example, if the data contains temperature, humidity, and wind speed, the input dimensions are .

[0042] 2. Hidden layer dimension: Select the hidden layer dimension based on the complexity of the task ,For example or .

[0043] 3. Time step: Determine the time step according to task requirements Default selection The load for the next hour is predicted using hour as the time step.

[0044] 4. Output dimension: Determine the output dimension based on the prediction target For example, to predict the load for the next hour, the output dimension is .

[0045] 5. Hyperparameter Selection: Choosing an Appropriate Learning Rate (like ), batch size and the number of training rounds .

[0046] Step S22: Model training and saving.

[0047] First, a sliding window method is used to generate short-term time series samples. For example, the input is Hourly data , the output is Hourly data Then, the model is trained using the training set data and the model parameters are adjusted to minimize the loss function. The mean square error (MSE) is used as the loss function, and the formula is as follows:

[0048] ;

[0049] in, represents the number of samples; For the The actual output value of each sample (such as the actual value of the grid load); For the The predicted output value of samples is generated by the time series model. Continue to use the Adam optimizer for parameter optimization:

[0050] ;

[0051] in, Trainable time series model parameters, including weight matrix, bias terms, etc. Represents the current model iteration number, used to track the optimization process; is the learning rate, which controls the step size of parameter updates; It is the gradient of the loss function with respect to the parameters, indicating the direction of parameter update.

[0052] This method uses a validation set to monitor model performance and avoid overfitting. It also visualizes training and validation losses to ensure model convergence. Finally, it saves the trained model to a file for easy loading and deployment.

[0053] This paper uses the Timer model based on the Transformer architecture to perform large-scale pre-training on time series data from different sources. For the pre-training of large models, the Timer time series large model is selected as the basic architecture. Its main features are GPT-like time series pre-training and efficient attention mechanism. First, the time series data from different sources are processed in a unified format, and the specified length is converted to The time series data is converted into a time series token sequence. The Timer time series model based on Transformer is used for self-supervised pre-training. Time series token prediction Time sequence token. The specific network structure is:

[0054] 1) Decoder: An 8-layer decoder architecture. Each layer contains 8 self-attention heads and a decoder attention mechanism, enabling efficient processing of long time series. The self-attention heads are used to capture the temporal dependencies of the decoder input and output the prediction for the next time series token. The calculation formula for the self-attention mechanism is:

[0055] ;

[0056] in, are query, key, and value matrices respectively, is the dimension of the key vector, Represents the transpose operation of a matrix.

[0057] 2) Input sequence length: The time step is designed according to the specific task, and a multi-dimensional time series token vector is generated after passing through the tokenizer.

[0058] This paper selects Timer as the time series model architecture, and specifies the input sequence length as 1k. The self-supervised pre-training task of predicting the next token from a time sequence token.

[0059] Step S3: Through knowledge distillation, the global features of the large time series model are transferred to the small model, where supervised learning is performed on the small model, and the parameters of the small model are fine-tuned. After the small model completes self-supervised learning, transfer learning is performed to fine-tune the small model to the specific power grid time series data prediction task. This step uses labeled historical time series data for supervised learning to optimize the model's performance for the specific task.

[0060] This paper uses knowledge distillation to extract global features (such as long-term trends and cyclical patterns) from a large time series model, further supervising a lightweight small model. Specifically, using the large model's predictions or features as soft labels, combined with true label supervision, the small model is guided to align global and local features, thereby improving short-term prediction accuracy.

[0061] The proposed small time series model uses supervised learning to extract features from the large time series model through knowledge distillation, focusing on short-term time series forecasting tasks, such as hourly load forecasting. During the distillation process, the small model learns the global feature representation from the large model and simultaneously adjusts parameters through supervised learning to ensure high-precision forecasts for short time series.

[0062] As an implementation method, the small model adjustment process is as follows:

[0063] 1. Customize time series data from different sources and design and train specialized time series models.

[0064] 2. Use the trained large time series model for fine-tuning, extract features from the large time series model through knowledge distillation, and further supervise the lightweight small model to optimize its performance in short-term time series prediction tasks. for:

[0065] ;

[0066] in, Output for the small model, Output for large models, is the true label, is the distillation weight. represents the mean square error loss function, represents the cross entropy loss function.

[0067] Step S4: Collect power time series data in real time, and output short-term prediction results through the fine-tuned small model first. If the prediction result deviates from the historical pattern or exceeds the preset threshold, the large time series model is triggered to perform a global analysis to obtain and output the corrected prediction result; if the prediction result does not deviate from the historical pattern or exceeds the preset threshold, the prediction result is directly output. The present invention implements a collaborative time series processing mechanism of large and small models. The small time series model will process short-term prediction tasks in real time, preliminarily predict and output specific values ​​within the future short time window; the Transformer-based Timer time series large model that has undergone self-supervised pre-training and task fine-tuning learns long-term rules, and uses the self-attention mechanism to capture global time series patterns. It triggers a global correction when it detects that the local prediction of the small model deviates from the historical pattern or exceeds the threshold. That is, the large model is responsible for correcting the task, and when the specified step length is reached, the large model is responsible for correcting the task. After that, through the time series similarity analysis, if the similarity detection result deviates from the similarity threshold set by all corresponding domain time series patterns The system calls the time series large model to output the revised long-term forecast results. The collaborative mechanism of the present invention realizes efficient model switching through a dynamic triggering strategy and outputs accurate load forecast results.

[0068] First, generate a batch of specified lengths based on the time series large model The time series token sequence pair, for example, the current prediction result sequence is , the historical sequence is ,but This constitutes a time series token sequence pair. Subsequently, the BERT-based bidirectional Transformer encoder, through self-supervised pre-training and supervised fine-tuning of the BERT pre-training model, takes the time series token sequence as input and calculates the cosine similarity of the two time series vectors to determine whether the similarity deviates from the set threshold (for example, the similarity threshold is set to 0.8; if the cosine similarity between the current predicted sequence and the historical pattern is <0.8, it is determined to deviate from the historical pattern). If the similarity deviates from the threshold, it is determined to deviate from the historical pattern, which will trigger the time series large model (Timer) to perform global analysis and correction. The cosine similarity calculation formula is as follows:

[0069] ;

[0070] in, and are vector representations of two time series. The loss function uses mean square error:

[0071] ;

[0072] in, is the true similarity score; Represents the total number of time series pairs used to calculate cosine similarity; It is used to identify the sequence pair currently being processed. ; Each time series pair contains two time series, Representative A vector representation of a time series.

[0073] These symbols are used together to calculate the similarity of time series pairs and evaluate the prediction accuracy of the model through a loss function. The loss function is used to measure the difference between the predicted similarity and the true similarity, thereby supporting the subsequent collaborative mechanism (dynamic triggering strategy).

[0074] The present invention first processes the time series data from different sources in a unified format, and uses the Transformer-based Timer time series large model for pre-training to learn the global time series pattern. At the same time, the training of the corresponding professional time series small model is designed and completed. In order to optimize the effect of the time series large model in the power field tasks, the labeled power time series data is further used for supervised fine-tuning to optimize the performance of the time series large model in power time series related tasks. In order to further fine-tune the accuracy of the small model in short-term time series prediction, the small model will perform supervised learning by migrating the features extracted from the time series large model through knowledge distillation. Finally, the time series similarity calculation model is used to verify whether the prediction of the small model deviates. When the deviation exceeds the threshold, the time series large model will be used for correction. The present invention solves the technical problems in the prior art that the small model has insufficient accuracy in long-term time series prediction and the large model has low prediction efficiency.

[0075] This invention effectively improves the global pattern capture capability and short-term prediction accuracy of time series data through pre-training of large time series models, fine-tuning of lightweight small models, and a dynamic collaborative mechanism. This framework can be applied to a variety of time series data prediction scenarios, including power grid load forecasting, weather forecasting, and equipment operating status monitoring. By effectively utilizing this time series data, prediction accuracy can be significantly improved, resource allocation can be optimized, and the system's real-time responsiveness can be enhanced.

[0076] The power time series data prediction system based on large and small model collaboration of the present invention includes:

[0077] Data preprocessing module: acquires multi-source power time series data and performs unified preprocessing to construct a standardized power time series dataset for training;

[0078] Model construction module: Build a small time series model using a long short-term memory network, build a large time series model using a timer architecture, and pre-train the small and large time series models using data from the power time series dataset;

[0079] Knowledge distillation module: transfers the global features of the large time series model to the small model through knowledge distillation, performs supervised learning on the small model, and fine-tunes the parameters of the small model;

[0080] Dynamic collaborative prediction module: collects power time series data in real time, and preferentially outputs short-term prediction results through a fine-tuned small model. If the prediction result deviates from the historical pattern or exceeds the preset threshold, the large time series model is triggered to perform a global analysis to obtain and output the revised prediction result; if the prediction result does not deviate from the historical pattern or exceed the preset threshold, the prediction result is directly output.

[0081] As an implementation method, the knowledge distillation module adopts the loss function of knowledge distillation for:

[0082] ;

[0083] in, is the output of the small time series model, Output of the time series large model, is the true label, is the distillation weight, represents the mean square error loss function, represents the cross entropy loss function.

[0084] As an implementation method, the time series large model adopts an 8-layer decoder architecture, each layer contains 8 self-attention heads and a decoder attention mechanism. The calculation formula of the self-attention mechanism is:

[0085] ;

[0086] in, are query, key, and value matrices respectively, is the dimension of the key vector, Represents the transpose operation of a matrix.

[0087] As an implementation method, the dynamic collaborative prediction module generates a time series token sequence pair of a preset length based on the prediction results and the historical pattern, calculates the cosine similarity of the two time series sequence vectors of the sequence pair, and determines whether the similarity deviates from the preset similarity threshold. If the similarity deviates from the threshold, it is determined to deviate from the historical pattern. The cosine similarity calculation formula is as follows:

[0088] ;

[0089] in, and are the vector representations of the two time series of the sequence pair, and the loss function uses the mean square error:

[0090] ;

[0091] in, is the true similarity score; Represents the total number of time series pairs used to calculate cosine similarity; It is used to identify the sequence pair currently being processed. ; Each time series pair contains two time series, Representative A vector representation of a time series.

[0092] These symbols are used together to calculate the similarity of time series pairs and evaluate the model's prediction accuracy through a loss function. The loss function is used to measure the difference between the predicted similarity and the true similarity, thereby supporting the subsequent collaborative mechanism (dynamic triggering strategy).

[0093] As an implementation method, the time series small model uses mean square error as the loss function, and the formula is as follows:

[0094] ;

[0095] in, represents the number of samples; For the The actual output value of each sample (such as the actual value of the grid load); For the The predicted output value of each sample is generated by the time series model.

[0096] The Adam optimizer is used to optimize the parameters of the time series model:

[0097] ;

[0098] in, Trainable time series model parameters, including weight matrix, bias terms, etc. Represents the current model iteration number, used to track the optimization process; is the learning rate, which controls the step size of parameter updates; It is the gradient of the loss function with respect to the parameters, indicating the direction of parameter update.

[0099] The present invention, through the collaborative application of large and small models, fully leverages the high-precision global pattern learning capabilities of the large model and the high-efficiency real-time prediction advantages of the small model, combining knowledge distillation with dynamic triggering strategies to achieve a balance between efficiency and accuracy. Specifically, data preprocessing unifies multi-source time series features, large models are pre-trained to learn long-term patterns, small models are distilled to optimize short-term predictions, and large model corrections are triggered when small model predictions deviate from historical patterns, ultimately generating accurate time series prediction results that take into account both real-time and global patterns. This addresses the shortcomings of existing models in terms of accuracy, coverage, and practicality.

[0100] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.

Claims

1. A method for predicting power time series data based on collaboration between large and small models, characterized in that: include: Step S1: Acquire multi-source power time series data and perform unified preprocessing to construct a standardized power time series dataset for training; Step S2: Build a small time series model using a long short-term memory network, build a large time series model using a timer architecture, and pre-train the small and large time series models using data from the power time series dataset; Step S3: Migrating the global features of the large time series model to the small model through knowledge distillation, performing supervised learning on the small model, and adjusting the parameters of the small model; Step S4: Real-time power time series data is collected, and short-term prediction results are outputted preferentially through the adjusted small model. If the prediction results deviate from the historical pattern or exceed the preset threshold, the time series large model is triggered to perform a global analysis to obtain and output the revised prediction results. If the prediction results do not deviate from the historical pattern or exceed the preset threshold, the prediction results are directly outputted. In step S1, the structured data, text data, and image data in the time series data are preprocessed to obtain data representations of the three modalities. The data format is then unified to eliminate noise and dimensional differences, providing standardized input for model training. The image data is supplemented with text descriptions extracted by the AIGC model. In step S4, a time series token sequence pair of a preset length is generated based on the prediction result and the historical pattern, and the cosine similarity of the two time series sequence vectors of the sequence pair is calculated to determine whether the similarity deviates from the preset similarity threshold. If the similarity deviates from the threshold, it is determined to deviate from the historical pattern. The cosine similarity calculation formula is as follows: ; in, and are the vector representations of the two time series of the sequence pair, and the mean square error is used as the loss function. The calculation formula is as follows: ; in, is the true similarity score; Represents the total number of time series pairs used to calculate cosine similarity; It is used to identify the sequence pair currently being processed. ; Each time series pair contains two time series, Representative A vector representation of a time series.

2. The power time series data prediction method based on large and small model collaboration according to claim 1, characterized in that: In step S3, the loss function of knowledge distillation is for: ; in, is the output of the small time series model, Output of the time series large model, is the true label, is the distillation weight, represents the mean square error loss function, represents the cross entropy loss function.

3. The power time series data prediction method based on large and small model collaboration according to claim 1, characterized in that: The time series large model adopts an 8-layer decoder architecture. Each layer contains 8 self-attention heads and a decoder attention mechanism. The calculation formula of the self-attention mechanism is: ; in, are query, key, and value matrices respectively, is the dimension of the key vector, Represents the transpose operation of a matrix.

4. The power time series data prediction method based on large and small model collaboration according to claim 1, characterized in that: In step S2, the time series model uses mean square error as the loss function, and the formula is as follows: ; in, Represents the loss function value, which is used to measure the difference between the predicted value and the true value; represents the number of samples; For the The true output value of the sample; For the The predicted output value of each sample is generated by the time series model; The Adam optimizer is used to optimize the parameters of the time series model: ; in, are the parameters of the trainable time series model, Represents the number of iterations of the current model, is the learning rate, is the gradient of the loss function with respect to the parameters.

5. A power time series data prediction system based on collaboration between large and small models, characterized in that: include: Data preprocessing module: acquires multi-source power time series data and performs unified preprocessing to construct a standardized power time series dataset for training; Model construction module: Build a small time series model using a long short-term memory network, build a large time series model using a timer architecture, and pre-train the small and large time series models using data from the power time series dataset; Knowledge distillation module: transfers the global features of the large time series model to the small model through knowledge distillation to perform supervised learning of the small model and adjust the parameters of the small model; Dynamic collaborative forecasting module: This module collects power time series data in real time and prioritizes outputting short-term forecast results through an adjusted small model. If the forecast results deviate from historical patterns or exceed a preset threshold, the large time series model is triggered to perform a global analysis to obtain and output a revised forecast result. If the forecast results do not deviate from historical patterns or exceed a preset threshold, the forecast results are directly output. The data preprocessing module preprocesses the structured data, text data, and image data in the time series data to obtain data representations of the three modalities. It then unifies the data format, eliminates noise and dimensional differences, and provides standardized input for model training. The image data is supplemented with text descriptions extracted by the AIGC model. The dynamic collaborative prediction module generates a time series token sequence pair of preset length based on the prediction results and historical patterns, calculates the cosine similarity of the two time series vectors of the sequence pair, and determines whether the similarity deviates from the preset similarity threshold. If the similarity deviates from the threshold, it is determined to deviate from the historical pattern. The cosine similarity calculation formula is as follows: ; in, and are the vector representations of the two time series of the sequence pair, and the mean square error is used as the loss function. The calculation formula is as follows: ; in, is the true similarity score; Represents the total number of time series pairs used to calculate cosine similarity; It is used to identify the sequence pair currently being processed. ; Each time series pair contains two time series, Representative A vector representation of a time series.

6. The power time series data prediction system based on large and small model collaboration according to claim 5, characterized in that: The loss function of knowledge distillation used by the knowledge distillation module for: ; in, is the output of the small time series model, Output of the time series large model, is the true label, is the distillation weight, represents the mean square error loss function, represents the cross entropy loss function.

7. The power time series data prediction system based on large and small model collaboration according to claim 6, characterized in that: The time series large model adopts an 8-layer decoder architecture. Each layer contains 8 self-attention heads and a decoder attention mechanism. The calculation formula of the self-attention mechanism is: ; in, are query, key, and value matrices respectively, is the dimension of the key vector, Represents the transpose operation of a matrix.

8. The power time series data prediction system based on large and small model collaboration according to claim 6, characterized in that: The time series model uses mean square error as the loss function, and the formula is as follows: ; in, Represents the loss function value, which is used to measure the difference between the predicted value and the true value; is the sample size; For the The true output value of the sample; For the The predicted output value of each sample is generated by the time series model; The Adam optimizer is used to optimize the parameters of the time series model: ; in, are the parameters of the small time series model, is the gradient of the loss function with respect to the parameters, Represents the number of iterations of the current model, is the learning rate.

Citation Information

Patent Citations

  • Unmanned aerial vehicle image cloud edge collaborative identification method based on knowledge distillation

    CN116563731A

  • Load prediction method and system based on variational auto-encoder and TIME-LLM model

    CN119994901A