Multi-scale time sequence power consumption prediction method for deep learning application

Through the multi-scale time series method and attention mechanism, the accuracy and calculation overhead problems of power consumption prediction in the deep learning task of intelligent computing center are solved, and efficient power consumption prediction is achieved.

CN120276933APending Publication Date: 2025-07-08GUILIN UNIVERSITY OF TECHNOLOGY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510285255.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-11
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

The existing server power consumption prediction methods are not high in prediction accuracy and have high computational overhead in deep learning tasks in intelligent computing centers, especially in the fact that they fail to effectively consider the power consumption characteristics of multiple scales.

Method used

By using the multi-scale time series method, by acquiring server performance and power consumption data, after data preprocessing, the multi-scale time series extraction layer is used to extract time series of different scales, decompose and remix long-term and short-term components, and combine linear layers and attention mechanisms to predict, and finally obtain the final power consumption prediction result.

Benefits of technology

It improves the power consumption prediction accuracy of deep learning tasks in smart computing centers, and at the same time reduces the computing overhead and improves the model's expression and generalization capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120276933A_ABST
    Figure CN120276933A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of machine learning, in particular to a deep learning application-oriented multi-scale time sequence power consumption prediction method, which comprises the following steps of: firstly, acquiring performance data and power consumption data during server operation through a representative deep learning task; the method comprises the following steps of: firstly, extracting time sequence data of different scales from power consumption data through a multi-scale time sequence extraction layer, and remixing long-term components and short-term components of time sequences of different scales by using a multi-power consumption mode decomposition and remixing layer, so as to improve the prediction precision of the model on the data; and finally, a final prediction result is obtained by using multi-scale time sequence prediction and a mixed layer. Compared with most existing power consumption prediction model methods based on common high-performance computing center tasks, the power consumption prediction precision of the deep learning tasks of the intelligent computing center can be effectively improved, and the computing overhead is low.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of machine learning, and specifically relates to a multi-scale time series power consumption prediction method for deep learning applications. Background Art

[0002] Due to the rapid development of artificial intelligence technology, general high-performance computing centers are increasingly unable to meet the growing artificial intelligence computing power requirements, so intelligent computing centers have been established. The intelligent computing center uses its highly optimized hardware architecture, distributed computing capabilities, software ecosystem support, and targeted service models to meet the computing power requirements of artificial intelligence technology. However, the high computing power of the intelligent computing center also brings higher power consumption overhead. The core hardware of the intelligent computing center, such as GPUs, TPUs, or AI-specific chips, has a single-card power consumption of up to 300-700 watts, which is much higher than that of CPUs used in traditional high-performance computing centers. Therefore, how to reduce the high energy consumption of the intelligent computing center is an urgent problem to be solved.

[0003] Power consumption prediction of servers is the basis for reducing energy consumption. Existing server power consumption prediction methods can be divided into two categories. One is the power consumption prediction method based on non-neural networks, which uses statistical methods to model power consumption and has the advantage of fast operation, but the disadvantage is low accuracy. The other is the power consumption prediction method based on neural networks. Compared with the non-neural network power consumption prediction method, it can better adapt to the complex non-linear relationship of server power consumption prediction, but the disadvantage is high computational overhead, and most of the existing neural network power consumption prediction methods are for traditional high-performance computing center tasks and do not consider the multi-scale power consumption characteristics of deep learning tasks in intelligent computing centers, so that their effects in deep learning task power consumption prediction are not good. Summary of the Invention

[0004] The purpose of the present invention is to provide a multi-scale time series power consumption prediction method for deep learning applications, which uses the multi-scale time series method to improve the prediction accuracy of deep learning tasks of servers in intelligent computing centers and reduce the computational overhead while ensuring the prediction accuracy.

[0005] To achieve the above purpose, the present invention provides a multi-scale time series power consumption prediction method for deep learning applications, including the following steps:

[0006] Step 1: Obtain server performance data and power consumption data during the execution of deep learning;

[0007] Step 2: Merge the server performance data and power consumption data and perform data preprocessing;

[0008] Step 3: Obtain multi-scale power consumption time series of different scales from the preprocessed data through a multi-scale time series extraction layer;

[0009] Step 4: Extract the long-term component and the short-term component, and remix them separately;

[0010] Step 5: Re-fuse the mixed long-term component and short-term component into a multi-scale time series;

[0011] Step 6: Use a linear layer to predict the future power consumption for each multi-scale time series;

[0012] Step 7: Use the attention mechanism to mix time series of different scales to obtain the final prediction result.

[0013] Optionally, during the process of data preprocessing in Step 2, use a preprocessing layer for preprocessing. The preprocessing layer is embedded in the neural network model and learnable parameters are added. The expression is as follows:

[0014]

[0015] x” = γ ⊙ x′ + β

[0016] where μ and σ are the mean and standard deviation of the input data respectively, which can be calculated along the feature dimension or the batch dimension. γ represents the scaling factor, β represents the offset term, which consists of neural network parameters and are the parameters to be optimized during the model training process. They will be updated with each iteration so that the network can improve its expressive ability and generalization ability while maintaining stability. ⊙ represents element-wise multiplication.

[0017] Optionally, during the execution of Step 3, first select 1, 2, 4, 6 as the window size of AvgPooling or Conv1D, then process the power consumption data respectively with the Padding operation, and finally obtain four time series with different sampling granularities. The expression is as follows:

[0018] A m = AvgPooling(Padding(x m ))

[0019] e m = Embedding(A m )

[0020] where, x m represents the data of the m-th scale of the input, A m represents the extracted multi-scale time series, Padding(·) represents the operation of padding the input data, AvgPooling(·) represents the average pooling layer, and e m represents the multi-scale power consumption time series after passing through the embedding layer.

[0021] Optionally, the execution process of Step 4 includes the following steps:

[0022] Step 4.1: Decompose the long-term components in the data of each scale through the AvgPooling layer or the Conv1D layer. The expression is as follows:

[0023] c m = AvgPooling(Padding(x m ))

[0024] where x m represents the data of the m-th scale of the input, and c m represents the long-term component extracted from the data of the m-th scale. Padding(·) represents the operation of padding the input data, and AvgPooling represents the average pooling layer;

[0025] Step 4.2: Calculate the short-term components in the data of each scale according to the obtained long-term components and the input data. The expression is as follows:

[0026] f m = x m - c m

[0027] where f m represents the short-term component extracted from the data of the m-th scale;

[0028] Step 4.3: Remix the extracted long-term components. The expression is as follows:

[0029] c m = W x (c m+1 ) + b x + c m

[0030] where c m represents the long-term component extracted from the m-th component, c m+1 represents the long-term component extracted from the (m + 1)-th component, W x represents the weight matrix, and b x represents the bias term;

[0031] Step 4.4: Remix the extracted short-term components to make their features easier to be discovered. The expression is as follows:

[0032] f m = Linear(f m+1 ) + f m

[0033] where f m represents the short-term component extracted from the m-th component, f m+1Denote the short-term component extracted from the (m + 1)-th component as the short-term component, and Linear(·) represents the linear layer. In the short-term component mixing when m equals 1, the KAN linear layer is used instead of the fully connected layer. Its expression is as follows:

[0034] x′ = x.reshape(H * W, A)

[0035]

[0036] x out = x B .reshape(H, W, B)

[0037] where x represents the short-term component with the finest-grained scale of one, H, W, and A respectively represent the sizes of the three dimensions of x, x′ is the data after dimension compression, α q and β q,p represent one-dimensional functions, i represents a sample point in x′, p represents the feature index in this sample point, q represents the index under the summation symbol, which is used to traverse a series of one-dimensional functions α q and β q,p , x.reshape(*) represents the dimension transformation operation on x, B represents the size of the last dimension of the output after passing through the KAN layer, and x out represents the output data after dimension reduction.

[0038] Optionally, in step 5, the mixed long-term component and short-term component are re-fused into a multi-scale time series, and the expression is as follows:

[0039] P m = W1 * c m + W2 * f m + W3 * e m

[0040] where, P m refers to the re-fused multi-scale time series, W1, W2, and W3 are learnable parameter matrices, c m represents the long-term component extracted from the data of the m-th scale, and f m represents the short-term component extracted from the data of the m-th scale.

[0041] Optionally, in step 6, the linear layer is used to predict the future power consumption respectively using the mixed data of each scale, and the expression is as follows:

[0042] O m = Predictor m (x m )

[0043] where, The future power consumption predicted using the data of the m-th scale, Predictor m (·) represents the predictor for the data of the m-th scale. In particular, for the scale with m equal to 1, the predictor uses a KAN linear layer instead of a fully connected layer to predict the future result.

[0044] Optionally, in the finest-grained mixing of the short-term component and in the process of predicting the future results of each scale, KAN is used instead of a fully connected layer. The expression is as follows:

[0045]

[0046] where f(x) is the objective function to be represented, and x1,...x n represent the n input variables of the function, and θ q,p represents a unary function with respect to a single input variable, and θ q represents an external function that takes the sum of the unary functions of the previous layer as input and outputs the final result.

[0047] Optionally, in step 7, an attention mechanism is used to mix time series of different scales to obtain the final prediction result. The expression is as follows:

[0048] W q = W x2 (tanh(W x1 x + b))

[0049] p out = W q * O m

[0050] where x is the prediction result of different scales, W q is the attention weight, b is the bias term, tanh is the activation function, and p out is the final prediction result, and W x2 and W x1 are the network weight parameters used to obtain the attention weight and can be automatically adjusted through backpropagation, and b is the bias term.

[0051] The present invention provides a multi-scale time series power consumption prediction method for deep learning applications. First, representative deep learning tasks are used to obtain the performance data and power consumption data during the operation of the server. Then, the collected data is preprocessed to improve the prediction accuracy of the model for the data. Next, different-scale time series data is extracted from the power consumption data through a multi-scale time series extraction layer. Then, a variety of power consumption mode decomposition and remixing layers are used to remix the long-term and short-term components of each scale time series, so that the model can more easily learn the features in the data. Finally, a multi-scale time series prediction and mixing layer is used to obtain the final prediction result. Compared with most existing power consumption prediction model methods based on ordinary high-performance computing center tasks, the present invention can effectively improve the power consumption prediction accuracy of deep learning tasks in intelligent computing centers and has a lower computational overhead. Description of the Drawings

[0052] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0053] Figure 1 It is a schematic flow chart of the specific steps of a multi-scale time series power consumption prediction method for deep learning applications of the present invention.

[0054] Figure 2 It is a comparison chart of the mean absolute error between the predicted value and the true value of the server power consumption predicted by other power consumption prediction algorithms for the next 96 seconds on the power consumption datasets of four deep learning tasks in a specific embodiment of the present invention.

[0055] Figure 3 It is a comparison chart of the mean relative error between the predicted value and the true value of the server power consumption predicted by other power consumption prediction algorithms for the next 96 seconds on the power consumption datasets of four deep learning tasks in a specific embodiment of the present invention.

[0056] Figure 4 It is a comparison chart between the predicted result and the true power consumption after predicting the power consumption for the next 96 seconds in the image classification task of the present invention. Detailed Description of the Embodiments

[0057] The following will describe in detail the embodiments of the present invention. The examples of the embodiments are shown in the drawings, where the same or similar reference numerals represent the same or similar elements or elements with the same or similar functions from beginning to end. The embodiments described below with reference to the drawings are exemplary and are intended to explain the present invention and should not be construed as a limitation of the present invention.

[0058] The present invention provides a multi-scale time series power consumption prediction method for deep learning applications, including the following steps:

[0059] Step 1: Obtain server performance data and power consumption data during the execution of deep learning;

[0060] Step 2: Merge the server performance data and power consumption data and perform data preprocessing;

[0061] Step 3: Pass the preprocessed data through a multi-scale time series extraction layer to obtain multi-scale power consumption time series of different scales;

[0062] Step 4: Extract the long-term component and the short-term component, and remix them respectively;

[0063] Step 5: Re-fuse the mixed long-term component and short-term component into a multi-scale time series;

[0064] Step 6: Use a linear layer to predict the future power consumption for each multi-scale time series respectively;

[0065] Step 7: Use an attention mechanism to mix time series of different scales to obtain the final prediction result.

[0066] The following is a further description in combination with specific implementation steps, as Figure 2 shown:

[0067] Step 1: Obtain server performance data and power consumption data during the execution of deep learning;

[0068] In Step 1, use the NVIDIA interface and Python library to obtain server performance data, and use an intelligent electricity meter to obtain server power consumption data.

[0069] Step 2: Merge the server performance data and power consumption data and perform data preprocessing;

[0070] Step 2 includes the following two processes:

[0071] Step 2.1 Align the acquired data at the time level:

[0072] Align and merge the performance data and power consumption data in seconds in terms of time

[0073] Step 2.2 Preprocess the merged data through a preprocessing layer:

[0074] The preprocessing layer is embedded in the neural network model and learnable parameters are added, and the expression is as follows:

[0075]

[0076] x” = γ⊙x′ + β

[0077] Where μ and σ are the mean and standard deviation of the input data respectively, which can be calculated along the feature dimension or batch dimension. γ represents the scaling factor, β represents the offset term, which is composed of neural network parameters and can be learned, and ⊙ represents element-wise multiplication.

[0078] Step 3: Pass the preprocessed data through the multi-scale time series extraction layer to obtain multi-scale power consumption time series at different scales;

[0079] Specifically, during the execution of Step 3, first select 1, 2, 4, 6 as the window sizes of AvgPooling or Conv1D, then process the power consumption data respectively with the Padding operation, and finally obtain four time series with different sampling granularities. The expressions are as follows:

[0080] A m = AvgPooling(Padding(x m ))

[0081] e m = Embedding(A m )

[0082] Where, x m represents the data of the m-th scale of the input, A m represents the extracted multi-scale time series, Padding(·) represents the operation of padding the input data, AvgPooling(·) represents the average pooling layer, and e m represents the multi-scale power consumption time series after passing through the embedding layer.

[0083] Step 4: Extract the long-term component and the short-term component, and remix them respectively;

[0084] The execution process of Step 4 includes the following steps:

[0085] Step 4.1: Decompose the long-term component in the data of each scale through the AvgPooling layer or Conv1D layer. The expression is as follows:

[0086] c m = AvgPooling(Padding(x m ))

[0087] Where, x m represents the data of the m-th scale of the input, c m represents the long-term component extracted from the data of the m-th scale, Padding(·) represents the operation of padding the input data, and AvgPooling(·) represents the average pooling layer;

[0088] Step 4.2: Calculate the short-term components in each scale of data based on the obtained long-term components and the input data. The expression is as follows:

[0089] f m = x m - c m

[0090] where f m represents the short-term component extracted from the m-th scale of data;

[0091] Step 4.3: Re-mix the extracted long-term components. The expression is as follows:

[0092] c m = Linear(c m+1 ) + c m

[0093] where c m represents the long-term component extracted from the m-th component, c m+1 represents the long-term component extracted from the (m + 1)-th component, and Linear(·) represents a linear layer;

[0094] Step 4.4: Re-mix the extracted short-term components to make their features easier to discover. The expression is as follows:

[0095] f m = Linear(f m+1 ) + f m

[0096] where f m represents the short-term component extracted from the m-th component, f m+1 represents the short-term component extracted from the (m + 1)-th component, Linear(·) represents a linear layer. In the mixing of the short-term component when m equals 1, the KAN linear layer is used instead of the fully connected layer. Its expression is as follows:

[0097] x′ = x.reshape(H * W, A)

[0098]

[0099] x out = x B .reshape(H, W, B)

[0100] where x represents the short-term component with the finest granularity of scale one, H, W, and A respectively represent the sizes of the three dimensions of x, x′ is the data after dimensional compression, α q and β q,prepresents a one-dimensional function, i represents a sample point in x′, p represents the feature index in this sample point, and q represents the index under the summation symbol, which is used to traverse a series of one-dimensional functions α q and β q,p , x.reshape(*) represents performing a dimensional transformation operation on x, B represents the size of the last dimension of the output after passing through the KAN layer, and x out represents the output data after dimensional reduction.

[0101] Step 5: Re-fuse the mixed long-term component and short-term component into a multi-scale time series;

[0102] In Step 5, the mixed long-term component and short-term component are re-fused into a multi-scale time series, and the expression is as follows:

[0103] P m = W1 * c m + W2 * f m + W3 * e m

[0104] where, P m refers to the re-fused multi-scale time series, W1, W2, and W3 are learnable parameter matrices, and c m represents the long-term component extracted from the data of the m-th scale, and f m represents the short-term component extracted from the data of the m-th scale.

[0105] Step 6: Use a linear layer to predict the future power consumption for each multi-scale time series respectively;

[0106] In Step 6, the linear layer is used to predict the future power consumption using the mixed data of each scale respectively, and the expression is as follows:

[0107] O m = Predictor m (x m )

[0108] where, represents the future power consumption predicted using the data of the m-th scale, and Predictor m (·) represents the predictor for the data of the m-th scale. In particular, the predictor for the scale where m equals 1 uses the KAN linear layer to replace the fully connected layer to predict the future result.

[0109] Optionally, in the mixing of the finest granularity of the short-term component and the prediction process of the future results of each scale, use KAN to replace the fully connected layer, and the expression is as follows:

[0110]

[0111] Among them, f(x) is the target function to be represented, and x1,... x n represent n input variables of the function, and θ q,p represents a unary function with respect to a single input variable, and θ q represents an external function that takes the sum of the previous-layer unary functions as input and outputs the final result.

[0112] Step 7: Use the attention mechanism to mix time series of different scales to obtain the final prediction result.

[0113] In Step 7, the attention mechanism is used to mix time series of different scales to obtain the final prediction result, and the expression is as follows:

[0114] W q = W x2 (tanh(W x1 x + b))

[0115] p out = W q * O m

[0116] Among them, x is the prediction result of different scales, W q is the attention weight, b is the bias term, tanh is the activation function, and p out is the final prediction result. W x2 and W x1 are network weight parameters used to obtain the attention weight and can be automatically adjusted through backpropagation, and b is the bias term.

[0117] Furthermore, please refer to Figures 2 to 4 , and the present invention also proposes specific embodiments. By comparing with other baseline methods, the usage results are used for auxiliary explanation:

[0118] Figure 2 is a comparison chart of the mean absolute error between the predicted value and the true value of the server power consumption predicted by other latest high-performance computing center power consumption prediction algorithms and the present invention on a server equipped with an RTX3080Ti graphics card for the next 96 seconds on the power consumption datasets of four deep learning tasks (time series prediction task, image classification task, natural language processing task, anomaly detection task). Among them, BP_PM represents a model composed of a four-layer BP neural network, D long short-term memory network_PM represents a two-layer long short-term memory network model, ITF_PM represents an iTransformer model, and MLR_PM represents a multiple linear regression model.

[0119] Figure 3It is a comparison chart of the average relative error between the predicted value and the true value of the server power consumption predicted by other latest high-performance computing center power consumption prediction algorithms and the server power consumption in the next 96 seconds on a server equipped with an RTX3080Ti graphics card in the power consumption datasets of four deep learning tasks (time series prediction task, image classification task, natural language processing task, anomaly detection task). Among them, BP_PM represents the model composed of a four-layer BP neural network, D long short-term memory network_PM represents a two-layer long short-term memory network model, ITF_PM represents the iTransformer model, and MLR_PM represents the multiple linear regression model.

[0120] Figure 4 It is a comparison chart of the prediction result and the true power consumption after the power consumption prediction in the next 96 seconds in the image classification task of the present invention. Among them, the light-colored line strip represents the predicted value, and the dark-colored line represents the true value.

[0121] Table 1 is a comparison of the time-consuming results of a multi-scale time series power consumption prediction method for deep learning applications of the present invention and four other intelligent computing center power consumption prediction algorithms in the time series prediction type deep learning task. Among them, BP_PM represents the model composed of a four-layer BP neural network, D long short-term memory network_PM represents a two-layer long short-term memory network model, ITF_PM represents the iTransformer model, and MLR_PM represents the multiple linear regression model.

[0122] Table 1 Comparison table of the duration results between the present invention and the remaining baseline methods

[0123]

[0124] Therefore, it can be seen that the method of the present invention does not cause a large amount of computing power consumption while improving the prediction accuracy under different datasets, and its consumption time is within an acceptable range.

[0125] In summary, the difference between the present invention and the existing research work is that it takes into account the multi-scale power consumption characteristics of deep learning tasks, integrates the standardization processing step into the neural network model as a layer of the model, applies the KAN network to the prediction module to enhance the expression ability of the model, and finally uses a fully connected layer to fuse the prediction results of all scales. The present invention proposes a multi-scale time series power consumption prediction method for deep learning applications. First, power consumption data and performance data of server deep learning tasks are obtained through intelligent electricity meters and program scripts. The collected data is directly input into the model after being processed. Through the normalization layer of the model, the input data is normalized, and scaling parameters are added to improve the prediction ability of the model. Then, multi-scale time series are extracted from the power consumption data through the multi-scale time series extraction layer, projected using the Embedding layer, and the short-term and long-term components of the power consumption are extracted and remixed through multiple power consumption mode decomposition and remixing layers. Specifically, the KAN layer is used to remix the short-term components with the smallest scale to improve the prediction accuracy of the model. Finally, the final result is obtained through the mixing of the multi-scale time series prediction and mixing layer through the attention mechanism.

[0126] The above-disclosed are only one or more preferred embodiments of the present invention. Of course, the scope of the rights of the present invention cannot be limited thereby. Those of ordinary skill in the art can understand all or part of the processes of implementing the above embodiments, and the equivalent changes made according to the claims of the present invention still fall within the scope covered by the invention.

Claims

1. A multi-scale time series power consumption prediction method for deep learning applications, characterized in that, It includes the following steps: Step 1: Obtain server performance data and power consumption data during the execution of deep learning; Step 2: Merge the server performance data and power consumption data and perform data preprocessing; Step 3: Pass the preprocessed data through a multi-scale time series extraction layer to obtain multi-scale power consumption time series of different scales; Step 4: Extract the long-term component and the short-term component, and remix them separately; Step 5: Re-fuse the mixed long-term component and short-term component into a multi-scale time series; Step 6: Use a linear layer to predict the future power consumption for each multi-scale time series respectively; Step 7: Use an attention mechanism to mix time series of different scales to obtain the final prediction result.

2. The multi-scale time series power consumption prediction method for deep learning applications as described in claim 1, wherein, During the data preprocessing in Step 2, a preprocessing layer is used for preprocessing. The preprocessing layer is embedded in the neural network model and learnable parameters are added. The expression is as follows: x”=γ⊙x′+β where μ and σ are the mean and standard deviation of the input data respectively, which can be calculated along the feature dimension or the batch dimension. γ represents the scaling factor, β represents the offset term, which is composed of neural network parameters, is the parameter to be optimized during the model training process, and will be updated with each iteration. ⊙ represents element-wise multiplication.

3. The multi-scale time series power consumption prediction method for deep learning applications as described in claim 2, wherein, During the execution of Step 3, first select 1, 2, 4, 6 as the window size of AvgPooling or Conv1D, then perform Padding operations on the power consumption data respectively, and finally obtain four time series with different sampling granularities. The expression is as follows: A m = AvgPooling(Padding(x m )) e m = Embedding(A m ) Among them, x m represents the data of the m-th scale of the input, A m represents the extracted multi-scale time series, Padding(·) represents the operation of padding the input data, Padding(·) represents the average pooling layer, e m represents the multi-scale power consumption time series after passing through the embedding layer.

4. The multi-scale time series power consumption prediction method for deep learning applications as described in claim 3, wherein, The execution process of Step 4 includes the following steps: Step 4.1: Decompose the long-term component in each scale data through the AvgPooling layer or the Conv1D layer. The expression is as follows: c m = AvgPooling(Padding(x m )) where x m represents the data of the m-th scale of the input, and c m represents the long-term component extracted from the data of the m-th scale. Padding(·) represents the operation of padding the input data, and Padding(·) represents the average pooling layer; Step 4.2: Calculate the short-term component in each scale data according to the obtained long-term component and the input data. The expression is as follows: f m = x m - c m where f m represents the short-term component extracted from the m-th scale data; Step 4.3: Remix the extracted long-term component. The expression is as follows: c m = W x (c m+1 ) + b x + c m Among them, c m represents the long-term component extracted from the m-th component, and c m+1 represents the long-term component extracted from the (m + 1)-th component, W x represents the weight matrix, and b x represents the bias term; Step 4.4: Remix the extracted short-term component to make its features more easily discovered. The expression is as follows: f m = Linear(f m+1 ) + f m Among them, f m represents the short-term component extracted from the m-th component, and f m+1 represents the short-term component extracted from the (m + 1)-th component. Linear(·) represents a linear layer. When m is not 1, Linear(·) represents a fully connected layer. Only in the short-term component mixing where m equals 1, the KAN linear layer is used instead of the fully connected layer. The expression is as follows: x′=x.reshape(H*W,A) x out = x B .reshape(H,W,B) where x represents the short-term component with the finest granularity of scale one, H, W, and A represent the sizes of the three dimensions of x respectively, x' is the data after dimensionality compression, α q and β q,p represent one-dimensional functions, i represents a sample point in x', p represents the feature index in this sample point, q represents the index under the summation symbol, used to traverse a series of one-dimensional functions α q and β q,p , x.reshape(*) represents the dimensionality transformation operation on x, B represents the size of the last dimension of the output after passing through the KAN layer, x out represents the output data after dimensionality restoration.

5. The multi-scale time series power consumption prediction method for deep learning applications as described in claim 4, wherein, In Step 5, the mixed long-term component and short-term component are re-fused into a multi-scale time series. The expression is as follows: P m = W1 * c m + W2 * f m + W3 * e m Among them, P m refers to the re - fused multi - scale time series. W1, W2, and W3 refer to learnable parameter matrices, and c m represents the long - term component extracted from the data of the m - th scale, and f m represents the short - term component extracted from the data of the m - th scale.

6. The multi-scale time series power consumption prediction method for deep learning applications as described in claim 5, wherein, In Step 6, a linear layer is used to predict the future power consumption respectively using the mixed data of each scale. The expression is as follows: O m = Predictor m (x m ) Among them, O m represents the future power consumption predicted using the data of the m-th scale, and Predictor m (·) represents the predictor for the data of the m-th scale. For the predictor of the scale where m equals 1, the KAN linear layer is used instead of the fully connected layer to predict the future result.

7. The multi-scale time series power consumption prediction method for deep learning applications as described in claim 6, wherein, In the process of the finest-grained mixing of short-term components and the prediction of future results at each scale, KAN is used instead of the fully connected layer, and the expression is as follows: Among them, f(x) is the target function to be represented, and x1,...x n represents n input variables of the function, and θ q,p represents a unary function with respect to a single input variable, and θ q represents an external function that takes the sum of the upper-layer unary functions as input and outputs the final result.

8. The multi-scale time series power consumption prediction method for deep learning applications according to claim 7, characterized in that In step 7, the attention mechanism is used to mix time series at different scales to obtain the final prediction result, and the expression is as follows: W q = W x2 (tanh(W x1 x + b)) p out = W q * O m Among them, x is the prediction result at different scales, W q is the attention weight, b is the bias term, tanh is the activation function, p out is the final prediction result, O m represents the future power consumption predicted using the data of the m-th scale, W x2 and W x1 are the network weight parameters, which can be automatically adjusted by backpropagation, and b is the bias term.