Frequency domain semantic alignment enhanced large language model time sequence prediction method

By employing the FreqLLM method and utilizing frequency domain semantic alignment technology, the signal matching problem in the time and frequency domains of LLM is solved, improving the accuracy and generalization ability of time series prediction and achieving better global and local pattern capture.

CN120805914APending Publication Date: 2025-10-17CHONGQING BUSINESS VOCATIONAL COLLEGE +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510237227.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-01
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

Existing time series prediction methods based on large language models (LLMs) suffer from a mismatch between time-domain and frequency-domain signals, which prevents the model from effectively capturing global correlations and local patterns, and redundant information weakens the correlation of the prediction results.

Method used

A large language model time series prediction method (FreqLLM) with frequency domain semantic alignment enhancement is adopted. The time series is converted into frequency domain embedding through dual-scale frequency encoding, and the multi-head cross-attention mechanism is used to align with the pre-trained word embedding to generate semantic examples as soft cues, thereby optimizing the global and local pattern capture capabilities of LLM.

Benefits of technology

It significantly improves the accuracy and generalization ability of the model in time series forecasting. By encoding global and local frequency domain signals, it enhances the LLM's understanding and prediction performance of time series data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120805914A_ABST
    Figure CN120805914A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of time series prediction, and particularly discloses a frequency domain semantic alignment enhanced large language model time series prediction method. According to the method, frequency domain information is integrated to enhance semantic alignment of large language model time series prediction, and the method aims at providing a wider global view angle which is more naturally aligned with an LLM data processing method by utilizing the frequency domain information. Experiments and analysis on multiple reference data sets show that the method is superior to existing methods.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of time series prediction, and specifically discloses a frequency domain semantic alignment enhanced large language model time series prediction method. BACKGROUND

[0002] Time series forecasting (TSF) has wide applications in network systems, covering various fields such as point of interest (POI) recommendation, network economic models, and microservice log analysis. Traditionally, models for specific tasks are developed from scratch to improve prediction accuracy, and widely adopted methods include statistical methods (e.g., ARIMA, Xgboost), convolutional neural networks (e.g., TimesNet, MICN), Transformer-based models (e.g., PatchTST, Crossformer), and linear-based methods (e.g., DLinear, TiDE). However, these models require a large amount of domain data for training, which limits their applicability.

[0003] The advancements in pattern recognition capabilities of large language models (LLMs) have significantly improved their performance in the field of time series forecasting (TSF). However, by examining the relationship between the embeddings generated by LLMs when processing time series data and the prediction results, it can be found that these embeddings represent the LLM's reinterpretation of the data, and their correlation with the prediction results represents the model's ability to capture relevant patterns. When using time-domain signals as input, the embeddings show diagonal correlation across limited time steps. This indicates that the prediction is influenced by a small number of embeddings, revealing that the model cannot capture more extensive global correlations. The presence of redundant information was also found in the examination results, which weakens the embedding's correlation with the prediction results.

[0004] Although frequency domain signals have been proven to be effective in capturing global features of time series and have been successfully applied in various non-LLM-based TSF models, their role in LLM-based TSF methods remains underexplored. To address this gap, frequency domain signals were integrated into LLM-based TSF models in previous studies. Analysis shows that incorporating frequency domain signals into the prompt and sequence input of LLMs causes the embeddings to exhibit a curvilinear correlation with the prediction results. This correlation spans longer time steps, enabling the model to consider more extensive information at each prediction time step. Additionally, the global weight distribution is more uniform, and there is no redundant information, indicating that the embeddings provide a stronger global perspective. This highlights the significant advantage of utilizing frequency domain signals in the prompt and sequence input, improving the LLM's interpretation of time series data and prediction accuracy.

[0005] While incorporating frequency domain signals into the prompt and sequence data part significantly improves the model's ability to capture key patterns and global features, it also presents several challenges. One fundamental issue stems from the mismatch between the discrete embedding operations of LLMs and the continuous nature of time and frequency domain signals. This discrepancy complicates the way LLMs process and interpret these signals. Furthermore, the pre-trained knowledge and reasoning capabilities of LLMs are not naturally suited to capturing complex time and frequency domain patterns in time series data, making it a continuous challenge to achieve accurate and generalizable performance in TSF tasks. SUMMARY

[0006] To address the issues raised in the background, the present invention proposes a frequency domain semantic alignment enhanced large language model time series prediction method, which integrates frequency domain information based on a novel model framework (FreqLLM) to enhance the semantic alignment of large language model time series prediction. The framework aims to utilize frequency domain information and provide a broader global perspective that is more naturally aligned with the data processing methods of LLMs.

[0007] The frequency domain semantic alignment enhanced large language model time series prediction method in the present invention includes the following steps:

[0008] Step 1, based on the selected pre-trained backbone LLM, the FreqLLM model is constructed according to the following strategies:

[0009] For a given normalized time series input X, the time series input is converted into a frequency domain embedding f using a double-scale frequency encoding fre At the same time, the time series input is taken as the time domain embedding;

[0010] Based on the pre-trained word embeddings from the backbone LLM, the semantic examples E' are generated according to the principle of filtering word vectors unrelated to the time series analysis task and consolidating relevant word vectors;

[0011] Based on the similarity score between the semantic examples and the frequency domain embedding, the top K semantic examples that best represent the frequency information are selected as soft prompts and provided to the pre-trained semantic example LLM;

[0012] The frequency domain embedding and the time domain embedding are aligned with the semantic examples using the patching and multi-head cross-attention mechanisms, respectively. The aligned frequency domain embedding, time domain embedding, and soft prompts are input into the pre-trained backbone LLM to generate time series predictions for the next L time steps

[0013] Step 2, train the FreqLLM model;

[0014] Step 3, use the trained FreqLLM model to perform time series prediction tasks.

[0015] Further, the double-scale frequency encoding includes: on the one hand, performing fast Fourier transform on the normalized time sequence X, applying a linear layer to filter out useful frequency information, and changing the dimension of the output vector to obtain a global frequency domain signal

[0016] On the other hand, X is divided by a sliding window, and FFT is applied to the sequence in each small window, and then reorganized into a matrix Where w represents the number of sliding windows, and b represents the size of the sliding window after frequency extraction, and then The linear mapping result of the matrix is used as the query matrix Key matrix And value matrix Then use the linear mapping result of the matrix Finally, a linear layer is used to extract the most important and useful local frequency information, and the dimension of the output vector is changed, and the above process can be represented as follows:

[0017]

[0018] Where, Represents the local frequency domain signal;

[0019] f global and f local are spliced into a frequency domain embedding f fre ∈R 1×D .

[0020] Further, in step 1, a linear probe is used to generate semantic examples E' ∈ R V′×D from the pre-trained word embedding E ∈ R V×D of the autonomous stem model, where V represents the number of pre-trained word embeddings of the stem LLM model, V' represents the number of semantic examples, V' << V, and D represents the dimension of the pre-trained word embedding, and:

[0021] E' = Linear(E).

[0022] Further, the similarity score is calculated as follows:

[0023]

[0024] Where e' is the nth semantic example in E', n = {1, 2,..., V'}; n 1×D

[0025] ​​Further, the process of obtaining the soft prompt includes selecting the top K semantic examples with the highest similarity scores and concatenating the K semantic examples to form the final soft prompt Prompt∈R K×D , where:

[0026] Prompt=Concat(e′ [1] ,e′ [2] ,…,e′ [K] ),

[0027] where e′ [K] represents the semantic example with the k-th largest similarity score.

[0028] Further, the process of obtaining the aligned time domain embedding S time and the frequency domain embedding S fre includes:

[0029] The time domain embedding X needs to be divided into overlapping or non-overlapping time blocks, each time block has a length of L p , and the total number of input time blocks is where S represents the horizontal sliding span, and then a simple linear layer is used to map the dimension to d m to obtain the blocked time domain signal

[0030] Similarly, the same operation is performed on the frequency domain signal f fre ∈R 1×D to obtain the blocked frequency domain signal

[0031] Based on the multi-head attention mechanism, the query matrix Q key matrix K and value matrix V are defined for each head, where m={1,2,…,M} represents the m-th head, M represents the total number of heads of multi-head attention, and the subscript singal={time,fre}, where:

[0032]

[0033] where is the corresponding network parameter in the m-th head,

[0034] Then, the following reprogramming operation is performed:

[0035]

[0036] Then, each of all heads is aggregated to obtain​

[0037] Finally, a linear linear mapping adjusts the dimensions to obtain the aligned time domain embedding and the frequency domain embedding

[0038] Further, in step 2, the optimization objective loss function used in each training iteration is calculated as follows:

[0039]

[0040] where λ ≥ 0 is a trade-off hyperparameter, Y l 、 are the true and predicted results at the l-th time step, respectively.

[0041] Further, in each training iteration, only the network parameters in the model used to generate the semantic example E', used to perform the dual-scale frequency encoding, and used to align the frequency domain embedding and the time domain embedding with the semantic example, as well as the network parameters in the linear mapping part of the backbone LLM used to output the final result, are updated.

[0042] The principle and effect of the present application are that

[0043] The present application encodes global and local frequency domain signals using dual-scale frequency domain encoding, ensuring that the embedding captures comprehensive information. By guiding the encoding process with global frequency trends and local patterns closest to the predicted time point, the model ensures that it retains long-term and immediate context.

[0044] The present application selects prompts that align with the model's pre-trained semantic knowledge in order to enhance the LLM's contextual understanding of frequency domain information. By inductively combining the TSF semantics of a specific task with the LLM's pre-trained word embeddings, the present application generates semantic examples as prompts. These prompts are closely aligned with the frequency domain embedding, helping the model better understand global and local trends in the data.

[0045] By reprogramming the time and frequency domain signals into embeddings optimized for the LLM's semantic understanding, the present application bridges the gap between time and frequency domain data. By adjusting the numerical scale of the embedding sequence and the changes in the two domains, the present application ensures that the LLM can effectively process and reason at global and local scales.

[0046] In summary, the FreqLLM framework proposed by the present application introduces dual-scale frequency domain signal encoding and aligns it with pre-trained word embeddings. Together with the reprogramming of time and frequency domain signals into optimized representations, the present application significantly enhances the model's ability to capture global patterns and improve prediction accuracy. Experiments and analysis on multiple benchmark datasets show that FreqLLM outperforms existing methods. BRIEF DESCRIPTION OF DRAWINGS

[0047] Figure 1 The overall framework of the FreqLLM model in the embodiment of the application.

[0048] Figure 2 The workflow of the dual-scale frequency encoding module in the embodiment of the application.

[0049] Figure 3 The workflow of the dual-domain re-encoding module in the embodiment of the application. DETAILED DESCRIPTION

[0050] In this example, the TSF task is defined as follows: given a time series X ∈ R N×T , representing N different univariate at T time steps, the goal is to learn a prediction module F(·) that can understand the input time series and generate accurate predictions for the next L time steps, denoted as , representing N different univariate at L time steps. The ultimate goal is to minimize the mean squared error (MSE) between the predicted values and the actual values Y, which is formulated as:

[0051]

[0052] In this example, the multivariate time series is divided into N univariate time series, and each of the N univariate time series is processed independently. The i-th sequence is denoted as X i ∈ R 1×T , where each input channel X i is first normalized using reversible instance normalization (RevIN) to have zero mean and unit standard deviation to mitigate distribution shifts in time series data. For ease of reading, X is used to represent each channel in the following.

[0053] In this example, based on the selected pre-trained backbone LLM, the FreqLLM model is constructed according to the following strategies:

[0054] For a given normalized time series input X, the dual-scale frequency encoding is used to convert the time series input into a frequency domain embedding f fre , while taking the time series input as a time domain embedding;

[0055] Based on the pre-trained word embeddings from the backbone LLM, the semantic examples E' are generated according to the principle of filtering word vectors irrelevant to the time series analysis task and consolidating relevant word vectors;

[0056] Based on the similarity scores between the semantic examples and the frequency domain embeddings, the top K semantic examples that best represent the frequency information are selected as soft prompts and provided to the pre-trained semantic example LLM.

[0057] The frequency domain embedding and the time domain embedding are respectively aligned with the semantic examples by using the patching and multi-head cross attention mechanism, and the aligned frequency domain embedding, time domain embedding and soft prompt are input into the pre-trained backbone LLM to generate time series prediction of L subsequent time steps

[0058] Specifically as follows:

[0059] The time series input is converted from the time domain signal to the frequency domain embedding by using the dual-scale frequency encoding module, wherein the global and local frequency domain signals are used to guide the encoding process. In time series prediction, both the global and short-term signals of the sequence are important: the global signal reflects long-term trends or periodic information, while the short-term signal captures changes in a shorter time. In this example, both the global and local frequency domain signals are encoded simultaneously to ensure that the frequency information of the entire time window is extracted while the understanding of recent changes is strengthened. The detailed process of this part is shown in Figure 2 .

[0060] As shown in the left path of Figure 2 , for the global signal, the fast Fourier transform (FFT) is first applied to the normalized time series X to obtain the global frequency domain signal. However, not all frequency domain information is useful, and the noise in the time series data leads to a long-tailed frequency distribution in the frequency domain, so after performing FFT, a linear layer needs to be applied to filter out the useful frequency information and change the dimension of the output vector, which can be represented as follows:

[0061] f global = Linear(FFT(X)) (2)

[0062] wherein represents the global frequency domain signal, and D represents the dimension of the pre-trained word embedding of the backbone model.

[0063] In order to extract the local signal, the dual-scale frequency encoding module in this example also adopts the target attention mechanism. The attention mechanism can dynamically handle the relationship between different time steps and focus on important time steps, so it is widely used in time series prediction. Target attention is developed from the basic attention mechanism and is widely used in recommendation systems. Specifically, target attention emphasizes the relationship between specific time steps and global signals, enabling it to highlight the connection between local and global signals. In detail, as shown in the right path of the total in Figure 2 , X is divided by sliding window, and FFT is applied to the sequence in each small window, and then reorganized into a matrix wherein w represents the number of sliding windows, b represents the size of the sliding window after frequency extraction, and then The linear mapping results of the sequence corresponding to the last window are used as the query matrix Contains local information closest to the prediction point. Key matrix of target attention And value matrix Then use all windows, i.e. the linear mapping of matrix Finally, a linear layer is adopted to help the model extract the most important and useful local frequency information during training and change the dimension of the output vector, and the above process can be represented as follows:

[0064]

[0065] Where, Represents the local frequency domain signal.

[0066] After obtaining the global and local frequency domain signals, we concatenate the two to generate the final frequency domain embedding f fre ∈R 1×D :

[0067] f fre = Concat(f global ;f local ). (4)

[0068] Prompt is a simple and effective method to activate large language models (LLMs) to perform better in specific domain tasks. In the field of time series analysis, existing work mainly focuses on template-based and fixed prompts. However, this approach ignores the fact that time series representation inherently lacks human semantics, and rigidly combining fixed prompts with sequence information may prevent LLMs from effectively understanding both. Some work also utilizes soft prompts, where a randomly initialized, trainable vector specific to the task is generated as a prompt and improved during training to guide the LLM. However, these methods are still fixed on generating soft prompts from time domain signals, ignoring the inherent instability and redundant local fluctuations in time series data, which can significantly reduce model performance. Therefore, in the model of this example, we generate soft prompts from frequency domain signals to obtain a broader global perspective and mitigate the shortcomings associated with time domain signals.

[0069] As shown in Figure 1 , this example represents the pre-trained word embedding from the backbone model as E∈R V×D , V represents the number of pre-trained word embeddings of the backbone model, and D represents the dimension of the pre-trained word embedding, and a simple linear probe is used to generate a semantic example E′, where E′∈R V′×D, where V' << V. This linear probing aims to filter out word vectors irrelevant to the time series analysis task and consolidate relevant word vectors, thus reducing computational cost and focusing semantic information. Next, the frequency domain embedding f fre is computed is used as the basis for selecting examples as prompts. The similarity score is computed as follows:

[0070] E' = Linear(E), (5)

[0071]

[0072] where e' n ∈ R 1×D is the n-th word embedding of E', n = {1, 2, …, V'}.

[0073] Based on the similarity score, in this example, K semantic examples that best represent the frequency domain information are selected, and these K examples are concatenated to form the final prompt Prompt ∈ R K×D , which has:

[0074] Prompt = Concat(e' [1] , e' [2] , …, e' [K] ), (6)

[0075] where e' [k] represents the k-th largest word embedding with the largest similarity score.

[0076] In this example, the model uses a dual-domain re-encoding module to reprogram the time domain and frequency domain signals into semantic examples, and then aligns the sequence patterns with natural language patterns. This alignment enables large language models (LLMs) to provide accurate and generalizable performance in different areas of time series prediction. As mentioned earlier, incorporating frequency domain signals into prompts is essential for capturing global and local sequence dynamics, while time domain signals help LLMs understand the range and average level of sequences. Therefore, both time domain and frequency domain signals need to be input. However, since LLMs operate on discrete embeddings, while time domain and frequency domain signals are essentially continuous, there is a fundamental mismatch in data representation. To address this issue, a multi-head cross-attention mechanism is used in this example to reprogram the signals using semantic examples, converting them into a format that LLMs can better understand. The detailed process of this part is shown in Figure 3 .

[0077] First, the time domain signal X needs to be divided into overlapping or non-overlapping time blocks, each with a length of L p . Therefore, the total number of input time blocks is​ Where S represents the horizontal sliding span, and then a simple linear layer is used to map the dimension to d m , get the block time domain signal

[0078] Similarly, for the frequency domain signal f fre ∈R 1×D Perform the same operation to obtain the frequency domain signal of the block

[0079] Next, we need to reprogram the signal into the semantic example space. In the process of constructing the semantic example alignment prompt, we have obtained a series of semantic examples E′. These examples retain the word embeddings related to time series analysis while alleviating the problem of semantic information dispersion. Therefore, this example uses these semantic examples to reprogram the signal into a format that is easier for LLM to understand. Specifically, we use Figure 3 This is achieved by using the multi-head cross-attention network structure shown in Figure 3.

[0080] Specifically, the query matrix is ​​defined for each head Bond Matrix Sum Matrix Where m = {1, 2, ..., M} represents the mth head, M represents the total number of heads of multi-head attention, and subscript singal = {time, fre}, so:

[0081]

[0082] in D is the word embedding dimension of the backbone model,

[0083] Then, the reprogramming operation of the signal is defined as follows:

[0084]

[0085] By aggregating all the heads get Then perform linear mapping to align the hidden dimension with the backbone model to obtain the aligned time domain embedding and frequency domain embedding

[0086] Finally, the cues, frequency domain embeddings, and time domain embeddings are integrated and fed into a frozen large language model (LLM) to obtain the output, which is passed through the final linear mapping layer to produce the final prediction have:

[0087]

[0088] After building the above model, it also needs to be trained, in this case, the optimization objective loss function of each training iteration is :

[0089]

[0090] where λ ≥ 0 is a trade-off hyperparameter. The first term in the loss function represents the prediction loss of the model, focusing on improving the performance of the model. The second term is the similarity score loss, which calculates the average of the similarity scores between the selected K semantic examples and the frequency domain embedding. The purpose of adding this term is to ensure that the selected semantic examples are aligned with the frequency domain embedding, thereby optimizing the soft prompt.

[0091] As shown in the figure, in each training iteration, only the network parameters of the linear probe part, the dual-scale frequency encoding module and the dual-domain re-encoding module, and the network parameters in the linear mapping part of the main LLM used to output the final result are updated. The main network parameters of the LLM and the initial word embedding part remain frozen.

[0092] The following section verifies the effectiveness and adaptability of the proposed FreqLLM method through experiments

[0093] Experimental setup

[0094] Dataset: For long-term prediction experiments, multiple datasets were used for testing, including the power transformer temperature (ETT) dataset, and weather and traffic datasets, which are widely used to evaluate the long-term prediction performance of time series models. For short-term experiments, the M4 benchmark dataset was mainly used, which includes annual, quarterly, monthly, etc. Time series data of categories, with large-scale, wide coverage and high quality. For specific dataset information, please refer to Table 1, where the dimension represents the number of time series variables, and the dataset size represents the size of the training, validation and test sets.

[0095] Table 1: Dataset statistics table

[0096]

[0097] The baseline includes a group of Transformer-based methods: PatchTST[1], FEDformer[2], Autoformer[3] and Informer[4]; In addition, a group of non-Transformer-based methods are also compared: DLinear[5] and TimesNet[6], finally, two LLM-based methods, GPT4TS[7] and Time-LLM[8] are included.

[0098] Implementation details: To avoid software and hardware environment errors during model execution, this experiment uses Pycharm as the platform for Python code editing and management. All codes are written using Python 3.6, and all deep learning models are built using the Pytorch 1.11.0 framework. In addition, the hardware environment includes CPU (i9-14900k) and GPU (GTX 4090). The training uses the MSE loss, and the Adam optimizer is used with an initial learning rate of 10 -2 . GPT-2 is used as the backbone model, and the backbone model is kept at 32 layers. The patch dimension d m is set to 16, the number of heads M is set to 8, the size of the semantic examples V' is set to 1000, the loss weight λ is set to 0.08, the length of the prompt K is set to 8, the dimension size b of the sliding window is set to 8, and the length L p of each time block is set to 16, and the horizontal sliding span S is set to 8. The above is only an exemplary model and parameter setting of the present application, but is not limited thereto. Those skilled in the art can widely select suitable backbone LLM network and parameter combination adapted thereto according to their own task needs.

[0099] Experiment 1: Long-term prediction experiment (RQ1)

[0100] Experiment setup: For long-term prediction, the input time series length is set to 512, and the performance of four different prediction ranges is evaluated: {96, 192, 336, 720}. The evaluation indicators include the mean squared error (SSE) and the mean absolute error (MAE).

[0101] Experiment results: The experimental results are summarized in Table 2, where lower indicator values indicate better model performance. The best result is in bold, and the suboptimal result is underlined. In most cases, FreqLLM outperforms all baselines. Compared with LLM-related models, FreqLLM improves the average performance of models without fine-tuning the backbone (such as Time-LLM) by 2.21% and improves the average performance of models with fine-tuning the backbone (such as GPT 4TS) by 3.88%. Compared with the optimal Transformer-based model PatchTST, FreqLLM improves the average performance by 3.83%. Compared with other models, the average performance is improved by 25.42%. This is because FreqLLM utilizes both time-domain and frequency-domain data, allowing the model to retain the numerical scale of the data while improving its ability to capture a global view, and inducing semantic examples related to the time series prediction task from the pre-trained word embeddings of LLM, further enhancing the representation of time series.

[0102] Table 2: Long-term prediction results

[0103]

[0104] Experiment 2: Short-term forecasting experiment (RQ2)

[0105] Experiment setup: For the short-term forecasting experiment, the M4 benchmark test was used as the test dataset, and three datasets with too small dataset length (M4-Weekly, M4-Daily, M4-Hourly) were merged into M4-Other. The prediction length was set to 6 to 48, and the input length was set to twice the prediction length. The evaluation metrics included Symmetric Mean Absolute Percentage Error (SMAPE), Mean Absolute Scaled Error (MASE), and Overall Weighted Average (OWA).

[0106] Experiment results: Table 3 summarizes the short-term forecasting results, where lower metric values indicate better model performance. The optimal results are in bold, and the suboptimal results are underlined. The performance of FreqLLM is always better than all baselines, with an improvement of 3.15% and 9.39% compared to Time-LLM and GPT4TS, respectively. Even compared to the SOTA model PatchTST, FreqLLM is still competitive. This can be attributed to the model's use of local frequency domain signals closest to the prediction time period, which helps identify information most relevant to the prediction time step and mitigates the loss of correlation caused by too short sequence length.

[0107] Table 3: Short-term forecasting results.

[0108]

[0109]

[0110] The above-mentioned is only an embodiment of the present application, and the common knowledge of specific technical solutions and / or characteristics in the scheme is not described in detail. It should be noted that for those skilled in the art, without departing from the technical solutions of the present application, a number of modifications and improvements can be made, which should also be considered as the protection scope of the present application, which will not affect the effect and practicality of the patent. The protection scope of the present application should be subject to the content of its claims, and the specific embodiments in the description can be used to explain the content of the claims.

[0111] References

[0112] [1] Yuqi Nie, Nam H. Nguyen, Phanwadee Sinthong, and Jayant Kalagnanam. 2023. A Time Series is Worth 64 Words: Long-term Forecasting with Transformers. In The Eleventh International Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, May 1-5, 2023. OpenReview.net. https: / / doi.org / 10.48550 / arXiv.2211.14730.

[0113] [2] Tian Zhou, Ziqing Ma, Qingsong Wen, Xue Wang, Liang Sun, and Rong Jin. 2022. Fedformer: Frequency enhanced decomposed transformer for long-term series forecasting. In International conference on machine learning. PMLR, 27268-27286.

[0114] [3] Haixu Wu, Jiehui Xu, Jianmin Wang, and Mingsheng Long. 2021. Autoformer: Decomposition transformers with auto-correlation for long-term series forecasting. Advances in Neural Information Processing Systems 34 (2021), 22419-22430. https: / / doi.org / 10.48550 / arXiv.2106.13008.

[0115] [4] Haoyi Zhou, Shanghang Zhang, Jieqi Peng, Shuai Zhang, Jianxin Li, Hui Xiong, and Wancai Zhang. 2021. Informer: Beyond efficient transformer for long sequence time-series forecasting. In Proceedings of the AAAI conference on artificial intelligence, Vol. 35. 11106-11115.

[0116] [5] Ailing Zeng, Muxi Chen, Lei Zhang, and Qiang Xu. 2023. Are transformers effective for time series forecasting?. In Proceedings of the AAAI conference on artificial intelligence, Vol. 37. 11121-11128.

[0117] [6] Haixu Wu, Tengge Hu, Yong Liu, Hang Zhou, Jianmin Wang, and Mingsheng Long. 2023. TimesNet: Temporal 2D-Variation Modeling for General Time Series Analysis. In The Eleventh International Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, May 1-5, 2023. OpenReview.net. https: / / openreview.net / forum?id=ju_Uqw384Oq

[0118] [7] Tian Zhou, Peisong Niu, Liang Sun, Rong Jin, et al. 2023. One fits all: Power general time series analysis by pretrained lm. Advances in neural information processing systems 36 (2023), 43322-43355.

[0119] [8] Ming Jin, Shiyu Wang, Lintao Ma, Zhixuan Chu, James Zhang, Xiaoming Shi, Pin Yu Chen, Yuxuan Liang, Yuan-fang Li, Shirui Pan, et al. 2024. Time-LLM: Time Series Forecasting by Reprogramming Large Language Models. In International Conference on Learning Representations.

Claims

1. A large language model time series prediction method enhanced by frequency domain semantic alignment, characterized in that: The following steps are involved: Step 1: Based on the selected pre-trained backbone large language model LLM, build the FreqLLM model according to the following strategy: For a given, normalized time series input X, dual-scale frequency encoding is used to convert the time series input into a frequency domain embedding f fre , while using the time series input as time domain embedding; Based on the pre-trained word embeddings from the backbone LLM, semantic examples E′ are generated according to the principle of filtering out word vectors irrelevant to the time series analysis task and consolidating relevant word vectors; Based on the similarity score between the semantic examples and the frequency domain embedding, the top K semantic examples that best represent the frequency information are selected as soft hints and provided to the pre-trained semantic example LLM; The frequency domain embedding and time domain embedding are aligned with the semantic examples respectively using patching and multi-head cross attention mechanisms. The aligned frequency domain embedding, time domain embedding and soft hints are input into the pre-trained backbone LLM to generate time series predictions for the next L time steps. Step 2: Train the FreqLLM model; Step 3: Use the trained FreqLLM model to perform time series forecasting tasks.

2. The method according to claim 1, characterized in that The dual-scale frequency coding includes: on the one hand, performing a fast Fourier transform on the normalized time series X, applying a linear layer to filter out useful frequency information, and changing the dimension of the output vector to obtain a global frequency domain signal On the other hand, X is divided into two parts by sliding windows, and FFT is applied to the sequence in each small window, and then reorganized into a matrix Where w represents the number of sliding windows, b represents the size of the sliding window after frequency extraction, and Use linear mapping; based on the target attention mechanism, the linear mapping result corresponding to the sequence in the last window is selected and used as the query matrix Bond Matrix Sum Matrix Then use the matrix The linear mapping result; finally, a linear layer is used to extract the most important and useful local frequency information and change the dimension of the output vector. The above process can be expressed as follows: in, Represents the local frequency domain signal; f global and f local Splicing into frequency domain embedding f fre ∈R 1×D .

3. The method according to claim 2, characterized in that In step 1, a linear probe is used to generate semantic examples E′∈R from the pre-trained word embeddings E∈R of the backbone model. V×D Here, V represents the number of pre-trained word embeddings of the backbone LLM model, V′ represents the number of semantic examples, with V′ << V, and D represents the dimension of the pre-trained word embeddings. We have: V′×D where V represents the number of pre-trained word embeddings of the backbone LLM model, V′ represents the number of semantic examples, with V′ << V, and D represents the dimension of the pre-trained word embeddings, and we have: E′=Linear(E).

4. The method according to claim 3, characterized in that Similarity score The specific calculation process is as follows: where e′ n ∈R 1×D is the nth semantic example in E′, n = {1, 2, …, V′}.

5. The method according to claim 4, characterized in that The process of obtaining the soft prompt includes selecting the top K semantic examples with the highest similarity scores and connecting these K semantic examples to form the final soft prompt Prompt∈R K×D ,have: Prompt=Concat(e′ [1] ,And' [2] ,…,And' [K] ), where e′ [K] Represents similarity The k-th largest semantic example.

6. The method according to claim 5, characterized in that ,, get the aligned time domain embedding S time and frequency domain embedding S fre The process includes: The time domain embedding X needs to be split into overlapping or non-overlapping time blocks, each of which has a length of L p , the total number of input time blocks is Where S represents the horizontal sliding span, and then a simple linear layer is used to map the dimension to d m , get the block time domain signal Similarly, for the frequency domain signal f fre ∈R 1×D Perform the same operation to obtain the frequency domain signal of the block Based on the multi-head attention mechanism, the query matrix is ​​defined for each head Bond Matrix Sum Matrix Where m = {1, 2, ..., M} represents the mth head, M represents the total number of heads of multi-head attention, and subscript singal = {time, fre}, we have: in is the corresponding network parameter in the mth header, Then, perform the following reprogramming operations: Then, by aggregating all the heads get Finally, linear mapping is performed to adjust the dimension to obtain the aligned time domain embedding and frequency domain embedding 7. The method according to claim 4, characterized in that: In step 2, the optimization objective loss function used in each training iteration is The calculation is as follows: Where λ≥0 is a compromise hyperparameter, Y l 、 The true result and predicted result of the lth time step respectively.

8. The method according to claim 7, characterized in that: In each training iteration, only the network parameters in the model used to generate semantic examples E′, for dual-scale frequency encoding and for aligning frequency domain embedding and time domain embedding with semantic examples, as well as the network parameters in the linear mapping part of the backbone LLM used to output the final result are updated.

9. The method according to claim 1, characterized in that: Use reversible instance normalization to normalize the original time series to obtain the normalized time series X.