Time series data prediction method and device, electronic equipment, storage medium and program product
By introducing large language models to predict time series data, the problem of the reduction in accuracy of traditional models when processing different distributions or novel pattern data is solved, and higher prediction accuracy and resource efficiency are achieved, and the model deployment and use process is simplified.
Patent Information
- Application Number
- CN202510178273.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-18
- Publication Date
- 2025-05-13
AI Technical Summary
The traditional time series data prediction model has reduced prediction accuracy when processing data with different distributions or novel patterns, and requires complex feature engineering and model parameter adjustment, which increases application difficulty and occupies a high level of system resources, especially in resource-constrained environments.
The semantic understanding ability of the large language model is used to predict the time series data, and the current continuous time series data is obtained for quantization and discretization, the large language model is used for prediction, and the prediction results are reverse discretized, which simplifies the model deployment and use process.
It improves the accuracy and generalization ability of time series data prediction, reduces system resource usage, simplifies model deployment and use, reduces technical barriers, and facilitates the maintenance and application of large language models.
Smart Images

Figure CN119990460A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a time series data prediction method, device, electronic device, storage medium and program product. Background Art
[0002] Traditional models often have deficiencies in the process of predicting time series data. For example, they often overfit specific types of data, resulting in a decrease in the model's prediction accuracy when encountering data with different distributions or novel patterns; for another example, they require complex feature engineering and model parameter adjustment processes, requiring experts to design models, perform feature engineering, and adjust parameters, which increases the difficulty of practical applications and limits the use of models by non-professionals; for another example, they require a large amount of computing resources, and their application in resource-constrained environments needs to be improved urgently. Summary of the invention
[0003] The present invention provides a time series data prediction method, device, electronic device, storage medium and program product, which utilize the semantic understanding ability of a large language model to improve the accuracy of time series data prediction, and also avoid excessive occupation of system resources by training a large number of models for a large amount of indicator data in the time series data, reduce the occupation of system resources, simplify the model deployment process, reduce the technical barriers to model use, facilitate the maintenance of the large language model, and improve the applicability and generalization of time series data prediction.
[0004] According to one aspect of the present invention, a time series data prediction method is provided, the method comprising:
[0005] Get the current continuous time series data for the current time period;
[0006] quantizing the current continuous time series data to obtain current discrete time series data of the current time period;
[0007] Using a large language model, predicting the current discrete time series data to obtain target discrete time series data for a target time period;
[0008] The target discrete time series data is de-discretized to obtain the target continuous time series data of the target time period.
[0009] According to another aspect of the present invention, there is provided a time series data prediction device, the device comprising:
[0010] The current continuous time series data acquisition module is used to acquire the current continuous time series data of the current time period;
[0011] A current continuous time series data quantization module is used to quantize the current continuous time series data to obtain the current discrete time series data of the current time period;
[0012] A target discrete time series data prediction module is used to use a large language model to predict the current discrete time series data to obtain target discrete time series data for a target time period;
[0013] The target discrete time series data dequantization module is used to de-discretize the target discrete time series data to obtain the target continuous time series data of the target time period.
[0014] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising:
[0015] at least one processor; and
[0016] a memory communicatively connected to the at least one processor; wherein,
[0017] The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the time series data prediction method described in any embodiment of the present invention.
[0018] According to another aspect of the present invention, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the time series data prediction method described in any embodiment of the present invention when executed.
[0019] According to another aspect of the present invention, a computer program product is provided. The computer program product includes a computer program. When the computer program is executed by a processor, the time series data prediction method described in any embodiment of the present invention is implemented.
[0020] The technical solution of the embodiment of the present invention obtains the current continuous time series data of the current time period, quantizes the current continuous time series data, and obtains the current discrete time series data of the current time period, thereby realizing the discretization of the current continuous time series, which is convenient for the large language model to predict the time series data. The large language model is adopted to predict the current discrete time series data to obtain the target discrete time series data of the target time period, and the target discrete time series data is de-discretized to obtain the target continuous time series data of the target time period. The large language model is introduced, and the semantic understanding ability of the large language model is utilized to improve the accuracy of time series data prediction. Moreover, the excessive occupation of system resources by training a large number of models for a large number of indicator data in the time series data is avoided, the occupation of system resources is reduced, the model deployment process is simplified, the technical barriers to the use of the model are reduced, the maintenance of the large language model is facilitated, and the applicability and generalization of time series data prediction are improved.
[0021] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present invention, nor are they intended to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0023] Figure 1 is a flowchart of a time series data prediction method provided according to Embodiment 1 of the present invention;
[0024] Figure 2 is a flowchart of a time series data prediction method provided according to Embodiment 2 of the present invention;
[0025] Figure 3 is a schematic diagram of the structure of a time series data prediction device provided according to Embodiment 3 of the present invention;
[0026] Figure 4 It is a schematic diagram of the structure of an electronic device for implementing the time series data prediction method of an embodiment of the present invention. DETAILED DESCRIPTION
[0027] In order to enable those skilled in the art to better understand the scheme of the present invention, the technical scheme in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present invention.
[0028] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0029] Embodiment 1
[0030] Figure 1 A flowchart of a time series data prediction method provided in Embodiment 1 of the present invention. The embodiment of the present invention is applicable to the case of predicting time series data, and the method can be performed by a time series data prediction device, which can be implemented in the form of hardware and / or software, and can be configured in an electronic device that carries the time series data prediction function, such as a client or a server.
[0031] See also Figure 1 The time series data forecasting method shown includes:
[0032] S110, obtaining current continuous time series data of the current time period.
[0033] The current time period may be the time period in which the current continuous time series data is located. The current continuous time series data may be the continuous time series data in the business system detected within the current time period. The current continuous time series data may be data with complex nonlinear relationships. It can be understood that the current continuous time series data has certain irregularities in its changes over time. Optionally, the current continuous time series data may include indicator data of multiple dimensions. Exemplarily, the current continuous time series data may include A indicator, A indicator compared to the beginning of the year, A indicator compared to the previous month, A indicator annual average and A indicator monthly average, etc.
[0034] In an optional embodiment of the present invention, the current continuous time series data may be the current market analysis continuous time series data. The current market analysis continuous time series data may be used to analyze and predict the market status of a certain business. Exemplarily, the current market analysis continuous time series data may include the amount of resource transfer in the business system, the amount of resource transfer compared to the beginning of the year, the amount of resource transfer compared to the previous month, the annual average of the amount of resource transfer, and the monthly average of the amount of resource transfer. By concretizing the current continuous time series data as the current market analysis continuous time series data, the applicability of time series data prediction can be improved.
[0035] Specifically, with authorization, an API (Application Programming Interface) may be used to obtain current continuous time series data for the current time period.
[0036] S120, quantizing the current continuous time series data to obtain the current discrete time series data of the current time period.
[0037] The current discrete time series data may be a quantized result of the current continuous time series data. In comparison, both the current discrete time series data and the current continuous time series data correspond to the current time period, but the continuity of the data is different.
[0038] Specifically, a quantization function may be used to quantize the current continuous time series data to obtain the current discrete time series data of the current time period.
[0039] For example, a series of quantile bin boundaries can be defined, and a unique corresponding discrete value can be determined based on the range of each group of quantile bin boundaries. Each value in the current continuous time series data can be quantified, and the discrete value corresponding to the continuous value of each data point in the current continuous time series can be determined, and each discrete value can be sorted in time order to obtain the current discrete time series data of the current time period.
[0040] S130 . Use a large language model to predict the current discrete time series data to obtain target discrete time series data for a target time period.
[0041] Large Language Models (LLMs) can be used to predict time series data. In comparison, traditional models (such as deep learning models) require experts to design, feature engineer, and adjust parameters, which increases the difficulty of practical application of the model and limits the use of the model by non-professionals. In addition, traditional models usually predict indicators of a single dimension. When predicting a large amount of time series data, a large number of models need to be trained and deployed, which has a high overall occupancy rate of hardware resources, limiting its application in resource-constrained environments. Large language models have high semantic understanding capabilities, and can use a single large language model to predict indicators of multiple dimensions without deploying a large number of models, which can reduce the ineffective occupation of system resources. At the same time, the technical difficulty of the training and use process of large language models is relatively low, which can improve the applicability and generalization of the model. In addition, it can also improve the accuracy of time series data prediction.
[0042] The target time period may be a predicted time period corresponding to the current time period. Optionally, the time span of the target time period may be the same as or different from the time span of the current time period. The target discrete time series data may be predicted data of the target time period obtained by the large language model based on the current discrete time series data.
[0043] Specifically, the current discrete time series data and the corresponding prompt words may be input into a large language model, and the large language model may be used to perform semantic understanding on the current discrete time series data to directly obtain the target discrete time series data of the target time period.
[0044] In an optional embodiment of the present invention, a large language model is used to predict the current discrete time series data to obtain target discrete time series data for a target time period, including: using a large language model to obtain a time span of the target time period; using a large language model to predict the current discrete time series data to obtain candidate discrete time series data for a candidate time period; using a large language model to detect whether the time span of the candidate time period reaches the time span of the target time period; using a large language model, when the time span of the candidate time period does not reach the time span of the target time period, updating the candidate time period to the current time period, updating the candidate discrete time series data to the current discrete time series data, and returning to execute the step of predicting the current discrete time series data to obtain the candidate discrete time series data for the candidate time period, until the time span of the candidate time period reaches the time span of the target time period, combining the candidate discrete time series data of each candidate time period in chronological order to obtain the target discrete time series data for the target time period.
[0045] The time span of the target time period may be the difference between the end time and the start time of the target time period. The time span of the current time period may be the difference between the end time and the start time of the current time period. Optionally, the time span of the target time period may be greater than or equal to the time span of the current time period. The time span of the alternative time period may be the time span obtained by training the large language model. Optionally, the time span of the alternative time period may be less than or equal to the time span of the target time period. When the time span of the alternative time period is less than the time span of the target time period, the target discrete time series data of the target time period may be obtained by recursive prediction of the large language model.
[0046] Specifically, a large language model can be used to obtain the time span of the target time period based on the corresponding prompt word. The current discrete time series data and the corresponding prompt word can be input into the large language model, and the large language model can be used to perform semantic understanding on the current discrete time series data to obtain the candidate discrete time series data of the candidate time period. The large language model can be used to compare the time span of the candidate time period with the time span of the target time period. When the time span of the candidate time period does not reach the time span of the target time period, the candidate time period is updated to the current time period, the candidate discrete time series data is updated to the current discrete time series data, and the step of predicting the current discrete time series data to obtain the candidate discrete time series data of the candidate time period is returned to execute, until the time span of the candidate time period reaches the time span of the target time period, the candidate discrete time series data of each candidate time period are combined in time order to obtain the target discrete time series data of the target time period.
[0047] This scheme realizes the recursive prediction of the large language model by introducing the comparison process of the time span of the target time period and the time span of the alternative time period, further improving the flexibility and applicability of time series data prediction.
[0048] S140 , de-discretizing the target discrete time series data to obtain target continuous time series data in a target time period.
[0049] The target continuous time series data may be a result of mapping the target discrete time series data in the target time period to a continuous numerical range.
[0050] Specifically, the target discrete time series data is de-discretized by using the de-quantization function corresponding to the quantization function to obtain the target continuous time series data in the target time period.
[0051] Exemplarily, a series of predefined quantile bin boundaries can be obtained, and a unique corresponding range midpoint value can be determined based on the quantile bin boundary range of each group. The corresponding range midpoint value can be determined according to the quantile bin boundary range corresponding to each value in the target discrete time series data, and the range midpoint value is used as the corresponding continuous value in the target continuous time series data, thereby obtaining the target continuous time series data for the target time period.
[0052] In an optional embodiment of the present invention, the current continuous time series data is quantized to obtain the current discrete time series data of the current time period, including: normalizing the current continuous time series data to obtain the current standard time series data of the current time period; quantizing the current standard time series data to obtain the current discrete time series data of the current time period; correspondingly, the target discrete time series data is de-discretized to obtain the target continuous time series data of the target time period, including: de-discretizing the target discrete time series data to obtain the target standard time series data of the target time period; de-normalizing the target standard time series data to obtain the target continuous time series data of the target time period.
[0053] The current normative time series data may be the normalized result of the current continuous time series data of the current time period. The target normative time series data may be the de-discretized result of the target discrete time series data. The target continuous time series data may be the de-normalized result of the target normative time series data. In comparison, the data magnitudes of the indicator data of different dimensions in the target normative time series data are uniform, while the data magnitudes of the indicator data of different dimensions in the target continuous time series data are the actual data magnitudes corresponding to the indicator data.
[0054] Specifically, a normalization function can be used to normalize the continuous values of each data point in the current continuous time series data to obtain each standard value in the current standard time series data of the current time period. A quantization function can be used to quantize each standard value in the current standard time series data to obtain each discrete value in the current discrete time series data of the current time period. Correspondingly, an inverse quantization function corresponding to the quantization function can be used to inversely discretize each discrete value in the target discrete time series data to obtain each standard value in the target standard time series data of the target time period. An inverse normalization function corresponding to the normalization function can be used to inversely normalize each standard value in the target standard time series data to obtain the target continuous time series data of the target time period.
[0055] This solution introduces a normalization process for the current continuous time series data and a corresponding denormalization process, which reduces the impact of indicator data of different magnitudes in the time series data on the prediction accuracy of the time series data, and further improves the accuracy of the time series data prediction.
[0056] The technical solution of the embodiment of the present invention obtains the current continuous time series data of the current time period, quantizes the current continuous time series data, and obtains the current discrete time series data of the current time period, thereby realizing the discretization of the current continuous time series, which is convenient for the large language model to predict the time series data. The large language model is adopted to predict the current discrete time series data to obtain the target discrete time series data of the target time period, and the target discrete time series data is de-discretized to obtain the target continuous time series data of the target time period. The large language model is introduced, and the semantic understanding ability of the large language model is utilized to improve the accuracy of time series data prediction. Moreover, the excessive occupation of system resources by training a large number of models for a large number of indicator data in the time series data is avoided, the occupation of system resources is reduced, the model deployment process is simplified, the technical barriers to the use of the model are reduced, the maintenance of the large language model is facilitated, and the applicability and generalization of time series data prediction are improved.
[0057] Embodiment 2
[0058] Figure 2 A flowchart of a time series data prediction method provided for the second embodiment of the present invention. Based on the above embodiments, the embodiment of the present invention further adds "obtaining historical continuous time series data; quantizing the historical continuous time series data to obtain historical discrete time series data; time slicing the historical discrete time series data to obtain historical discrete time series data for each historical time period; for a single historical time period, using a large language model to predict the historical discrete time series data for the historical time period to obtain predicted discrete time series data for the predicted time period; for a single historical time period, adjusting the parameters of the large language model according to the difference between the predicted discrete time series data for the predicted time period and the historical discrete time series data for the predicted time period", introducing a pre-training process of the large language model, which can improve the prediction efficiency of time series data. It should be noted that for the parts not described in detail in the embodiment of the present invention, please refer to the description of other embodiments.
[0059] See also Figure 2 The time series data forecasting method shown includes:
[0060] S210, obtaining historical continuous time series data.
[0061] The historical continuous time series data may be continuous time series data that existed in the business system in the past. The historical continuous time series data may be used to pre-train a large language model.
[0062] Specifically, with authorization, you can use the API to obtain historical continuous time series data.
[0063] S220. Quantify the historical continuous time series data to obtain historical discrete time series data.
[0064] Historical discrete time series data can be the quantified result of historical continuous time series data. In comparison, both historical discrete time series data and historical continuous time series data correspond to the current time period, but the continuity of the data is different.
[0065] Specifically, a quantization function may be used to quantize historical continuous time series data to obtain historical discrete time series data.
[0066] For example, a series of quantile bin boundaries can be defined, and a unique corresponding discrete value can be determined based on the range of each group of quantile bin boundaries. The continuous value of each data point in the historical continuous time series data can be quantified to determine the discrete value corresponding to each continuous value, thereby obtaining the historical discrete time series data.
[0067] S230 , time-slicing the historical discrete time series data to obtain historical discrete time series data for each historical time period.
[0068] The time span of the historical time period can be preset and adjusted by the technicians. The time span of the historical time period affects the time span of the input data of the large language model.
[0069] Specifically, the historical discrete time series data may be time-sliced by using a preset time span of a historical time period to obtain the historical discrete time series data of each historical event segment.
[0070] Optionally, the historical discrete time series data may be time-sliced by adding a start mark and an end mark to the historical discrete time series data. Exemplarily, the start mark may be "t_start"; and the end mark may be "t_end".
[0071] In an optional embodiment of the present invention, historical discrete time series data are time-sliced to obtain historical discrete time series data for each historical time period, including: clustering the historical discrete time series data to obtain at least one historical discrete time series cluster center data and historical discrete time series cluster cluster data corresponding to each historical discrete time series cluster center data; selecting at least two historical discrete time series from the historical discrete time series cluster cluster data corresponding to each historical discrete time series cluster center data; obtaining the time span of the historical time period; for a single historical discrete time series, slicing the historical discrete time series according to the time span of the historical time period to obtain historical discrete time sub-sequences corresponding to each historical time period; combining the historical discrete time sub-sequences corresponding to each historical time period of each historical discrete time series to obtain each enhanced discrete time series; determining the historical discrete time series data for each historical time period according to the historical discrete time sub-sequences corresponding to each historical time period of each historical discrete time series and the historical discrete time sub-sequences corresponding to each historical time period of each enhanced discrete time series.
[0072] The historical discrete time series cluster center data may be the cluster center data of the historical discrete time series data. The historical discrete time series cluster cluster data corresponding to the historical discrete time series cluster center data may be the clustering result of the historical discrete time series data. The historical discrete time series may be part of the historical discrete time series cluster cluster data. The historical discrete time subsequence may be the time slicing result of the historical discrete time series. The enhanced discrete time series may be the combination result between different historical discrete time subsequences.
[0073] Specifically, a clustering algorithm can be used to cluster the historical discrete time series data to obtain at least one historical discrete time series cluster center data and historical discrete time series cluster cluster data corresponding to each historical discrete time series cluster center data. Based on the number of samples contained in the historical discrete time series, at least two historical discrete time series can be randomly selected from the historical discrete time series cluster cluster data corresponding to each historical discrete time series cluster center data. The time span of the pre-set historical time period can be obtained. For a single historical discrete time series, the historical discrete time series is fragmented according to the time span of the historical time period to obtain the historical discrete time subsequences corresponding to each historical time period. The historical discrete time subsequences corresponding to each historical time period of each historical discrete time series are combined to obtain each enhanced discrete time series. The historical discrete time subsequences corresponding to each historical time period of each historical discrete time series and the historical discrete time subsequences corresponding to each historical time period of each enhanced discrete time series can be used to determine the historical discrete time series data of each historical time period. Among them, the clustering algorithm can be a K-means (K-Means Clustering Algorithm, K-means clustering) algorithm.
[0074] This scheme introduces the clustering and enhancement process of historical discrete time series data, takes into account the data characteristics of historical discrete time series data, avoids the impact of too little data on the prediction accuracy of the large language model, and further improves the prediction accuracy of the large language model.
[0075] S240. For a single historical time period, a large language model is used to predict the historical discrete time series data of the historical time period to obtain predicted discrete time series data of the predicted time period.
[0076] The time span of the prediction time period can be preset and adjusted by the technician. The time span of the prediction time period affects the time span of the output result of the large language model. Optionally, the time span of the prediction time period can be the same as or different from the time span of the historical time period. The predicted discrete time series data can be the predicted data of the prediction time period obtained by the large language model based on the historical discrete time series data of the historical time period.
[0077] Specifically, the historical discrete time series data of the historical time period and the corresponding prompt words can be input into the large language model, and the large language model is used to perform semantic understanding on the historical discrete time series data of the historical time period to obtain the predicted discrete time series data of the predicted time period.
[0078] S250. For a single historical time period, adjust the parameters of the large language model according to the difference between the predicted discrete time series data of the predicted time period and the historical discrete time series data of the predicted time period.
[0079] The predicted discrete time series data of the prediction time period can be used to characterize the prediction results of the large language model. The historical discrete time series data of the prediction time period can be used to characterize the actual results. The difference between the predicted discrete time series data of the prediction time period and the historical discrete time series data of the prediction time period can be used to characterize the difference between the prediction results and the actual results.
[0080] Specifically, for a single historical time period, the loss function corresponding to the large language model can be used to calculate the difference between the predicted discrete time series data of the predicted time period and the historical discrete time series data of the predicted time period. An optimization algorithm can be used to adjust the parameters of the large language model with the goal of minimizing the loss function. The loss function can be a cross entropy loss function; the optimization algorithm can be a gradient descent method or a variant thereof.
[0081] S260: Obtain current continuous time series data of the current time period.
[0082] S270: quantize the current continuous time series data to obtain the current discrete time series data of the current time period.
[0083] S280: Use a large language model to predict the current discrete time series data to obtain target discrete time series data for a target time period.
[0084] S290, de-discretize the target discrete time series data to obtain target continuous time series data in the target time period.
[0085] The technical solution of the embodiment of the present invention obtains historical continuous time series data, quantizes the historical continuous time series data, obtains historical discrete time series data, time slices the historical discrete time series data, obtains historical discrete time series data for each historical time period, and for a single historical time period, uses a large language model to predict the historical discrete time series data for the historical time period to obtain predicted discrete time series data for the predicted time period, and for a single historical time period, adjusts the parameters of the large language model according to the difference between the predicted discrete time series data for the predicted time period and the historical discrete time series data for the predicted time period, introduces a pre-training process of the large language model, and can improve the prediction efficiency of time series data.
[0086] At present, time series data prediction mainly relies on the following methods:
[0087] 1. Statistical models: such as the autoregressive moving average model, the autoregressive integrated moving average model, or the seasonal difference autoregressive moving average model, etc. However, statistical models often overfit specific types of data, resulting in a decrease in the model's predictive performance when encountering data with different distributions or novel patterns. For example, statistical models are suitable for linear and static time series data, and perform poorly for data with complex nonlinear relationships. For another example, the autoregressive integrated moving average model requires the data to meet the stationary assumption, while the data in actual applications is often non-stationary, so the predictive performance in the actual prediction process is poor.
[0088] 2. Machine learning models: such as support vector machines (SVM), random forests (RF) or gradient boosted trees (GBT), etc. However, machine learning models can handle a certain degree of nonlinear relationships, but usually require complex feature engineering and model parameter adjustment processes, requiring experts to perform model design, feature engineering and parameter adjustment, which increases the difficulty of practical application and limits the use of models by non-professionals.
[0089] 3. Deep learning models: such as recurrent neural networks (RNNs), long short-term memory networks (LSTMs) or gated recurrent units (GRUs), etc. These deep learning models perform well in processing time series data. However, the training process of deep learning models usually requires a lot of computing resources and a large amount of data, and there may also be computing efficiency issues in the inference stage, which limits their application in resource-constrained environments and is very sensitive to hyperparameters.
[0090] In order to solve the above problems, based on the above embodiments, this solution proposes a preferred embodiment, that is, this solution proposes a time series data prediction method based on large language models (LLMs), which specifically includes the following steps:
[0091] Step 1: Data preprocessing.
[0092] Data preprocessing can ensure that the original continuous time series data (ie, historical continuous time series data) can be converted into the format required by large language models (LLMs).
[0093] Step 1-1 Data scaling (ie normalization).
[0094] In order to make the historical continuous time series data suitable for subsequent data quantization and data tokenization, the historical continuous time series data is first scaled, thereby normalizing each data in the historical continuous time series data to a unified range and reducing the prediction error caused by the numerical value.
[0095] Specifically, given each data point x in the historical continuous time series data i, where i = 1, 2, ..., N, and N is the total number of data points in the continuous time series data.
[0096] The data points in the historical continuous time series data can be converted to normalized values using mean scaling:
[0097]
[0098] In the formula, is the standard value; x i is each data point in the historical continuous time series data; x i are the data points in the historical continuous time series data; m is the mean of the data points in the historical continuous time series data; s is the average absolute deviation of the data points in the historical continuous time series data.
[0099] Among them, the following formula can be used to calculate the mean of each data point in the historical continuous time series data:
[0100]
[0101] Where m is the mean of each data point in the historical continuous time series data; N is the total number of data points in the historical continuous time series data; x i are the data points in the historical continuous time series data.
[0102] The following formula can be used to calculate the average absolute deviation of each data point in the historical continuous time series data:
[0103]
[0104] Where s is the average absolute deviation of each data point in the historical continuous time series data; N is the total number of data points in the historical continuous time series data; x i are the data points in the historical continuous time series data; m is the mean of the data points in the historical continuous time series data.
[0105] Step 1-2 Data quantification (i.e. discretization).
[0106] After scaling, because large language models cannot directly process continuous data, continuous values in historical normative time series data need to be converted into discrete token values. The data quantization step can map the normative values after data scaling to a finite set of discrete values.
[0107] First, we can define a series of quantile bin boundaries b1, b2, ..., b n-1 , where b j j+1 And j = 1, 2, ..., n-1. Each bin [bj ,b j+1 ) corresponds to a unique discrete value t j .
[0108] The quantization function can be expressed by the following formula:
[0109]
[0110] Where Q is the quantization function; is the standard value in the historical standard time series data; t j is the discrete value corresponding to the standard value in the historical standard time series data; b j is the upper bound of the bin; b j+1 is the lower bound of the bin.
[0111] These bin boundaries can be used to quantify the data distribution based on the historical normative time series data (for example, uniform binning or binning based on a specific quantile of the data set can be used). In this way, the continuous historical normative time series data can be mapped to the historical discrete time series data T = {t1, t2, ..., t n-1}.
[0112] After the quantization step, each data point in the historical normative time series data is assigned a discrete value. These discrete values can be organized into historical discrete time series data in chronological order. The historical discrete time series data is then used to train the large language model. In addition, in order to adapt to the processing requirements of the large language model, special tokens (i.e., start and end markers) can be introduced, including padding tokens and end of sequence tokens (EOS). The above special tokens help the large language model to identify the boundaries of the historical discrete time series data.
[0113] The data preprocessing method in the above steps ensures that the original historical continuous time series data can be efficiently converted into the format required by the large language model, thus laying a solid foundation for the subsequent training and reasoning of the large language model.
[0114] Step 2: Data clustering and data augmentation.
[0115] In order to improve the generalization ability of large language models on different time series data, this solution adopts a series of data clustering and data enhancement strategies.
[0116] Step 2-1 Data clustering.
[0117] First, the K-means algorithm can be used to cluster the historical discrete time series data. In this way, data points with similar characteristics in the historical discrete time series data can be grouped into the same cluster. The goal of the K-means algorithm is to minimize the sum of squares from each data point to its cluster center (centroid).
[0118] Specifically, for the historical discrete time series data T = {t1, t2, ..., t n-1}, find a set of centroids {μ1,μ2,…,μ K}, so that the following objective function is minimized:
[0119]
[0120] Where J is the objective function; T is the historical discrete time series data; r jk is an indicator variable, indicating data point t j Does it belong to cluster k? If the data point t j belongs to cluster k, then r jk =1; otherwise, r jk =0.
[0121] Thus, at least one historical discrete time series cluster center data and historical discrete time series cluster cluster data corresponding to each historical discrete time series cluster center data are obtained.
[0122] Step 2-2 Data enhancement.
[0123] Based on data clustering, data enhancement can be used to enrich the training set of the large language model. This method generates new enhanced discrete time series by combining historical discrete time subsequences of different historical time periods, thereby obtaining new training samples. The specific steps are as follows:
[0124] Step 2-2-1 Randomly select at least two historical discrete time series from each historical discrete time series cluster data to form O = {o1, o2, ..., o E}.
[0125] Step 2-2-2 For each selected historical discrete time series, further randomly select a time span l of a historical time period, and extract the historical discrete time subsequence O corresponding to each historical time period from each sequence. h ={o 1h ,o 2h ,……,o ah}, where h represents the hth historical discrete time series; h = 1, 2, ..., E. a represents the ath historical discrete time subsequence, a = L / l; L is the time span of the historical discrete time series; l is the time span of the historical discrete time subsequence.
[0126] Step 2-2-3 Generate a set of weights ω1, ω2, ..., ω according to the symmetric Dirichlet distribution a ,satisfy And ω i ≥ 0. The parameter α of the symmetric Dirichlet distribution can control the distribution shape of the weights and can be adjusted according to actual needs.
[0127] Step 2-2-3 uses the generated weights to perform convex combination on the selected historical discrete time subsequences to obtain a new enhanced discrete time series:
[0128]
[0129] In the formula, S aug To enhance discrete time series; ω i is the weight of the i-th historical discrete time subsequence; o ih is the i-th historical discrete time subsequence selected from the h-th historical discrete time series.
[0130] Through this data augmentation method, not only can a large number of diverse training samples for large language models be generated, but the temporal continuity and intrinsic correlation in the original historical discrete time series data can also be preserved.
[0131] Step 3: Training and reasoning of large language models.
[0132] The final stage of this solution is to train and infer the large language model so that it can accurately predict time series data. This stage mainly includes the following steps:
[0133] Step 3-1 Model training.
[0134] After completing data clustering and data enhancement, the enhanced discrete time series is obtained. According to the historical discrete time subsequences corresponding to each historical time period of each historical discrete time series and the historical discrete time subsequences corresponding to each historical time period of each enhanced discrete time series, the historical discrete time series data of each historical time period are determined, and these historical discrete time series data are then used to train the large language model. The specific process is as follows:
[0135] Step 3-1-1 Loss function definition.
[0136] Choose a suitable loss function to evaluate the difference between the prediction results of the large language model and the actual results. Preferably, a cross entropy loss function can be selected to measure the difference between the probability assigned by the large language model to each possible output and the true label distribution. For a multi-class classification problem with C categories, the cross entropy loss function is defined as:
[0137]
[0138] In the formula, H is the cross entropy loss function; y is the actual result; is the prediction result of the large language model; i is the predicted category; C is the total number of categories.
[0139] Step 3-1-2 Optimization algorithm selection.
[0140] Gradient Descent or its variant algorithms can be used as the optimization algorithm to minimize the loss function and update the parameters of the large language model.
[0141] Step 3-1-3 Training process of large language model.
[0142] In each training iteration of the large language model, the large language model receives the historical discrete time subsequence of a historical time period as input data, and predicts the probability distribution of each value in the next prediction time period, and the value with the highest probability constitutes the predicted discrete time series data of the prediction time period. The loss value is calculated by comparing the predicted discrete time series data with the historical discrete time subsequence of the prediction time period, and the optimization algorithm is used to adjust the large language model parameters.
[0143] Step 3-2 Reasoning and post-processing of the large language model.
[0144] After the model training is completed, you can use the trained large language model for inference. The inference process is as follows:
[0145] Step 3-2-1 generates alternative discrete time series data for alternative time periods.
[0146] The large language model can recursively generate candidate discrete time series data for the next candidate time period according to the current discrete time series data for the current time period, until the target discrete time series data for the target time period is reached.
[0147] Step 3-2-2 Data dequantization.
[0148] Map the target discrete time series data of the target time period output by the large language model back to a continuous numerical range.
[0149] The following formula can be used to express the inverse quantization function:
[0150]
[0151] In the formula, Q -1 is the inverse quantization function; t j is the discrete value in the target discrete time series data; b j is the upper bound of the bin; bj+1 is the lower bound of the bin; this discrete value t j , determined as the midpoint of the upper and lower bounds of the bin.
[0152] Step 3-2-3: Denormalize the data.
[0153] Through data dequantization, the normative values in the target normative time series data are restored to the original scale. For each normative value in the target normative time series data, the following formula can be used for denormalization:
[0154]
[0155] In the formula, x i are the continuous values in the historical continuous time series data; is the standard value in the target standard time series data; s is the average absolute deviation of each data point in the historical continuous time series data; m is the mean of each data point in the historical continuous time series data.
[0156] Through this step, the target continuous time series data of the final target time period can be obtained.
[0157] Step 3-3 Post-processing.
[0158] In some cases, additional data post-processing steps can be performed to ensure that the forecast results (i.e., the target continuous time series data for the target time period) meet specific business rules or constraints. For example, for financial time series data, it may be necessary to ensure that the predicted resource transfer amount is not less than a certain minimum value; for weather forecast data, it may be necessary to ensure that the humidity value is between 0 and 1.
[0159] Through the above training and reasoning process, large language models can be effectively applied to the prediction tasks of time series data, which not only improves the prediction accuracy of large language models, but also simplifies the entire data processing and model training process, so that non-experts can easily use this time series data prediction method. By simplifying the data preprocessing process, improving the generalization ability and prediction accuracy of the model, and reducing the complexity of model deployment and maintenance, it has brought improvements to the field of time series data analysis.
[0160] Embodiment 3
[0161] Figure 3 This is a schematic diagram of the structure of a time series data prediction device provided in Embodiment 3 of the present invention. The embodiment of the present invention can be applied to the case of predicting time series data, the device can execute the time series data prediction method, the device can be implemented in the form of hardware and / or software, and the device can be configured in an electronic device that carries the time series data prediction function, such as a client or a server.
[0162] See also Figure 3 The time series data prediction device shown includes: a current continuous time series data acquisition module 310, a current continuous time series data quantization module 320, a target discrete time series data prediction module 330 and a target discrete time series data dequantization module 340. Among them, the current continuous time series data acquisition module 310 is used to obtain the current continuous time series data of the current time period; the current continuous time series data quantization module 320 is used to quantize the current continuous time series data to obtain the current discrete time series data of the current time period; the target discrete time series data prediction module 330 is used to use a large language model to predict the current discrete time series data to obtain the target discrete time series data of the target time period; the target discrete time series data dequantization module 340 is used to de-discretize the target discrete time series data to obtain the target continuous time series data of the target time period.
[0163] The technical solution of the embodiment of the present invention obtains the current continuous time series data of the current time period, quantizes the current continuous time series data, and obtains the current discrete time series data of the current time period, thereby realizing the discretization of the current continuous time series, which is convenient for the large language model to predict the time series data. The large language model is adopted to predict the current discrete time series data to obtain the target discrete time series data of the target time period, and the target discrete time series data is de-discretized to obtain the target continuous time series data of the target time period. The large language model is introduced, and the semantic understanding ability of the large language model is utilized to improve the accuracy of time series data prediction. Moreover, the excessive occupation of system resources by training a large number of models for a large number of indicator data in the time series data is avoided, the occupation of system resources is reduced, the model deployment process is simplified, the technical barriers to the use of the model are reduced, the maintenance of the large language model is facilitated, and the applicability and generalization of time series data prediction are improved.
[0164] In an optional embodiment of the present invention, the target discrete time series data prediction module 330 includes: a target time span acquisition unit, which is used to use a large language model to acquire the time span of the target time period; an alternative discrete time series data prediction unit, which is used to use the large language model to predict the current discrete time series data to obtain the alternative discrete time series data of the alternative time period; a time span comparison unit, which is used to use the large language model to detect whether the time span of the alternative time period reaches the time span of the target time period; a target discrete time series data generation unit, which is used to use the large language model to update the alternative time period to the current time period and the alternative discrete time series data to the current discrete time series data when the time span of the alternative time period does not reach the time span of the target time period, and return to execute the step of predicting the current discrete time series data to obtain the alternative discrete time series data of the alternative time period, until the time span of the alternative time period reaches the time span of the target time period, the alternative discrete time series data of each of the alternative time periods are combined in time order to obtain the target discrete time series data of the target time period.
[0165] In an optional embodiment of the present invention, the current continuous time series data quantization module 320 includes: a current continuous time series data normalization unit, which is used to normalize the current continuous time series data to obtain the current standard time series data of the current time period; a current standard time series data quantization unit, which is used to quantize the current standard time series data to obtain the current discrete time series data of the current time period; accordingly, the target discrete time series data dequantization module 340 includes: a target discrete time series data de-discretization unit, which is used to de-discretize the target discrete time series data to obtain the target standard time series data of the target time period; a target standard time series data de-normalization unit, which is used to de-normalize the target standard time series data to obtain the target continuous time series data of the target time period.
[0166] In an optional embodiment of the present invention, the device also includes: a historical continuous time series data acquisition module, which is used to acquire historical continuous time series data before acquiring the current continuous time series data of the current time period; a historical continuous time series data quantization module, which is used to quantize the historical continuous time series data to obtain historical discrete time series data; a historical discrete time series data time slicing module, which is used to time slice the historical discrete time series data to obtain historical discrete time series data for each historical time period; a predicted discrete time series data prediction module, which is used to use a large language model to predict the historical discrete time series data of a single historical time period, and obtain predicted discrete time series data of the predicted time period; a large language model parameter adjustment module, which is used to adjust the parameters of the large language model according to the difference between the predicted discrete time series data of the predicted time period and the historical discrete time series data of the predicted time period.
[0167] In an optional embodiment of the present invention, the historical discrete time series data time slicing module includes: a historical discrete time series data clustering unit, which is used to cluster the historical discrete time series data to obtain at least one historical discrete time series cluster center data and historical discrete time series cluster cluster data corresponding to each of the historical discrete time series cluster center data; a historical discrete time series selection unit, which is used to select at least two historical discrete time series from the historical discrete time series cluster cluster data corresponding to each of the historical discrete time series cluster center data; a historical time period time span acquisition unit, which is used to acquire the time span of the historical time period; a historical discrete time series slicing unit, which is used to select at least two historical discrete time series from the historical discrete time series cluster cluster data corresponding to each of the historical discrete time series cluster center data; The historical discrete time series are segmented according to the time span of the historical time periods to obtain historical discrete time subsequences corresponding to the historical time periods; an enhanced discrete time series generating unit is used to combine the historical discrete time subsequences corresponding to the historical time periods of each historical discrete time series to obtain each enhanced discrete time series; a historical discrete time series data generating unit is used to determine the historical discrete time series data of each historical time period according to the historical discrete time subsequences corresponding to the historical time periods of each historical discrete time series and the historical discrete time subsequences corresponding to the historical time periods of each enhanced discrete time series.
[0168] In an optional embodiment of the present invention, the current continuous time series data is current market analysis continuous time series data.
[0169] The time series data prediction device provided in the embodiment of the present invention can execute the time series data prediction method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.
[0170] In the technical solution of the embodiment of the present invention, the collected information is information and data authorized by the user or fully authorized by all parties, and the collection, storage, use, processing, transmission, provision, disclosure and application of the relevant data comply with the relevant laws, regulations and standards of the relevant countries and regions, take necessary confidentiality measures, do not violate public order and good morals, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0171] Embodiment 4
[0172] According to an embodiment of the present invention, the present invention also provides an electronic device, a readable storage medium and a computer program product.
[0173] Figure 4 A schematic diagram of the structure of an electronic device 400 that can be used to implement an embodiment of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or required herein.
[0174] like Figure 4 As shown, the electronic device 400 includes at least one processor 401, and a memory connected to the at least one processor 401 in communication, such as a read-only memory (ROM) 402, a random access memory (RAM) 403, etc., wherein the memory stores a computer program that can be executed by at least one processor, and the processor 401 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 402 or the computer program loaded from the storage unit 408 to the random access memory (RAM) 403. In the RAM 403, various programs and data required for the operation of the electronic device 400 can also be stored. The processor 401, the ROM 402, and the RAM 403 are connected to each other via a bus 404. An input / output (I / O) interface 405 is also connected to the bus 404.
[0175] Multiple components in the electronic device 400 are connected to the I / O interface 405, including: an input unit 406, such as a keyboard, a mouse, etc.; an output unit 407, such as various types of displays, speakers, etc.; a storage unit 408, such as a disk, an optical disk, etc.; and a communication unit 409, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 409 allows the electronic device 400 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.
[0176] Processor 401 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of processor 401 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. Processor 401 performs the various methods and processes described above, such as a time series data prediction method.
[0177] In some embodiments, the time series data prediction method may be implemented as a computer program, which is tangibly contained in a computer-readable storage medium, such as a storage unit 408. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 400 via the ROM 402 and / or the communication unit 409. When the computer program is loaded into the RAM 403 and executed by the processor 401, one or more steps of the time series data prediction method described above may be performed. Alternatively, in other embodiments, the processor 401 may be configured to perform the time series data prediction method in any other appropriate manner (e.g., by means of firmware).
[0178] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0179] Computer programs for implementing the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that when the computer program is executed by the processor, the functions / operations specified in the flow chart and / or block diagram are implemented. The computer program may be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a stand-alone software package, or entirely on a remote machine or server.
[0180] In the context of the present invention, a computer-readable storage medium may be a tangible medium that may contain or store a computer program for use by or in combination with an instruction execution system, device or equipment. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices or equipment, or any suitable combination of the foregoing. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0181] To provide interaction with a user, the systems and techniques described herein may be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or trackball) through which the user can provide input to the electronic device. Other types of devices may also be used to provide interaction with the user; for example, the feedback provided to the user may be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user may be received in any form (including acoustic input, voice input, or tactile input).
[0182] The systems and techniques described herein may be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.
[0183] A computing system may include a client and a server. The client and the server are generally remote from each other and usually interact through a communication network. The relationship between the client and the server is generated by computer programs running on the corresponding computers and having a client-server relationship with each other. The server may be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system to solve the defects of difficult management and weak business scalability in traditional physical hosts and VPS (Virtual Private Server) services.
[0184] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps described in the present invention can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solution of the present invention can be achieved, and this document does not limit this.
[0185] The above specific implementations do not constitute a limitation on the protection scope of the present invention. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modification, equivalent substitution and improvement made within the spirit and principle of the present invention should be included in the protection scope of the present invention.
Claims
1. A time series data prediction method, characterized in that: The method comprises: Get the current continuous time series data for the current time period; quantizing the current continuous time series data to obtain current discrete time series data of the current time period; Using a large language model, predicting the current discrete time series data to obtain target discrete time series data for a target time period; The target discrete time series data is de-discretized to obtain the target continuous time series data of the target time period.
2. The method according to claim 1, characterized in that The method of using a large language model to predict the current discrete time series data to obtain target discrete time series data for a target time period includes: Use a large language model to obtain the time span of the target time period; Using the large language model, predicting the current discrete time series data to obtain candidate discrete time series data for a candidate time period; Using the large language model, detecting whether the time span of the candidate time period reaches the time span of the target time period; By using the large language model, when the time span of the alternative time period does not reach the time span of the target time period, the alternative time period is updated to the current time period, the alternative discrete time series data is updated to the current discrete time series data, and the step of predicting the current discrete time series data to obtain the alternative discrete time series data of the alternative time period is returned to execute, until the time span of the alternative time period reaches the time span of the target time period, the alternative discrete time series data of each of the alternative time periods are combined in chronological order to obtain the target discrete time series data of the target time period.
3. The method according to claim 1, characterized in that The step of quantizing the current continuous time series data to obtain the current discrete time series data of the current time period includes: Normalizing the current continuous time series data to obtain current standard time series data for the current time period; quantizing the current standard time series data to obtain current discrete time series data of the current time period; Accordingly, the de-discretizing the target discrete time series data to obtain the target continuous time series data of the target time period includes: De-discretizing the target discrete time series data to obtain target standard time series data for the target time period; The target normative time series data is denormalized to obtain the target continuous time series data of the target time period.
4. The method according to claim 1, characterized in that Before obtaining the current continuous time series data of the current time period, the method further includes: Get historical continuous time series data; quantifying the historical continuous time series data to obtain historical discrete time series data; Performing time slicing on the historical discrete time series data to obtain historical discrete time series data for each historical time period; For a single historical time period, a large language model is used to predict the historical discrete time series data of the historical time period to obtain predicted discrete time series data of the predicted time period; For a single historical time period, the large language model is adjusted according to the difference between the predicted discrete time series data of the predicted time period and the historical discrete time series data of the predicted time period.
5. The method according to claim 4, characterized in that The step of time slicing the historical discrete time series data to obtain the historical discrete time series data for each historical time period includes: Clustering the historical discrete time series data to obtain at least one historical discrete time series cluster center data and historical discrete time series cluster cluster data corresponding to each of the historical discrete time series cluster center data; Select at least two historical discrete time series from the historical discrete time series cluster cluster data corresponding to each of the historical discrete time series cluster center data; Get the time span of the historical time period; For a single historical discrete time series, the historical discrete time series is segmented according to the time span of the historical time period to obtain historical discrete time subsequences corresponding to each historical time period; Combining the historical discrete time subsequences corresponding to the historical time periods of the historical discrete time series to obtain enhanced discrete time series; According to the historical discrete time subsequences corresponding to the historical time periods of the historical discrete time series and the historical discrete time subsequences corresponding to the historical time periods of the enhanced discrete time series, the historical discrete time series data of each historical time period are determined.
6. The method according to claim 1, characterized in that The current continuous time series data is current market analysis continuous time series data.
7. A time series data prediction device, characterized in that: The device comprises: The current continuous time series data acquisition module is used to acquire the current continuous time series data of the current time period; A current continuous time series data quantization module is used to quantize the current continuous time series data to obtain the current discrete time series data of the current time period; A target discrete time series data prediction module is used to use a large language model to predict the current discrete time series data to obtain target discrete time series data for a target time period; The target discrete time series data dequantization module is used to de-discretize the target discrete time series data to obtain the target continuous time series data of the target time period.
8. An electronic device, characterized in that: The electronic device comprises: at least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the time series data prediction method according to any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the time series data prediction method according to any one of claims 1 to 6 when executed.
10. A computer program product, characterized in that The computer program product comprises a computer program, which, when executed by a processor, implements the time series data prediction method according to any one of claims 1 to 6.