Data prediction method and device, equipment, storage medium and product

Through the attention layer and alternative prediction layer in the Transformer model, the problem of high hardware resources and management costs caused by training multiple models in the prior art is solved, and time series predictions of flexible adaptation to different time lengths are achieved.

CN120277387APending Publication Date: 2025-07-08BEIJING 360 INTELLIGENT TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510320762.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-18
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

Existing deep learning models are difficult to flexibly adapt to time series prediction requirements of different time lengths, resulting in the need to train multiple models, increasing hardware resource overhead and management and maintenance costs.

Method used

Using the attention layer and multiple alternative prediction layers in the transformer Transformer model, a single model can achieve time series prediction of different time lengths by filtering and calling suitable prediction layers.

Benefits of technology

A single model is realized to adapt to short-term, medium-term and long-term forecast needs, reducing the management and maintenance costs of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120277387A_ABST
    Figure CN120277387A_ABST
Patent Text Reader

Abstract

The invention discloses a data prediction method and device, equipment, a storage medium and a product, relates to the technical field of artificial intelligence, and discloses a method for extracting a reference time sequence and a target prediction number from an instruction in response to the data prediction instruction, the target prediction quantity represents the quantity of to-be-predicted target observation values; obtaining feature vectors of the plurality of reference observation values through an attention layer in a Transform model of a converter; based on the reference prediction numbers corresponding to various alternative prediction layers in the Transform model, determining a target prediction layer of which the reference prediction number is matched with the target prediction number from the various alternative prediction layers; and through a target prediction layer, performing prediction based on the feature vectors of the plurality of reference observation values to obtain a target prediction number of target observation values. According to the method, time series prediction of different time durations can be realized through a single model, so that the cost is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular, to a data prediction method, apparatus, device, storage medium, and product. Background Art

[0002] Time series prediction is a key task in many fields and is widely used in fields such as finance, energy, and meteorology. In recent years, deep learning models have made significant progress in the field of time series prediction, greatly promoting the technological development of related fields. However, most of the existing deep learning models are for specific prediction ranges and are difficult to flexibly adapt to the needs of different prediction lengths.

[0003] To enable the model to perform time series predictions of different time lengths, the general approach is to train multiple models for different prediction ranges such as short-term, medium-term, and long-term. However, training multiple models incurs high hardware resource overhead, requires more hardware support during prediction, and has extremely high management and maintenance costs.

[0004] The above content is only used to assist in understanding the technical solution of this application and does not represent an admission that the above content is prior art. Summary of the Invention

[0005] The main purpose of this application is to provide a data prediction method, apparatus, device, storage medium, and product, which can achieve time series predictions of different time lengths through a single model, thereby reducing costs.

[0006] To achieve the above object, this application proposes a data prediction method, and the method includes:

[0007] In response to a data prediction instruction, extract a reference time series and a target prediction quantity from the data prediction instruction, where the reference time series includes a plurality of reference observations arranged in chronological order, and the target prediction quantity represents the number of target observations to be predicted after the reference time series;

[0008] Obtain feature vectors of the plurality of reference observations through an attention layer in a Transformer model;

[0009] Based on the reference prediction quantities respectively corresponding to multiple alternative prediction layers in the Transformer model, determine a target prediction layer whose reference prediction quantity matches the target prediction quantity from the multiple alternative prediction layers;

[0010] Perform a prediction based on the feature vectors of the plurality of reference observations through the target prediction layer to obtain target observations of the target prediction quantity.

[0011] Optionally, determining a target prediction layer with a reference prediction quantity matching the target prediction quantity from the multiple alternative prediction layers in the Transformer model, includes:

[0012] Based on the reference prediction quantities respectively corresponding to the multiple alternative prediction layers in the Transformer model, determining a target set, the target set includes at least one target prediction layer determined from the multiple alternative prediction layers and the quantities of the determined various target prediction layers, and the sum of the reference prediction quantities corresponding to all the target prediction layers in the target set is equal to the target prediction quantity.

[0013] Optionally, determining the target set based on the reference prediction quantities respectively corresponding to the multiple alternative prediction layers in the Transformer model, includes:

[0014] Initializing the predicted quantity to 0, and looping through the prediction layer screening process to update the predicted quantity until the predicted quantity reaches the target prediction quantity, and forming the target set with the screened various target prediction layers and the quantities of the screened various target prediction layers, the prediction layer screening process includes:

[0015] Determining the judgment parameter of each alternative prediction layer in turn according to the corresponding reference prediction quantity from large to small, and comparing the judgment parameter with the target prediction quantity, the judgment parameter is the sum of the reference prediction quantity and the predicted quantity;

[0016] If the judgment parameter of the currently judged alternative prediction layer is greater than the target prediction quantity, then judge the next alternative prediction layer with the second largest reference prediction quantity;

[0017] If the judgment parameter of the currently judged alternative prediction layer is less than or equal to the target prediction quantity, then determine the currently judged alternative prediction layer as an available target prediction layer, and update the predicted quantity to the sum of the reference prediction quantities corresponding to all the current target prediction layers.

[0018] Optionally, predicting, through the target prediction layer, a target observation value of the target prediction quantity based on the feature vectors of the multiple reference observation values, includes:

[0019] Determining the calling order of the target prediction layers in the target set;

[0020] In the order of the calls, sequentially pass through the target prediction layers in the target set, make predictions based on the feature vectors of multiple reference observations in the current reference time series, obtain the target observations of the reference prediction quantity corresponding to the target prediction layer, and add the target observations of the reference prediction quantity after the multiple reference observations to obtain an updated reference time series until all the target prediction layers in the target set participate in the prediction to obtain the target observations of the target prediction quantity.

[0021] Optionally, determining the call order of the target prediction layers in the target set includes:

[0022] Sort the various target prediction layers in the target set in descending order according to the corresponding reference prediction quantity;

[0023] In the case where there are multiple target prediction layers of the same type, the call order of the multiple target prediction layers of the same type is randomly arranged.

[0024] Optionally, the training process of the Transformer model includes:

[0025] Obtain a sample time series, where the sample time series includes multiple sample reference observations and multiple sample target observations arranged in chronological order;

[0026] Through the attention layer, obtain the feature vectors of the multiple sample reference observations;

[0027] Through the various alternative prediction layers, make predictions based on the feature vectors respectively to obtain the sample prediction observations of the reference prediction quantity corresponding to the various alternative prediction layers;

[0028] Train the Transformer model based on each sample prediction observation and each sample target observation.

[0029] Optionally, training the Transformer model based on each sample prediction observation and each sample target observation includes:

[0030] Based on the sample prediction observations corresponding to the various alternative prediction layers and the sample target observations corresponding to each sample prediction observation, determine the prediction losses of the various alternative prediction layers;

[0031] Determine the average prediction loss of the multiple alternative prediction layers as the model prediction loss;

[0032] Based on the model prediction loss, adjust the model parameters to reduce the model prediction loss.

[0033] Optionally, determining the prediction losses of the various alternative prediction layers based on the respective sample prediction observations corresponding to the various alternative prediction layers and the sample target observations corresponding to the respective sample prediction observations includes:

[0034] For any sample prediction observation corresponding to any alternative prediction layer, determining a first prediction loss based on the difference between the sample prediction observation and the sample target observation corresponding to the sample prediction observation;

[0035] Dividing the first prediction loss by the number of repeated predictions, where the number of repeated predictions is the number of times the sample target observation is repeatedly predicted by multiple alternative prediction layers in the Transformer model, to obtain a second prediction loss;

[0036] Adding up the second prediction losses corresponding to all the sample prediction observations corresponding to any alternative prediction layer to obtain the prediction loss of the alternative prediction layer.

[0037] Optionally, obtaining the sample time series includes:

[0038] Obtaining a sample sequence set, where the sample sequence set contains multiple sample time series, and the number of observations included in the multiple sample time series is different;

[0039] Sequentially obtaining the sample time series from the sample sequence set in ascending order of the number of included observations.

[0040] Optionally, the reference prediction quantities corresponding to the multiple alternative prediction layers in the Transformer model are 1, 4, 8, 16, 32, and 64 respectively.

[0041] In addition, to achieve the above object, the present application also proposes a data prediction device, and the device includes:

[0042] A prediction instruction response module, configured to extract a reference time series and a target prediction quantity from the data prediction instruction in response to the data prediction instruction, where the reference time series includes multiple reference observations arranged in chronological order, and the target prediction quantity represents the number of target observations to be predicted after the reference time series;

[0043] A feature vector acquisition module, configured to acquire the feature vectors of the multiple reference observations through the attention layer in the Transformer model;

[0044] A prediction layer selection module, configured to determine a target prediction layer whose reference prediction quantity matches the target prediction quantity from the multiple alternative prediction layers based on the reference prediction quantities corresponding to the multiple alternative prediction layers in the Transformer model;

[0045] An observation value prediction module, configured to perform prediction based on the feature vectors of the multiple reference observation values through the target prediction layer, and obtain the target observation values of the target prediction quantity.

[0046] Optionally, the prediction layer selection module is configured to determine a target set based on the reference prediction quantities respectively corresponding to multiple alternative prediction layers in the Transformer model. The target set includes at least one target prediction layer determined from the multiple alternative prediction layers and the quantities of the determined various target prediction layers, and the sum of the reference prediction quantities corresponding to all the target prediction layers in the target set is equal to the target prediction quantity.

[0047] Optionally, the prediction layer selection module is configured to initialize the predicted quantity to 0, and repeatedly execute a prediction layer screening process to update the predicted quantity until the predicted quantity reaches the target prediction quantity. The target set is composed of the various target prediction layers screened out and the quantities of the various target prediction layers screened out. The prediction layer screening process includes:

[0048] Determine the judgment parameter of each alternative prediction layer in turn according to the corresponding reference prediction quantity from largest to smallest, and compare the judgment parameter with the target prediction quantity. The judgment parameter is the sum of the reference prediction quantity and the predicted quantity.

[0049] If the judgment parameter of the currently judged alternative prediction layer is greater than the target prediction quantity, then judge the alternative prediction layer with the next largest reference prediction quantity.

[0050] If the judgment parameter of the currently judged alternative prediction layer is less than or equal to the target prediction quantity, then determine the currently judged alternative prediction layer as an available target prediction layer, and update the predicted quantity to the sum of the reference prediction quantities corresponding to all the current target prediction layers.

[0051] Optionally, the observation value prediction module includes:

[0052] An order determination unit, configured to determine the call order of the target prediction layers in the target set;

[0053] An observation value prediction unit, configured to sequentially perform prediction based on the feature vectors of multiple reference observation values in the current reference time series through the target prediction layers in the target set according to the call order, obtain the target observation values of the reference prediction quantity corresponding to the target prediction layer, and add the target observation values of the reference prediction quantity after the multiple reference observation values to obtain an updated reference time series until all the target prediction layers in the target set participate in the prediction, and obtain the target observation values of the target prediction quantity.

[0054] Optionally, the order determination unit is configured to sort various target prediction layers in the target set in descending order according to the corresponding reference prediction quantities; in the case of multiple target prediction layers of the same type, the call order of the multiple target prediction layers of the same type is randomly arranged.

[0055] Optionally, the apparatus further includes:

[0056] A sample acquisition module, configured to acquire a sample time series, where the sample time series includes a plurality of sample reference observation values and a plurality of sample target observation values arranged in chronological order;

[0057] The feature vector acquisition module is further configured to, through the attention layer, acquire feature vectors of the plurality of sample reference observation values;

[0058] The observation value prediction module is further configured to, through the multiple alternative prediction layers, respectively perform predictions based on the feature vectors to obtain sample prediction observation values corresponding to the reference prediction quantities of the respective alternative prediction layers;

[0059] A model training module, configured to train the Transformer model based on each sample prediction observation value and each sample target observation value.

[0060] Optionally, the model training module includes:

[0061] A layer loss determination unit, configured to determine the prediction loss of each alternative prediction layer based on each sample prediction observation value corresponding to each alternative prediction layer and the sample target observation value corresponding to each sample prediction observation value;

[0062] A model loss determination unit, configured to determine the average prediction loss of the multiple alternative prediction layers as the model prediction loss;

[0063] A parameter adjustment unit, configured to adjust model parameters based on the model prediction loss to reduce the model prediction loss.

[0064] Optionally, for any sample prediction observation value corresponding to any alternative prediction layer, the layer loss determination unit is configured to determine a first prediction loss based on the difference between the sample prediction observation value and the sample target observation value corresponding to the sample prediction observation value; divide the first prediction loss by the number of repeated predictions, where the number of repeated predictions is the number of times the sample target observation value is repeatedly predicted by the multiple alternative prediction layers in the Transformer model; and add up the second prediction losses corresponding to all sample prediction observation values corresponding to any alternative prediction layer to obtain the prediction loss of the alternative prediction layer.

[0065] Optionally, the sample acquisition module is configured to acquire a sample sequence set, where the sample sequence set includes a plurality of sample time series, and the number of observation values included in the plurality of sample time series is different; the sample time series are sequentially acquired from the sample sequence set in ascending order of the number of included observation values.

[0066] Optionally, the reference prediction quantities corresponding to multiple alternative prediction layers in the Transformer model are 1, 4, 8, 16, 32, and 64 respectively.

[0067] In addition, to achieve the above object, the present application further provides a data prediction device, where the device includes: a memory, a processor, and a computer program stored on the memory and executable on the processor, and the computer program is configured to implement the steps of the data prediction method as described above.

[0068] In addition, to achieve the above object, the present application further provides a storage medium, where the storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium, and when the computer program is executed by a processor, the steps of the data prediction method as described above are implemented.

[0069] In addition, to achieve the above object, the present application further provides a computer program product, where the computer program product includes a computer program, and when the computer program is executed by a processor, the steps of the data prediction method as described above are implemented.

[0070] One or more technical solutions proposed by the present application have at least the following technical effects:

[0071] The data prediction solution provided by this application extracts a reference time series and a target prediction quantity from a data prediction instruction in response to the data prediction instruction. The target prediction quantity represents the number of target observation values to be predicted after the reference time series and represents the time length required for prediction. Then, through the attention layer in the Transformer model, the feature vectors of multiple reference observations are obtained, providing effective information for subsequent prediction. Next, based on the reference prediction quantities corresponding to multiple alternative prediction layers in the Transformer model, a target prediction layer whose reference prediction quantity matches the target prediction quantity is determined from the multiple alternative prediction layers. That is, multiple alternative prediction layers applicable to different prediction lengths are integrated in the Transformer model, and a suitable target prediction layer is flexibly selected according to the specific target prediction quantity. Next, through the target prediction layer, predictions are made based on the feature vectors of multiple reference observations, and the target observation values of the target prediction quantity can be predicted. This solution realizes time series prediction of different time lengths through a single Transformer model containing multiple alternative prediction layers, and can flexibly adapt to short-term, medium-term, and long-term prediction requirements. Therefore, only this one model needs to be trained and maintained, greatly reducing the cost of managing and maintaining the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0072] The drawings herein are incorporated into the specification and form a part of this specification, showing embodiments consistent with this application, and are used together with the specification to explain the principles of this application.

[0073] In order to more clearly illustrate the technical solutions in the embodiments of this application or in the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, other drawings can also be obtained based on these drawings without creative efforts.

[0074] Figure 1 It is a schematic diagram of an implementation environment of a data prediction method of this application;

[0075] Figure 2 It is a schematic flowchart provided by the first embodiment of the data prediction method of this application;

[0076] Figure 3 It is a schematic flowchart provided by the second embodiment of the data prediction method of this application;

[0077] Figure 4 It is a schematic flowchart provided by the third embodiment of the data prediction method of this application;

[0078] Figure 5 It is a schematic flowchart of the newly added steps in the fourth embodiment of the data prediction method of this application;

[0079] Figure 6 It is a detailed schematic diagram of step S04 in the fifth embodiment of the data prediction method of this application;

[0080] Figure 7 It is a schematic structural diagram of a Transformer model provided by this application;

[0081] Figure 8 It is a schematic module structure diagram of the data prediction device in the embodiment of this application;

[0082] Figure 9 It is a schematic device structure diagram of the hardware operating environment involved in the data prediction method in the embodiment of this application.

[0083] The realization of the purpose, functional features and advantages of this application will be further described in combination with the embodiments with reference to the accompanying drawings. Detailed implementation manners

[0084] It should be understood that the specific embodiments described herein are only used to explain the technical solutions of this application and are not used to limit this application.

[0085] In order to better understand the technical solutions of this application, the following will be described in detail in combination with the accompanying drawings of the specification and specific implementation manners.

[0086] Figure 1 It is a schematic diagram of an implementation environment provided by an embodiment of the present disclosure. Refer to Figure 1 , this implementation environment includes a terminal 101 and a server 102. The terminal 101 and the server 102 are connected through a wireless or wired network. Exemplarily, a target application provided by the server 102 is installed on the terminal 101, and the terminal 101 can implement functions such as data transmission and message interaction through this target application.

[0087] Exemplarily, the terminal 101 is a computer, a mobile phone, a tablet computer or other terminals. Exemplarily, the target application is a target application in the operating system of the terminal 101, or a target application provided by a third party. For example, the target application is an energy management application, a financial data analysis application, a weather query application, etc. Exemplarily, the server 102 is the background server corresponding to this target application. Correspondingly, the server 102 is an energy management application server, a financial data analysis application server, a weather query application server, etc.

[0088] In this application, the terminal 101 is used to send a data prediction instruction to the server 102. The data prediction instruction includes a reference time series and a target prediction quantity. The reference time series includes a plurality of reference observation values arranged in chronological order, and the target prediction quantity represents the number of target observation values to be predicted after the reference time series. The server 102 is used to extract the reference time series and the target prediction quantity from the data prediction instruction in response to the data prediction instruction. Through the attention layer in the Transformer model (a model name), the feature vectors of the plurality of reference observation values are obtained. Based on the reference prediction quantities corresponding to the multiple alternative prediction layers in the Transformer model, the target prediction layer whose reference prediction quantity matches the target prediction quantity is determined from the multiple alternative prediction layers. Then, through the target prediction layer, prediction is performed based on the feature vectors of the plurality of reference observation values to obtain the target observation values of the target prediction quantity. Next, the server 102 is further used to send the target observation values of the target prediction quantity to the terminal 101. The terminal 101 is further used to receive and display the target observation values of the target prediction quantity.

[0089] Alternatively, the above data prediction process can also be completed by the terminal 101 alone. For example, it is completed by the target application installed on the terminal 101. The embodiments of this application do not limit this.

[0090] The data prediction method provided in this application is applicable to various scenarios. For example, in the scenario of weather prediction. The server extracts the reference time series of meteorological data and the target prediction quantity from the meteorological prediction instruction in response to the meteorological prediction instruction. The reference time series includes a plurality of meteorological observation values arranged in chronological order, such as the meteorological observation values corresponding to each day within a month. The target prediction quantity is the number of meteorological observation values to be predicted after the reference time series, such as 7. Then, the weather of each day within a week after the reference time series can be predicted by the method provided in this application. Another example is that the data prediction method can also be applied to the scenario of predicting the trend of financial data. The server extracts the reference time series of financial data and the target prediction quantity from the financial data prediction instruction in response to the financial data prediction instruction. The reference time series includes a plurality of financial data observation values arranged in chronological order, such as the financial data observation values corresponding to each hour within a day. The target prediction quantity is the number of financial data observation values to be predicted after the reference time series, such as 24. Then, the trend of financial data within 24 hours after the reference time series can be predicted by the method provided in this application.

[0091] Figure 2 It is a schematic flowchart of the first embodiment of the data prediction method of this application. Referring to Figure 2 , taking the execution entity as the terminal as an example, the data prediction method includes the following steps S10 to S40:

[0092] Step S10, in response to a data prediction instruction, extract a reference time series and a target prediction quantity from the data prediction instruction. The reference time series includes a plurality of reference observations arranged in chronological order, and the target prediction quantity represents the number of target observations to be predicted after the reference time series.

[0093] Among them, the data prediction instruction is issued by the terminal to the server, which is a command for predicting the future trend of specific data. This instruction contains key information about the prediction task, such as historical data providing information reference for prediction, the data range to be predicted, etc. The data prediction instruction is the trigger signal for the entire prediction process.

[0094] Exemplarily, in the stock price prediction scenario, the user can issue a command "predict the price of a certain stock in the next 5 days" to the terminal, and the terminal will generate a corresponding data prediction instruction, which includes the historical price of the stock and the number of days "5" of the price to be predicted. Then, send this data prediction instruction to the server.

[0095] The reference time series is a sequence composed of a plurality of reference observations arranged in chronological order. These observations are historical data and are the basic basis for prediction. For example, in weather prediction, the daily temperature values in the past week are reference observations, and arranging the daily temperature values of this week in order constitutes a reference time series.

[0096] The target prediction quantity refers to the number of target observations to be predicted after the reference time series. It stipulates the scope and length of the prediction. For example, to predict the sales volume in the next 3 months, the "3" among them is the target prediction quantity.

[0097] Step S20, through the attention layer in the Transformer model, obtain the feature vectors of a plurality of reference observations.

[0098] The Transformer model is a deep learning model based on the attention mechanism and has wide applications in the fields of natural language processing and time series prediction. The Transformer model mainly includes an attention layer and a prediction layer. The attention layer is used to extract data features, and the attention mechanism it adopts enables it to capture the dependencies between different positions in the sequence data and better understand the context information of the sequence. The prediction layer mainly makes predictions based on the data features extracted by the attention layer.

[0099] The attention layer is a core component in the Transformer model, and its key role is to capture the relationships between various elements in the sequence. Specifically, the attention layer assigns corresponding attention weights to each element in the sequence by calculating the degree of association between each element and all other elements in the sequence. These weights reflect the relative importance between different elements, enabling the model to dynamically focus on the information that is most crucial for the current element. Based on these determined attention weights, the attention layer further integrates the feature vectors of each element in the sequence. This process not only preserves the features of each element itself but also incorporates the information of other elements associated with it. In this way, the feature vector of each element is enriched, enabling it to more comprehensively reflect its context information in the sequence during subsequent processing.

[0100] The feature vector of the reference observation value is the vector representation obtained by processing the reference observation value through the attention layer. It contains the important feature information of the reference observation value and is the key input for subsequent prediction.

[0101] It should be noted that in the embodiments of this application, the basic structure before the prediction layer in the Transformer model is collectively referred to as the attention layer. Its core component is the attention mechanism, which is mainly responsible for extracting data features and providing information support for the prediction layer. Regarding the structure of this attention layer, the embodiments of this application do not make any restrictions and can be flexibly set according to needs.

[0102] Step S30: Based on the reference prediction quantities respectively corresponding to multiple alternative prediction layers in the Transformer model, determine a target prediction layer whose reference prediction quantity matches the target prediction quantity from the multiple alternative prediction layers.

[0103] The alternative prediction layers are the prediction layers pre-set in the Transformer model, and each alternative prediction layer corresponds to a specific reference prediction quantity. These alternative prediction layers can be selected according to different prediction requirements to achieve flexible prediction. The reference prediction quantity is the number of observation values that each alternative prediction layer can predict. For example, if the reference prediction quantity of a certain alternative prediction layer is 2, it means that it can predict the observation values at the next 2 time points.

[0104] Exemplarily, the alternative prediction layer is a feed-forward neural network (FFN), which can generate prediction results based on the input hidden state. In the embodiments of this application, the hidden state is the feature vector corresponding to the reference time series.

[0105] Optionally, the reference prediction quantities corresponding to multiple alternative prediction layers in the Transformer model are 1, 4, 8, 16, 32, and 64 respectively. That is to say, there are 6 different alternative prediction layers configured in the Transformer model, and the numbers of observed values at future time points that can be predicted are 1, 4, 8, 16, 32, and 64 respectively.

[0106] The target prediction layer is an alternative prediction layer selected from multiple alternative prediction layers, and the reference prediction quantity of which matches the target prediction quantity. It is responsible for performing actual prediction operations based on the feature vectors of reference observations. The matching of the reference prediction quantity of the alternative prediction layer and the target prediction quantity includes the following two cases:

[0107] The first case is that there is an alternative prediction layer whose reference prediction quantity is exactly equal to the target prediction quantity, and this alternative prediction layer is the target prediction layer. For example, to predict the future price trend of a certain stock, the target prediction quantity is 4 days, and there is exactly one alternative prediction layer in the Transformer model whose corresponding reference prediction quantity is 4, then this alternative prediction layer can be used as the target prediction layer.

[0108] The second case is that there is no alternative prediction layer whose reference prediction quantity is exactly equal to the target prediction quantity, but there are multiple alternative prediction layers whose corresponding reference prediction quantities sum up to the target prediction quantity. Then these multiple alternative prediction layers can be used as the target prediction layers. For example, to predict the future price trend of a certain stock, the target prediction quantity is 10 days, there is an alternative prediction layer A in the Transformer model whose corresponding reference prediction quantity is 8, and there is also an alternative prediction layer B whose corresponding reference prediction quantity is 1, then these two alternative prediction layers can be used as the target prediction layers. When predicting, first predict the stock price for the next 8 days through the alternative prediction layer A, then predict the stock price for the next 1 day through the alternative prediction layer B, and then predict the stock price for the next 1 day through the alternative prediction layer B, so as to complete the prediction of the stock price for the next 10 days.

[0109] It can be understood that the method of selecting the target prediction layer from multiple alternative prediction layers is flexible and diverse, as long as it is ensured that the reference prediction quantities corresponding to at least one determined target prediction layer can be combined into the target prediction quantity.

[0110] Step S40: Through the target prediction layer, perform a prediction based on the feature vectors of multiple reference observations to obtain the target observations with the target prediction quantity.

[0111] Exemplarily, when the number of target prediction layers is 1, the reference prediction quantity corresponding to the target prediction layer is equal to the target prediction quantity. Then, through this target prediction layer, prediction is performed based on the feature vectors of multiple reference observations, and thus the target observation value of the target prediction quantity can be obtained. When the number of target prediction layers is multiple, the sum of the reference prediction quantities corresponding to the multiple target prediction layers is equal to the target prediction quantity. Then, through the multiple target prediction layers, prediction is performed sequentially, and thus the target observation value of the target prediction quantity can be obtained.

[0112] The target observation value is the result obtained by prediction through the target prediction layer and is the predicted value of a specific number of observations after the reference time series. For example, when predicting electricity consumption, the target observation value is the electricity consumption value predicted for the next few days.

[0113] The data prediction solution provided in this application extracts the reference time series and the target prediction quantity from the data prediction instruction in response to the data prediction instruction. The target prediction quantity represents the number of target observation values to be predicted after the reference time series and represents the time length required for prediction. Then, through the attention layer in the Transformer model, the feature vectors of multiple reference observations are obtained to provide effective information for subsequent prediction. Next, based on the reference prediction quantities respectively corresponding to multiple alternative prediction layers in the Transformer model, a target prediction layer whose reference prediction quantity matches the target prediction quantity is determined from the multiple alternative prediction layers. That is, multiple alternative prediction layers suitable for different prediction lengths are integrated in the Transformer model, and a suitable target prediction layer is flexibly selected according to the specific target prediction quantity. Next, through the target prediction layer, prediction is performed based on the feature vectors of multiple reference observations, and thus the target observation value of the target prediction quantity can be predicted. This solution realizes time series prediction of different time lengths through a single Transformer model including multiple alternative prediction layers, and can flexibly adapt to short-term, medium-term, and long-term prediction requirements. Therefore, only this one model needs to be trained and maintained, greatly reducing the cost of managing and maintaining the model.

[0114] Based on the above first embodiment, the second embodiment of this application is proposed. For the same or similar content as the first embodiment, reference can be made to the above introduction and will not be repeated hereinafter. Referring to Figure 3 In the second embodiment, step S30 is refined into step S301:

[0115] Step S301: Based on the reference prediction quantities respectively corresponding to multiple alternative prediction layers in the Transformer model, a target set is determined. The target set includes at least one target prediction layer determined from the multiple alternative prediction layers and the quantities of the determined various target prediction layers, and the sum of the reference prediction quantities corresponding to all the target prediction layers in the target set is equal to the target prediction quantity.

[0116] The target set is a set composed of target prediction layers selected from multiple alternative prediction layers. This set not only contains at least one selected target prediction layer, but also contains the usage quantity of each target prediction layer. The core requirement of the target set is that the sum of the reference prediction quantities corresponding to all target prediction layers in it is equal to the target prediction quantity, so as to ensure that the prediction task of the specified length can be completed. That is, the quantity of each target prediction layer in the target set multiplied by the reference prediction quantity corresponding to each target prediction layer is equal to the total prediction quantity of each target prediction layer. Adding up the total prediction quantities of each target prediction layer in the target set is equal to the target prediction quantity.

[0117] Optionally, a greedy scheduling algorithm is used to screen the target prediction layers. Correspondingly, based on the reference prediction quantities respectively corresponding to multiple alternative prediction layers in the Transformer model, a target set is determined, including: initializing the predicted quantity to 0, and looping through the prediction layer screening process to update the predicted quantity until the predicted quantity reaches the target prediction quantity, and forming a target set with the selected various target prediction layers and the quantities of the selected various target prediction layers.

[0118] Among them, the prediction layer screening process includes: determining the judgment parameter of each alternative prediction layer in turn according to the order of the corresponding reference prediction quantity from large to small, and comparing the judgment parameter with the target prediction quantity. Among them, the judgment parameter is the sum of the reference prediction quantity and the predicted quantity. If the judgment parameter of the currently judged alternative prediction layer is greater than the target prediction quantity, then judge the next alternative prediction layer with the second largest reference prediction quantity; if the judgment parameter of the currently judged alternative prediction layer is less than or equal to the target prediction quantity, then determine the currently judged alternative prediction layer as an available target prediction layer, and update the predicted quantity to the sum of the reference prediction quantities corresponding to all current target prediction layers.

[0119] The predicted quantity is the sum of the reference prediction quantities corresponding to the target prediction layers that have been determined to be available for prediction during the process of screening the target prediction layers. It is initially set to 0, and as the screening process progresses, it will be continuously updated to reflect the prediction length that has been determined to be completed currently.

[0120] The prediction layer screening process is a series of operation steps executed to determine appropriate target prediction layers from multiple alternative prediction layers. This process gradually screens out the target prediction layers that can be combined to meet the target prediction quantity through continuous comparison and judgment.

[0121] The judgment parameter is a parameter calculated for each alternative prediction layer during the prediction layer screening process to determine whether the prediction layer is available. Its calculation method is to add the reference prediction quantity of the current alternative prediction layer to the predicted quantity. By comparing this judgment parameter with the target prediction quantity, it is decided whether to select the alternative prediction layer.

[0122] Suppose the Transformer model integrates 6 alternative prediction layers, namely the 1-step prediction layer, 4-step prediction layer, 8-step prediction layer, 16-step prediction layer, 32-step prediction layer, and 64-step prediction layer. The corresponding reference prediction quantities of the 6 alternative prediction layers are 1, 4, 8, 16, 32, and 64 respectively. The target prediction quantity Ht of the prediction task is 40. The prediction layer screening process includes:

[0123] (1) Perform data initialization:

[0124] The predicted quantity H = 0, indicating that no prediction has been made at the beginning.

[0125] The target set J = {}, indicating that no alternative prediction layer has been selected at the beginning.

[0126] (2) Perform loop prediction until H >= Ht, that is, the predicted quantity reaches the target prediction quantity.

[0127] The first loop:

[0128] Start trying from the alternative prediction layer with the largest corresponding reference prediction quantity. First, try the 64-step prediction layer.

[0129] Since H + 64 > Ht (0 + 64 > 40), the 64-step prediction layer is not applicable.

[0130] Try the 32-step prediction layer. Since H + 32 <= Ht (0 + 32 <= 40), the 32-step prediction layer is applicable.

[0131] Select the 32-step prediction layer and perform data update: H = 32, J = {32-step prediction layer}.

[0132] Since H < Ht (32 < 40), enter the second loop:

[0133] Start trying from the alternative prediction layer with the largest corresponding reference prediction quantity. First, try the 64-step prediction layer.

[0134] Since H + 64 > Ht (32 + 64 > 40), the 64-step prediction layer is not applicable.

[0135] Try the 32-step prediction layer. Since H + 32 > Ht (32 + 32 > 40), the 32-step prediction layer is not applicable.

[0136] Attempt the 16-step prediction layer. Since H + 16 > Ht (32 + 16 > 40), the 16-step prediction layer is not applicable.

[0137] Attempt the 8-step prediction layer. Since H + 8 <= Ht (32 + 8 <= 40), the 8-step prediction layer is applicable.

[0138] Select the 8-step prediction layer and perform data update: H = 40, J = {32-step prediction layer, 8-step prediction layer}.

[0139] Since H >= Ht (40 >= 4), the loop ends.

[0140] The finally selected target set J = {32-step prediction layer, 8-step prediction layer}.

[0141] Assume the target prediction quantity Ht = 70, then the target set J determined according to the above method = {64-step prediction layer, 4-step prediction layer, 1-step prediction layer, 1-step prediction layer}. Assume the target prediction quantity Ht = 10, then the target set J determined according to the above method = {8-step prediction layer, 1-step prediction layer, 1-step prediction layer}.

[0142] In the embodiment of the present application, screening is performed in the order from largest to smallest of the reference prediction quantities, and preference is given to the alternative prediction layers with larger reference prediction quantities. This can enable the sum of the reference prediction quantities of the screened target prediction layers to approach the target prediction quantity faster, thereby quickly finding a suitable combination of target prediction layers, reducing unnecessary comparison and screening times, and improving the screening efficiency.

[0143] In the embodiment of the present application, a target set is determined based on the reference prediction quantities respectively corresponding to multiple alternative prediction layers, and subsequently, all the target prediction layers in the target set jointly complete the data prediction task. This means that multiple alternative prediction layers can be flexibly combined according to different target prediction quantities, so that the model can adapt to various prediction requirements of different lengths. Whether it is short-term prediction or long-term prediction, it can be achieved by reasonably selecting the target prediction layers and their quantities, greatly improving the versatility and adaptability of the model. And this solution avoids training a separate model for each possible prediction length, thereby reducing the hardware resources and computing costs required for training and storing multiple models.

[0144] Based on the above second embodiment of the present application, the third embodiment of the present application is proposed. For the same or similar content as the second embodiment, reference can be made to the above introduction and will not be elaborated hereinafter. Refer to Figure 4 , in the third embodiment, step S40 includes step S401 and step S402.

[0145] Step S401, determine the call order of the target prediction layers in the target set.

[0146] The call order is the execution order determined for the target prediction layers in the target set. In this order, the target prediction layers are called sequentially for prediction operations to ensure the orderly progress of the prediction process.

[0147] Optionally, determining the call order of the target prediction layers in the target set includes: sorting various target prediction layers in the target set in descending order according to the corresponding reference prediction quantity; in the case of multiple target prediction layers of the same type, the call order of the multiple target prediction layers of the same type is randomly arranged.

[0148] Target prediction layers of the same type are target prediction layers in the target set that have the same reference prediction quantity. For example, if there are two prediction layers in the target set and their reference prediction quantities are both 5, then these two prediction layers belong to the target prediction layers of the same type.

[0149] In the embodiments of the present application, the target prediction layers are sorted in descending order according to the reference prediction quantity, and the prediction layer that can predict more observed values is preferentially used. This can approach the target prediction quantity faster during the prediction process.

[0150] Step S402, according to the call order, sequentially pass through the target prediction layers in the target set, and perform prediction based on the feature vectors of multiple reference observed values in the current reference time series to obtain target observed values corresponding to the reference prediction quantity of the target prediction layer, and add the target observed values of the reference prediction quantity after the multiple reference observed values to obtain an updated reference time series until all target prediction layers in the target set participate in the prediction to obtain target observed values of the target prediction quantity.

[0151] The reference time series is a series composed of multiple reference observed values arranged in chronological order. It is the starting point of prediction. As the prediction process progresses, the reference time series will be continuously updated to incorporate newly predicted target observed values.

[0152] The updated reference time series is a new reference time series formed by adding the target observed values corresponding to the reference prediction quantity of the target prediction layer obtained after each prediction through the target prediction layer to the original multiple reference observed values. It is used for the prediction of subsequent target prediction layers, enabling the model to make more accurate predictions using the latest information.

[0153] Exemplarily, each time the reference time series is updated based on the target observed values predicted by the target prediction layer, the feature vectors of multiple reference observed values in the updated reference time series are obtained through the attention layer in the Transformer model, and then the next target prediction layer performs prediction based on the multiple reference observed values in the updated reference time series until all target prediction layers in the target set participate in the prediction to obtain target observed values of the target prediction quantity.

[0154] Exemplary, in order to further improve the prediction efficiency, after receiving the input reference time series, each target prediction layer can generate multiple target observations that match the number of its reference predictions at once. This means that when the model performs a single prediction operation, it can batch output multiple prediction results instead of predicting each target observation one by one. This design significantly speeds up the prediction speed. Especially when dealing with large-scale time series data, it can greatly reduce the calculation time and improve the overall performance.

[0155] In the embodiment of the present application, by sequentially calling the target prediction layers in the target set and updating the reference time series after each prediction, the model can use the latest prediction results as the input for subsequent predictions. This allows the model to continuously consider new information during the prediction process, capture the dynamic changes in the data sequence, and thus improve the prediction accuracy of the target observations. Moreover, by determining the call order and gradually updating the reference time series, this solution can complete the prediction in an orderly manner even in the face of complex prediction scenarios, ensuring the integrity and coherence of the prediction results.

[0156] Based on the first embodiment of the present application described above, the fourth embodiment of the present application is proposed. The same or similar content as the first embodiment can be referred to the above introduction and will not be repeated hereinafter. Referring to Figure 5 , in the fourth embodiment, before step S10, steps S01 to S04 are included.

[0157] Step S01, obtain a sample time series, which includes a plurality of sample reference observations and a plurality of sample target observations arranged in chronological order.

[0158] The sample time series is the time series data used to train the Transformer model. It consists of a plurality of sample reference observations and a plurality of sample target observations arranged in chronological order. The sample reference observations are historical data and serve as the basis for model input. The sample target observations are the future true values corresponding to the sample reference observations and are used to evaluate the prediction accuracy of the model.

[0159] Optionally, obtaining the sample time series includes: obtaining a sample sequence set, where the sample sequence set contains a plurality of sample time series and the number of observations included in the plurality of sample time series is different; sequentially obtaining the sample time series from the sample sequence set in ascending order of the number of included observations.

[0160] A set of sample sequences is a set composed of multiple sample time series. These sample time series can come from different time periods, different data sources, or different scenarios. The number of observations contained in each sample time series in the set may be different, which reflects the diversity and complexity of data in practical applications.

[0161] In the embodiments of the present application, the sample time series are obtained in ascending order of the number of observations for training. The model can start learning from simple and short-sequence data and gradually adapt to more complex and long-sequence data. This progressive learning method helps the model better understand the characteristics and laws of the data and avoid problems such as unstable training or difficult convergence caused by overly complex data at the initial stage of training. If long-sequence data containing a large number of observations is used for training at the beginning, the model may fall into a local optimal solution because it is difficult to capture the long-term dependence relationships in the data; while starting from short-sequence data for training, the model can first master some basic patterns and laws and then gradually process more complex long-sequence data, thereby improving the stability and convergence speed of training.

[0162] Step S02: Obtain the feature vectors of multiple sample reference observations through the attention layer.

[0163] Step S03: Through multiple alternative prediction layers, perform predictions respectively based on the feature vector to obtain sample predicted observations corresponding to the reference prediction quantities of the respective alternative prediction layers.

[0164] The sample predicted observations are the results obtained by each alternative prediction layer making predictions based on the feature vector and are the model's estimates of future observations. Each alternative prediction layer will obtain sample predicted observations corresponding to its reference prediction quantity.

[0165] Exemplarily, the sample predicted observations of various alternative prediction layers are determined through the following relational expression (1).

[0166] Relational expression (1):

[0167] where t represents the number of sample reference observations in the sample time series, h t represents the feature vector of multiple sample reference observations input to each alternative prediction layer, f θ represents the j-th alternative prediction layer, θ is the learnable parameter in the j-th alternative prediction layer, p j represents the reference prediction quantity corresponding to the j-th alternative prediction layer, represents the sequence composed of the sample predicted observations predicted by the j-th alternative prediction layer.

[0168] Exemplarily, assume that there are three alternative prediction layers in the model, namely alternative prediction layer A, alternative prediction layer B, and alternative prediction layer C. The corresponding reference prediction quantities are 1, 4, and 8 respectively. The prediction task is to predict the power load in different future time periods. The feature vector is a vector with a length of 128, which contains the key information of the power load in the past period. Then in step S03, based on this vector with a length of 128, prediction is performed through alternative prediction layer A, and 1 sample prediction observation value is obtained, that is, the predicted value of the power load in the next 1 hour. Based on this vector with a length of 128, prediction is performed through alternative prediction layer B, and 4 sample prediction observation values are obtained, corresponding to the predicted values of the power load in the 1st hour, 2nd hour, 3rd hour, and 4th hour in the future respectively. Based on this vector with a length of 128, prediction is performed through alternative prediction layer C, and 8 sample prediction observation values are obtained, corresponding to the predicted values of the power load from the 1st hour to the 8th hour in the future respectively.

[0169] Step S04: Train the Transformer model based on each sample prediction observation value and each sample target observation value.

[0170] Training the Transformer model based on each sample prediction observation value and each sample target observation value means: adjusting the model parameters to reduce the difference between each sample prediction observation value and each sample target observation value, so as to improve the prediction accuracy of the model.

[0171] In the embodiment of the present application, through multiple alternative prediction layers, predictions are respectively performed based on the feature vector, and sample prediction observation values corresponding to the reference prediction quantities of various alternative prediction layers are obtained. Based on each sample prediction observation value and each sample target observation value, the Transformer model is trained. This means that multiple alternative prediction layers are trained synchronously, which enables the model to learn prediction patterns of different lengths. Subsequently, in actual applications, no matter what the target prediction quantity is, the model can select an appropriate prediction layer from multiple alternative prediction layers for prediction, so as to adapt to different prediction requirements and improve the versatility and flexibility of the model.

[0172] Based on the fourth embodiment of the present application above, the fifth embodiment of the present application is proposed. For the same or similar content as the fourth embodiment, reference can be made to the above introduction and will not be repeated hereinafter. Refer to Figure 6 , in the fifth embodiment, step S04 includes steps S041 to S043.

[0173] Step S041: Determine the prediction loss of each alternative prediction layer based on each sample prediction observation value corresponding to each alternative prediction layer and the sample target observation value corresponding to each sample prediction observation value.

[0174] The prediction loss is used to measure the degree of difference between the prediction results of each alternative prediction layer and the true values. By comparing the sample prediction observations of a certain alternative prediction layer with the corresponding sample target observations, using specific loss functions, such as mean squared error loss function, cross-entropy loss function, etc., the prediction loss can be calculated. The smaller the prediction loss, the closer the prediction result of the alternative prediction layer is to the true value, and the stronger the prediction ability.

[0175] Optionally, based on the sample prediction observations corresponding to various alternative prediction layers and the sample target observations corresponding to the sample prediction observations, determine the prediction losses of the alternative prediction layers, including: for any sample prediction observation corresponding to any alternative prediction layer, determine the first prediction loss based on the difference between the sample prediction observation and the sample target observation corresponding to the sample prediction observation; divide the first prediction loss by the number of repeated predictions, where the number of repeated predictions is the number of times the sample target observation is repeatedly predicted by multiple alternative prediction layers in the Transformer model; add up the second prediction losses corresponding to all sample prediction observations of any alternative prediction layer to obtain the prediction loss of the alternative prediction layer.

[0176] The first prediction loss is the loss obtained by calculating the difference between any sample prediction observation corresponding to a certain alternative prediction layer and its corresponding sample target observation. Loss functions such as mean squared error and absolute error can be used to calculate it, which is used to preliminarily evaluate the deviation between a single predicted value and the true value.

[0177] The number of repeated predictions is the number of times the sample target observation is repeatedly predicted by multiple alternative prediction layers in the Transformer model. Since different alternative prediction layers are responsible for predicting different numbers of observations, it is possible that they will all predict the same sample target observation, so there will be a situation of repeated prediction.

[0178] For example, assume there are two types of alternative prediction layers in the model: the first alternative prediction layer has a reference prediction quantity of 3, that is, it predicts the observations of the next 3 time steps. The second alternative prediction layer has a reference prediction quantity of 2, that is, it predicts the observations of the next 2 time steps. If the sample target observation is the observation of the first time step in the future or the observation of the second time step in the future, it will be predicted by both alternative prediction layers at the same time, so it is repeatedly predicted twice.

[0179] The second prediction loss is the result obtained by dividing the first prediction loss by the number of repeated predictions. By this way, the first prediction loss is adjusted to balance the influence brought by repeated prediction.

[0180] In the embodiments of the present application, multiple alternative prediction layers in the model may predict the same sample target observation value multiple times. If the differences between all sample predicted observations and the sample target observations are directly accumulated to calculate the loss, the influence of some sample target observations will be over-amplified, resulting in bias in model training. By dividing the first prediction loss by the number of repeated predictions to obtain the second prediction loss, the influence brought by repeated predictions can be effectively eliminated, enabling each sample target observation value to have a reasonable weight in loss calculation, and thus calculating a more reasonable prediction loss. This enables the model to more accurately learn the features and patterns in the data when adjusting the parameters based on these losses, thereby improving the prediction accuracy.

[0181] Step S042: Determine the average prediction loss of multiple alternative prediction layers as the model prediction loss.

[0182] The model prediction loss is the loss value obtained by averaging the prediction losses of multiple alternative prediction layers. It comprehensively reflects the prediction performance of the entire Transformer model in the current training stage. The model prediction loss is the basis for adjusting the model parameters, and the training objective of the model is to continuously reduce this loss value to improve the prediction accuracy of the model.

[0183] Exemplarily, the model prediction loss is determined through the following relational expression (2).

[0184] Relational expression (2):

[0185] Among them, loss represents the model prediction loss, j represents the serial number of the alternative prediction layer, p represents the total number of alternative prediction layers in the model, represents the sequence composed of the sample predicted observations predicted by the j-th alternative prediction layer, is the sequence composed of the sample target observations corresponding to the sample predicted observations predicted by the j-th alternative prediction layer, is the prediction loss of the j-th alternative prediction layer.

[0186] Step S043: Adjust the model parameters based on the model prediction loss to reduce the model prediction loss.

[0187] The model parameters are the learnable parameters in the Transformer model, including the weights and biases of the attention layer, as well as the weights and biases of each alternative prediction layer, etc. These parameters determine the specific structure and function of the model. During the training process, these parameters are adjusted according to the model prediction loss through the backpropagation algorithm, enabling the model to continuously learn and optimize to better adapt to the training data and prediction tasks.

[0188] In the embodiments of the present application, since different alternative prediction layers may perform differently on prediction tasks of different lengths, considering only the loss of a single prediction layer cannot accurately reflect the true ability of the model. Therefore, the average prediction loss of multiple alternative prediction layers is calculated as the model prediction loss, which can comprehensively and integrally evaluate the overall performance of the Transformer model. During the training process, the model parameters are adjusted based on the model prediction loss, enabling the model to consider the prediction effects of multiple alternative prediction layers simultaneously. This allows the model to learn more general features and patterns, thereby improving the generalization ability of the model on prediction tasks of different lengths.

[0189] Figure 7 is a schematic structural diagram of a Transformer model provided by the present application. Refer to Figure 7 In this model, multiple alternative prediction layers are integrated. Each alternative prediction layer corresponds to a different reference prediction quantity, and each alternative prediction layer is connected to the attention layer. After inputting the reference time series into the Transformer model, the feature vector corresponding to the reference time series is obtained through the attention layer in the Transformer model, and then the alternative prediction layer can be flexibly selected according to the target prediction quantity to complete the time series prediction task. It can be understood that compared with the traditional solution, that is, training multiple models to separately implement time series predictions of different lengths. The present application proposes an innovative solution, which realizes wide support for various time series prediction tasks of different lengths by sharing the core structure of the Transformer model, such as the attention layer. This design can not only achieve a higher model capacity and prediction accuracy with fewer computing resources, but also significantly improve the flexibility and scalability of the model. Specifically, the model structure can be easily extended to larger datasets and wider model scales to adapt to the requirements of different time series features. In addition, this unified architecture greatly reduces the model management and maintenance costs, providing convenience for practical applications.

[0190] Another point to note is that the above examples are only for understanding the present application and do not constitute a limitation on the data prediction method of the present application. Any simple transformation in more forms based on this technical concept is within the protection scope of the present application.

[0191] The present application also provides a data prediction device. Please refer to Figure 8 The data prediction device includes:

[0192] A prediction instruction response module 10, configured to extract a reference time series and a target prediction quantity from the data prediction instruction in response to the data prediction instruction. The reference time series includes multiple reference observations arranged in chronological order, and the target prediction quantity represents the number of target observations to be predicted after the reference time series;

[0193] The feature vector acquisition module 20 is used to obtain the feature vectors of multiple reference observations through the attention layer in the Transformer model;

[0194] The prediction layer selection module 30 is used to determine the target prediction layer with the reference prediction quantity matching the target prediction quantity from multiple alternative prediction layers based on the reference prediction quantities respectively corresponding to the multiple alternative prediction layers in the Transformer model;

[0195] The observation prediction module 40 is used to predict based on the feature vectors of multiple reference observations through the target prediction layer to obtain the target observations of the target prediction quantity.

[0196] Optionally, the prediction layer selection module 30 is used to determine the target set based on the reference prediction quantities respectively corresponding to the multiple alternative prediction layers in the Transformer model. The target set includes at least one target prediction layer determined from the multiple alternative prediction layers and the quantities of the determined various target prediction layers, and the sum of the reference prediction quantities corresponding to all the target prediction layers in the target set is equal to the target prediction quantity.

[0197] Optionally, the prediction layer selection module 30 is used to initialize the predicted quantity to 0 and repeatedly execute the prediction layer screening process to update the predicted quantity until the predicted quantity reaches the target prediction quantity, and form the target set with the selected various target prediction layers and the quantities of the selected various target prediction layers. The prediction layer screening process includes:

[0198] Determine the judgment parameter of each alternative prediction layer in turn according to the corresponding reference prediction quantity from large to small, and compare the judgment parameter with the target prediction quantity. The judgment parameter is the sum of the reference prediction quantity and the predicted quantity;

[0199] If the judgment parameter of the currently judged alternative prediction layer is greater than the target prediction quantity, then judge the alternative prediction layer with the second largest reference prediction quantity;

[0200] If the judgment parameter of the currently judged alternative prediction layer is less than or equal to the target prediction quantity, then determine the currently judged alternative prediction layer as the available target prediction layer, and update the predicted quantity to the sum of the reference prediction quantities corresponding to all the current target prediction layers.

[0201] Optionally, the observation prediction module 40 includes:

[0202] The sequential determination unit is used to determine the call order of the target prediction layers in the target set;

[0203] The observation value prediction unit is used to sequentially pass through the target prediction layers in the target set according to the call order, make predictions based on the feature vectors of multiple reference observation values in the current reference time series, obtain the target observation values corresponding to the reference prediction quantity of the target prediction layer, and add the target observation values of the reference prediction quantity after the multiple reference observation values to obtain an updated reference time series until all the target prediction layers in the target set participate in the prediction to obtain the target observation values of the target prediction quantity.

[0204] Optionally, the order determination unit is used to sort the various target prediction layers in the target set in descending order according to the corresponding reference prediction quantity; in the case where there are multiple target prediction layers of the same type, the call order of the multiple target prediction layers of the same type is randomly arranged.

[0205] Optionally, the device further includes:

[0206] The sample acquisition module is used to acquire a sample time series, and the sample time series includes multiple sample reference observation values and multiple sample target observation values arranged in chronological order;

[0207] The feature vector acquisition module 20 is further used to obtain the feature vectors of multiple sample reference observation values through the attention layer;

[0208] The observation value prediction module 40 is further used to make predictions respectively based on the feature vectors through various alternative prediction layers to obtain the sample prediction observation values corresponding to the reference prediction quantity of each alternative prediction layer;

[0209] The model training module is used to train the Transformer model based on each sample prediction observation value and each sample target observation value.

[0210] Optionally, the model training module includes:

[0211] The layer loss determination unit is used to determine the prediction loss of each alternative prediction layer based on the sample prediction observation values corresponding to each alternative prediction layer and the sample target observation values corresponding to each sample prediction observation value;

[0212] The model loss determination unit is used to determine the average prediction loss of multiple alternative prediction layers as the model prediction loss;

[0213] The parameter adjustment unit is used to adjust the model parameters based on the model prediction loss to reduce the model prediction loss.

[0214] Optionally, a layer loss determination unit is configured to, for any sample prediction observation value corresponding to any alternative prediction layer, determine a first prediction loss based on the difference between the sample prediction observation value and the sample target observation value corresponding to the sample prediction observation value; divide the first prediction loss by the number of repeated predictions, where the number of repeated predictions is the number of times the sample target observation value is repeatedly predicted by multiple alternative prediction layers in the Transformer model, to obtain a second prediction loss; and add up the second prediction losses corresponding to all sample prediction observation values corresponding to any alternative prediction layer to obtain the prediction loss of the alternative prediction layer.

[0215] Optionally, a sample acquisition module is configured to acquire a sample sequence set, where the sample sequence set includes multiple sample time series, and the number of observation values included in the multiple sample time series is different; and sequentially acquire the sample time series from the sample sequence set in the order of increasing number of included observation values.

[0216] Optionally, the reference prediction quantities corresponding to multiple alternative prediction layers in the Transformer model are 1, 4, 8, 16, 32, and 64 respectively.

[0217] The data prediction device provided by the present application adopts the data prediction method in the above embodiment, and can solve the technical problem in the related art that training multiple models for time series prediction of different time lengths results in high management and maintenance costs of the models. Compared with the prior art, the beneficial effects of the data prediction device provided by the present application are the same as those of the data prediction method provided by the above embodiment, and other technical features in the data prediction device are the same as those disclosed in the method of the above embodiment, and will not be elaborated here.

[0218] The present application provides a data prediction device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the data prediction method in the first embodiment above.

[0219] Next, refer to Figure 9, which shows a schematic structural diagram of a data prediction device suitable for implementing the embodiments of the present application. The data prediction device in the embodiments of the present application may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistant), PADs (Portable Application Description), PMPs (Portable Media Player), in-vehicle terminals (such as in-vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 9 The shown data prediction device is merely an example and should not impose any limitation on the functions and usage scope of the embodiments of the present application.

[0220] As Figure 9 shown, the data prediction device may include a processing device 1001 (such as a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM: Read Only Memory) 1002 or the program loaded from the storage device 1003 into the random access memory (RAM: Random Access Memory) 1004. In the RAM 1004, various programs and data required for the operation of the data prediction device are also stored. The processing device 1001, the ROM 1002, and the RAM 1004 are connected to each other through a bus 1005. The input / output (I / O) interface 1006 is also connected to the bus. Generally, the following systems may be connected to the I / O interface 1006: an input device 1007 including, for example, a touch screen, a touch pad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output device 1008 including, for example, a liquid crystal display (LCD: Liquid Crystal Display), a speaker, a vibrator, etc.; a storage device 1003 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1009. The communication device 1009 can allow the data prediction device to communicate with other devices wirelessly or wiredly to exchange data. Although the figure shows a data prediction device with various systems, it should be understood that it is not required to implement or have all the shown systems. Instead, more or fewer systems may be implemented or had.

[0221] In particular, according to the embodiments disclosed in the present application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments disclosed in the present application include a computer program product that includes a computer program carried on a computer-readable medium, and the computer program contains program codes for executing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from a network through a communication device, or installed from a storage device 1003, or installed from a ROM 1002. When the computer program is executed by a processing device 1001, the above-mentioned functions defined in the methods of the embodiments disclosed in the present application are executed.

[0222] The data prediction device provided in the present application adopts the data prediction method in the above embodiment, and can solve the technical problem in the related art that training multiple models for time series prediction of different time lengths results in high management and maintenance costs of the models. Compared with the prior art, the beneficial effects of the data prediction device provided in the present application are the same as those of the data prediction method provided in the above embodiment, and other technical features in the data prediction device are the same as the features disclosed in the method of the previous embodiment, and will not be described in detail here.

[0223] It should be understood that each part disclosed in the present application can be implemented by hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in a suitable manner in any one or more embodiments or examples.

[0224] The above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed in the present application, and all of them should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

[0225] The present application provides a computer-readable storage medium having computer-readable program instructions (i.e., computer programs) stored thereon, and the computer-readable program instructions are used to execute the data prediction method in the above embodiment.

[0226] The computer-readable storage medium provided by this application can be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or components, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections with one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM) or flash memory, optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In this embodiment, the computer-readable storage medium can be any tangible medium that contains or stores a program, which can be used by or in conjunction with an instruction execution system, device, or component. The program code contained on the computer-readable storage medium can be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination of the above.

[0227] The above computer-readable storage medium can be included in the data prediction device; it can also exist independently without being assembled into the data prediction device.

[0228] The above computer-readable storage medium carries one or more programs. When the one or more programs are executed by the data prediction device, the data prediction device is caused to: in response to a data prediction instruction, extract a reference time series and a target prediction quantity from the data prediction instruction, where the reference time series includes a plurality of reference observations arranged in chronological order, and the target prediction quantity represents the number of target observations to be predicted after the reference time series; obtain the feature vectors of the plurality of reference observations through the attention layer in the Transformer model; based on the reference prediction quantities respectively corresponding to multiple alternative prediction layers in the Transformer model, determine a target prediction layer whose reference prediction quantity matches the target prediction quantity from the multiple alternative prediction layers; and through the target prediction layer, perform a prediction based on the feature vectors of the plurality of reference observations to obtain the target observations of the target prediction quantity.

[0229] Computer program code for performing the operations of this application can be written in one or more programming languages or combinations thereof. The above-mentioned programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any kind of network, including a local area network (LAN: Local Area Network) or a wide area network (WAN: Wide Area Network), or it can be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).

[0230] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutively represented blocks may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, as well as the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0231] The modules described in the embodiments of this application can be implemented in software or in hardware. Among them, the name of the module does not constitute a limitation to the unit itself in some cases.

[0232] The readable storage medium provided by this application is a computer-readable storage medium. The computer-readable storage medium stores computer-readable program instructions (i.e., computer programs) for performing the above-mentioned data prediction method, and can solve the technical problem in the related art that training multiple models for time series prediction with different time lengths results in high model management and maintenance costs. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided by this application are the same as those of the data prediction method provided by the above embodiments, and will not be elaborated here.

[0233] The present application also provides a computer program product, including a computer program which, when executed by a processor, implements the steps of the data prediction method as described above.

[0234] The computer program product provided by the present application can solve the technical problem in the related art that training multiple models for time series prediction of different time lengths results in high costs for model management and maintenance. Compared with the prior art, the beneficial effects of the computer program product provided by the present application are the same as those of the data prediction method provided in the above embodiments, and will not be elaborated here.

[0235] The foregoing are only partial embodiments of the present application, and thus do not limit the patent scope of the present application. Any equivalent structural transformation made by using the content of the specification and drawings of the present application under the technical concept of the present application, or any direct / indirect application in other related technical fields, is included in the patent protection scope of the present application.

Claims

1. A data prediction method, characterized in that, The method includes: In response to a data prediction instruction, extracting a reference time series and a target prediction quantity from the data prediction instruction, where the reference time series includes a plurality of reference observations arranged in chronological order, and the target prediction quantity represents the number of target observations to be predicted after the reference time series; Obtaining feature vectors of the plurality of reference observations through an attention layer in a Transformer model; Based on the reference prediction quantities respectively corresponding to multiple alternative prediction layers in the Transformer model, determining a target prediction layer whose reference prediction quantity matches the target prediction quantity from the multiple alternative prediction layers; Through the target prediction layer, predicting based on the feature vectors of the plurality of reference observations to obtain target observations of the target prediction quantity.

2. The method according to claim 1, characterized in that, The determining, based on the reference prediction quantities respectively corresponding to multiple alternative prediction layers in the Transformer model, a target prediction layer whose reference prediction quantity matches the target prediction quantity from the multiple alternative prediction layers includes: Based on the reference prediction quantities respectively corresponding to multiple alternative prediction layers in the Transformer model, determining a target set, where the target set includes at least one target prediction layer determined from the multiple alternative prediction layers and the quantities of the determined various target prediction layers, and the sum of the reference prediction quantities corresponding to all target prediction layers in the target set is equal to the target prediction quantity.

3. The method according to claim 2, characterized in that, The determining, based on the reference prediction quantities respectively corresponding to multiple alternative prediction layers in the Transformer model, a target set includes: Initializing the predicted quantity to 0, and repeatedly executing a prediction layer screening process to update the predicted quantity until the predicted quantity reaches the target prediction quantity, and forming the target set with the screened various target prediction layers and the quantities of the screened various target prediction layers, where the prediction layer screening process includes: Sequentially determining a judgment parameter for each alternative prediction layer in descending order of the corresponding reference prediction quantity, and comparing the judgment parameter with the target prediction quantity, where the judgment parameter is the sum of the reference prediction quantity and the predicted quantity; If the judgment parameter of the currently judged alternative prediction layer is greater than the target prediction quantity, then judging the next alternative prediction layer with the second largest reference prediction quantity; If the judgment parameter of the currently judged alternative prediction layer is less than or equal to the target prediction quantity, then determining the currently judged alternative prediction layer as an available target prediction layer, and updating the predicted quantity to the sum of the reference prediction quantities corresponding to all current target prediction layers.

4. The method according to claim 2, characterized in that, The obtaining, through the target prediction layer, target observations of the target prediction quantity based on the feature vectors of the plurality of reference observations includes: Determining the call order of the target prediction layers in the target set; In the described calling order, sequentially pass through the target prediction layers in the target set, make predictions based on the feature vectors of multiple reference observations in the current reference time series, obtain target observations of the reference prediction quantity corresponding to the target prediction layer, and add the target observations of the reference prediction quantity after the multiple reference observations to obtain an updated reference time series until all the target prediction layers in the target set participate in the prediction to obtain the target observations of the target prediction quantity.

5. The method according to claim 4, characterized in that, Determining the calling order of the target prediction layers in the target set includes: Sorting various target prediction layers in the target set in descending order according to the corresponding reference prediction quantity; In the case where there are multiple target prediction layers of the same type, the calling order of the multiple target prediction layers of the same type is randomly arranged.

6. The method according to claim 1, characterized in that, The training process of the Transformer model includes: Obtaining a sample time series, which includes multiple sample reference observations and multiple sample target observations arranged in chronological order; Obtaining the feature vectors of the multiple sample reference observations through the attention layer; Making predictions respectively based on the feature vectors through the multiple alternative prediction layers to obtain sample predicted observations of the reference prediction quantity corresponding to each alternative prediction layer; Training the Transformer model based on each sample predicted observation and each sample target observation.

7. A data prediction device, characterized in that The device includes: A prediction instruction response module, configured to, in response to a data prediction instruction, extract a reference time series and a target prediction quantity from the data prediction instruction, where the reference time series includes multiple reference observations arranged in chronological order, and the target prediction quantity represents the quantity of target observations to be predicted after the reference time series; A feature vector acquisition module, configured to obtain the feature vectors of the multiple reference observations through the attention layer in the Transformer model; A prediction layer selection module, configured to determine a target prediction layer whose reference prediction quantity matches the target prediction quantity from the multiple alternative prediction layers based on the reference prediction quantities corresponding to the multiple alternative prediction layers in the Transformer model; An observation prediction module, configured to make predictions based on the feature vectors of the multiple reference observations through the target prediction layer to obtain the target observations of the target prediction quantity.

8. A data prediction device, characterized in that, The device includes: a memory, a processor, and a computer program stored on the memory and executable on the processor, where the computer program is configured to implement the steps of the data prediction method according to any one of claims 1 to 6.

9. A storage medium, characterized in that, The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium, and when the computer program is executed by a processor, it implements the steps of the data prediction method according to any one of claims 1 to 6.

10. A computer program product, characterized in that, The computer program product includes a computer program, and when the computer program is executed by a processor, it implements the steps of the data prediction method according to any one of claims 1 to 6.