Data prediction method and device, equipment, storage medium and product

By extracting and updating the extreme value feature vectors in the Transformer model, and using weight control parameters to enhance attention weights, the problem of low prediction accuracy at extreme values ​​is solved, and higher prediction accuracy and capture ability of extreme values ​​are achieved.

CN120105055APending Publication Date: 2025-06-06BEIJING 360 INTELLIGENT TECHNOLOGY CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510157313.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-12
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

When processing time series data, the existing Transformer model is difficult to effectively learn and capture the complex patterns and dynamic laws behind extreme values, resulting in low prediction accuracy at extreme values.

Method used

By extracting the extreme eigenvectors in the reference time series and using weight control parameters to increase the attention weight between the extreme eigenvectors and other eigenvectors, the eigenvectors are updated to increase the model's attention to extreme values.

Benefits of technology

The prediction accuracy of the Transformer model at the extreme value is improved, so that the model can better learn and capture the characteristics and laws of the extreme value.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120105055A_ABST
    Figure CN120105055A_ABST
Patent Text Reader

Abstract

The invention discloses a data prediction method and device, equipment, a storage medium and a product, relates to the technical field of artificial intelligence, and discloses a method for extracting a reference time sequence from a data prediction instruction in response to the data prediction instruction, the reference time sequence comprising observation values corresponding to a plurality of time points, and the plurality of observation values comprising an extreme value; determining feature vectors of the plurality of observation values; through an attention layer in a Transform model, attention weights between each feature vector and all feature vectors are determined based on weight regulation and control parameters, each feature vector is updated based on the determined attention weights, and the weight regulation and control parameters are used for increasing the attention weight between each feature vector and the feature vector of the extreme value; and through a prediction layer in the Transform model, based on the updated feature vectors, predicting observation values after the plurality of time points. According to the method, the prediction precision of the Transform model at the extreme value can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and in particular to a data prediction method, device, equipment, storage medium and product. Background Art

[0002] In today's digital age, time series data is widely used in various fields, such as energy management, weather forecasting, financial market analysis, etc. Time series data is a sequence of observations arranged in chronological order. It contains the laws and characteristics of the evolution of things over time, and can help people understand and predict the development trends of various phenomena.

[0003] As an important sequence prediction model, the Transformer model aims to predict future values ​​based on time series data and provide strong support for decision-making. However, the existing Transformer model faces a significant challenge in practical applications: the extreme values ​​in time series data, such as peaks and troughs, are rare, and it is difficult for the model to fully learn and capture the complex patterns and dynamic laws behind these extreme values, resulting in generally low prediction accuracy at extreme values.

[0004] The above contents are only used to assist in understanding the technical solution of the present application and do not constitute an admission that the above contents are prior art. Summary of the invention

[0005] The main purpose of this application is to provide a data prediction method, device, equipment, storage medium and product that can improve the prediction accuracy of the Transformer model at extreme values.

[0006] In response to the data prediction instruction, extracting a reference time series from the data prediction instruction, the reference time series including observation values ​​corresponding to a plurality of time points, and the plurality of observation values ​​including an extreme value;

[0007] determining a feature vector of the plurality of observations;

[0008] Determine, through an attention layer in a transformer model, an attention weight between each feature vector and all feature vectors based on a weight control parameter, and update each feature vector based on the determined attention weight, wherein the weight control parameter is used to increase the attention weight between each feature vector and the feature vector of the extreme value;

[0009] The prediction layer in the Transformer model predicts the observation values ​​after the multiple time points based on the updated feature vector.

[0010] Optionally, determining the attention weight between each feature vector and all feature vectors based on a weight control parameter through an attention layer in a transformer model includes:

[0011] Determine the similarity between each feature vector and all feature vectors through the attention layer;

[0012] Through the attention layer, the similarity between each feature vector and all feature vectors is added to the weight control parameter between each feature vector and all feature vectors to obtain the attention weight between each feature vector and all feature vectors;

[0013] The weight control parameter between each eigenvector and the eigenvector of the extreme value is greater than the weight control parameter between each eigenvector and other eigenvectors.

[0014] Optionally, determining the similarity between each feature vector and all feature vectors through the attention layer includes:

[0015] Through the attention layer, each feature vector is converted into a query vector and a key vector respectively;

[0016] A dot product operation is performed on each query vector and all key vectors to obtain a similarity matrix, wherein each element (i, j) in the similarity matrix represents the similarity between the i-th query vector and the j-th key vector, wherein i and j are both natural numbers greater than 0.

[0017] Optionally, the similarity between each feature vector and all feature vectors is added to a weight control parameter between each feature vector and all feature vectors through the attention layer to obtain an attention weight between each feature vector and all feature vectors, including:

[0018] Obtain a weight control matrix, where the weight control matrix is ​​a binary matrix of the same size as the similarity matrix, each element (i, j) in the weight control matrix represents a weight control parameter between the i-th query vector and the j-th key vector, and when the j-th key vector corresponds to an extreme value, the element (i, j) is a natural number other than 0, otherwise the element (i, j) is 0;

[0019] Through the attention layer, the similarity matrix and the weight control matrix are added and converted into a probability distribution to obtain an attention matrix, wherein each element (i, j) in the attention matrix represents the attention weight between the i-th query vector and the j-th key vector.

[0020] Optionally, after predicting the observation values ​​after the plurality of time points based on the updated feature vectors through the prediction layer in the Transformer model, the method further comprises:

[0021] Determining a model prediction error based on a prediction result of the Transformer model;

[0022] Based on the model prediction error, the Transformer model is trained to reduce the model prediction error, wherein training the Transformer model includes adjusting the weight control parameter.

[0023] Optionally, determining the feature vectors of the multiple observations includes:

[0024] Generate feature vectors of the plurality of observations through an embedding layer in the Transformer model, and fuse the indication information into the feature vectors;

[0025] The indication information is used to indicate that the corresponding observation value is the extreme value or the background value, and the background value is the observation value other than the extreme value among the multiple observation values.

[0026] Optionally, generating feature vectors of the plurality of observations through an embedding layer in the Transformer model, and fusing indication information into the feature vectors, comprises:

[0027] Generate content embedding vectors and position codes of the plurality of observation values ​​through the embedding layer, and add the indication information to the position code of each observation value;

[0028] The position code of each observation value is concatenated with the content embedding vector to obtain a feature vector of each observation value.

[0029] Optionally, after extracting a reference time series from the data prediction instruction in response to the data prediction instruction, the method further comprises:

[0030] Placing a sliding window at the starting position of the reference time series so that the sliding window covers part of the observations in the reference time series;

[0031] Determine extreme values ​​in the sliding window, wherein the extreme values ​​in the sliding window include a maximum value, a minimum value, and a target number of observation values ​​around the maximum value and the minimum value in the sliding window;

[0032] After adjusting the position of the sliding window, the extreme values ​​in the sliding window are re-determined until all the extreme values ​​in the reference time series are determined.

[0033] Optionally, after extracting a reference time series from the data prediction instruction in response to the data prediction instruction, the method further comprises:

[0034] Based on a plurality of observations in the reference time series, determining a rate of change corresponding to each observation;

[0035] Filter observations whose corresponding rate of change has a different sign from the rate of change corresponding to the previous observation;

[0036] The screened observations and a target number of observations surrounding the observations are determined as the extreme values.

[0037] Optionally, after predicting the observation values ​​after the plurality of time points based on the updated feature vectors through the prediction layer in the Transformer model, the method further comprises:

[0038] Obtaining observation values ​​corresponding to a plurality of predicted time points after the plurality of time points, wherein the observation values ​​corresponding to the plurality of predicted time points include extreme values ​​and background values ​​other than the extreme values;

[0039] Determine an extreme value prediction error and a background value prediction error based on the observed values ​​corresponding to the multiple prediction time points and the predicted values ​​corresponding to the multiple prediction time points in the model prediction result;

[0040] Performing a weighted summation on the extreme value prediction error and the background value prediction error to obtain a model prediction error;

[0041] The Transformer model is trained based on the model prediction error.

[0042] Optionally, performing weighted summation on the extreme value prediction error and the background value prediction error to obtain a model prediction error includes:

[0043] Based on the size of the extreme value prediction error, the weight of the extreme value prediction error used in the previous round of training is updated, and the size of the extreme value prediction error is positively correlated with the size of the weight of the extreme value prediction error;

[0044] Based on the updated weights, the extreme value prediction error and the background value prediction error are weightedly summed to obtain the model prediction error.

[0045] In addition, to achieve the above-mentioned purpose, the present application also proposes a data prediction device, the device comprising:

[0046] An instruction response module, configured to respond to a data prediction instruction and extract a reference time series from the data prediction instruction, wherein the reference time series includes observation values ​​corresponding to a plurality of time points, and the plurality of observation values ​​include extreme values;

[0047] A vector determination module, used to determine the characteristic vectors of the plurality of observations;

[0048] A vector update module, for determining an attention weight between each feature vector and all feature vectors based on a weight control parameter through an attention layer in a transformer model, and updating each feature vector based on the determined attention weight, wherein the weight control parameter is used to increase the attention weight between each feature vector and the feature vector of the extreme value;

[0049] A data prediction module is used to predict the observation values ​​after the multiple time points based on the updated feature vectors through the prediction layer in the Transformer model.

[0050] Optionally, the vector updating module includes:

[0051] a similarity determination unit, configured to determine the similarity between each feature vector and all feature vectors through the attention layer;

[0052] an attention determination unit, configured to add the similarity between each feature vector and all feature vectors to a weight control parameter between each feature vector and all feature vectors through the attention layer to obtain an attention weight between each feature vector and all feature vectors;

[0053] The weight control parameter between each eigenvector and the eigenvector of the extreme value is greater than the weight control parameter between each eigenvector and other eigenvectors.

[0054] Optionally, the similarity determination unit is used to convert each feature vector into a query vector and a key vector respectively through the attention layer; perform a dot product operation on each query vector and all key vectors to obtain a similarity matrix, wherein each element (i, j) in the similarity matrix represents the similarity between the i-th query vector and the j-th key vector, where i and j are both natural numbers greater than 0.

[0055] Optionally, the attention determination unit is used to obtain a weight control matrix, which is a binary matrix of the same size as the similarity matrix, and each element (i, j) in the weight control matrix represents a weight control parameter between the i-th query vector and the j-th key vector, and when the j-th key vector corresponds to an extreme value, the element (i, j) is a natural number other than 0, otherwise the element (i, j) is 0; through the attention layer, the similarity matrix and the weight control matrix are added and converted into a probability distribution to obtain an attention matrix, and each element (i, j) in the attention matrix represents the attention weight between the i-th query vector and the j-th key vector.

[0056] Optionally, the device further comprises:

[0057] A parameter adjustment module is used to determine a model prediction error based on a prediction result of the Transformer model; based on the model prediction error, train the Transformer model to reduce the model prediction error, wherein training the Transformer model includes adjusting the weight control parameter.

[0058] Optionally, the vector determination module is used to generate feature vectors of the multiple observations through an embedding layer in the Transformer model, and fuse the indication information into the feature vectors;

[0059] The indication information is used to indicate that the corresponding observation value is the extreme value or the background value, and the background value is the observation value other than the extreme value among the multiple observation values.

[0060] Optionally, the vector determination module is used to generate content embedding vectors and position codes of the multiple observation values ​​through the embedding layer, and add the indication information to the position code of each observation value; splicing the position code of each observation value with the content embedding vector to obtain a feature vector of each observation value.

[0061] Optionally, the device further comprises:

[0062] The first extreme value determination module is used to place a sliding window at the starting position of the reference time series so that the sliding window covers part of the observations in the reference time series; determine the extreme values ​​in the sliding window, the extreme values ​​in the sliding window include the maximum value, the minimum value and the target number of observations around the maximum value and the minimum value in the sliding window; after adjusting the position of the sliding window, re-determine the extreme values ​​in the sliding window until all the extreme values ​​in the reference time series are determined.

[0063] Optionally, the device further comprises:

[0064] The second extreme value determination module is used to determine the change rate corresponding to each observation value based on multiple observation values ​​in the reference time series; screen observation values ​​whose corresponding change rate has a different sign from the change rate corresponding to the previous observation value; and determine the screened observation value and a target number of observation values ​​around the observation value as the extreme value.

[0065] Optionally, the device further includes a model training module, and the model training module includes:

[0066] An observation value acquisition unit, used to acquire observation values ​​corresponding to a plurality of predicted time points after the plurality of time points, wherein the observation values ​​corresponding to the plurality of predicted time points include extreme values ​​and background values ​​other than the extreme values;

[0067] A first error determination unit, configured to determine an extreme value prediction error and a background value prediction error based on the observation values ​​corresponding to the multiple prediction time points and the prediction values ​​corresponding to the multiple prediction time points in the model prediction result;

[0068] A second error determination unit, configured to perform a weighted summation of the extreme value prediction error and the background value prediction error to obtain a model prediction error;

[0069] A model training unit is used to train the Transformer model based on the model prediction error.

[0070] Optionally, the second error determination unit is used to update the weight of the extreme value prediction error used in the previous round of training based on the size of the extreme value prediction error, and the size of the extreme value prediction error is positively correlated with the size of the weight of the extreme value prediction error; based on the updated weight, the extreme value prediction error and the background value prediction error are weightedly summed to obtain the model prediction error.

[0071] In addition, to achieve the above-mentioned purpose, the present application also proposes a data prediction device, which includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the data prediction method described above.

[0072] In addition, to achieve the above-mentioned purpose, the present application also proposes a storage medium, which is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, the steps of the data prediction method described above are implemented.

[0073] In addition, to achieve the above-mentioned purpose, the present application also provides a computer program product, which includes a computer program, and when the computer program is executed by a processor, the steps of the data prediction method described above are implemented.

[0074] One or more technical solutions proposed in this application have at least the following technical effects:

[0075] The data prediction scheme provided by the present application enables the model to better learn the characteristics and laws of extreme values ​​by enhancing the model's attention to the eigenvectors of extreme values, thereby improving the prediction accuracy at extreme values ​​such as peaks and troughs. Among them, in response to the data prediction instruction, a reference time series is extracted from the data prediction instruction, and the reference time series includes observations corresponding to multiple time points, and the multiple observations include extreme values. Then, the eigenvectors of multiple observations are determined. Through the attention layer in the Transformer model, the attention weight between each eigenvector and all eigenvectors is determined based on the weight control parameter. Since the weight control parameter can increase the attention weight between each eigenvector and the eigenvector of the extreme value, when the eigenvector is updated based on the attention weight, the model can pay more attention to the eigenvector corresponding to the extreme value. That is, the model can allocate more attention to capture the complex patterns and dynamic laws behind the extreme value during the learning process. After updating each eigenvector based on the determined attention weight, the updated eigenvector contains more information about the extreme value and the association information between the extreme value and other observations. Therefore, the prediction layer can more accurately predict the observations after multiple time points based on the updated eigenvector, especially can improve the prediction accuracy at the extreme value. BRIEF DESCRIPTION OF THE DRAWINGS

[0076] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.

[0077] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0078] Figure 1 A schematic diagram of an implementation environment of the data prediction method of the present application;

[0079] Figure 2 A schematic diagram of a flow chart provided for the first embodiment of the data prediction method of this application;

[0080] Figure 3This is a detailed schematic diagram of step S30 in the second embodiment of the data prediction method of this application;

[0081] Figure 4 This is a detailed schematic diagram of step S20 in the third embodiment of the data prediction method of this application;

[0082] Figure 5 This is a schematic diagram of the newly added steps in the fourth embodiment of the data prediction method of this application;

[0083] Figure 6 This is a schematic diagram of the newly added steps in the fifth embodiment of the data prediction method of this application;

[0084] Figure 7 This is a schematic diagram of the newly added steps in the sixth embodiment of the data prediction method of this application;

[0085] Figure 8 This is a schematic diagram of the module structure of the data prediction device according to an embodiment of the present application;

[0086] Fig. 9 Schematic diagram of the device structure of the hardware operating environment involved in the data prediction method in the embodiment of the present application.

[0087] The purpose, features and advantages of this application will be further described in conjunction with the embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION

[0088] It should be understood that the specific embodiments described herein are only used to explain the technical solutions of the present application and are not used to limit the present application.

[0089] In order to better understand the technical solution of the present application, a detailed description will be given below in conjunction with the accompanying drawings and specific implementation methods.

[0090] Figure 1 is a schematic diagram of an implementation environment provided by an embodiment of the present disclosure. Figure 1 , the implementation environment includes a terminal 101 and a server 102. The terminal 101 and the server 102 are connected via a wireless or wired network. Exemplarily, the terminal 101 is installed with a target application provided by the server 102, and the terminal 101 can implement functions such as data transmission and message interaction through the target application.

[0091] Exemplarily, the terminal 101 is a computer, a mobile phone, a tablet computer or other terminals. Exemplarily, the target application is a target application in the operating system of the terminal 101, or a target application provided by a third party. For example, the target application is an energy management application, a financial data analysis application, a weather query application, etc. Exemplarily, the server 102 is a background server corresponding to the target application. Accordingly, the server 102 is an energy management application server, a financial data analysis application server, a weather query application server, etc.

[0092] In the present application, terminal 101 is used to respond to a data prediction instruction and extract a reference time series from the data prediction instruction, wherein the reference time series includes observations corresponding to multiple time points, and the multiple observations include extreme values. The reference time series is then sent to server 102. Server 102 is used to receive the reference time series and determine the feature vectors of multiple observations in the reference time series. Through the attention layer in the Transformer model, the attention weight between each feature vector and all feature vectors is determined based on the weight control parameter, and each feature vector is updated based on the determined attention weight, wherein the weight control parameter is used to increase the attention weight between each feature vector and the feature vector of the extreme value. Then, through the prediction layer in the Transformer model, based on the updated feature vector, the observations after multiple time points are predicted.

[0093] Alternatively, the data prediction process may also be completed by the terminal 101 or the server 102 alone. Alternatively, the terminal 101 completes the process through the installed target application. This embodiment of the present application does not limit this.

[0094] The data prediction method provided in the present application is applicable to a variety of scenarios. For example, the scenario of predicting the weather. The terminal responds to the weather forecast instruction and extracts the reference time series of weather data from the weather forecast instruction. The reference time series includes weather observation values ​​corresponding to multiple time points, such as weather observation values ​​corresponding to each day of a month. Then the weather after these multiple time points is predicted by the method provided in the present application. For another example, the data prediction method can also be applied to the scenario of predicting the trend of financial data. The terminal responds to the financial data prediction instruction and extracts the reference time series of financial data from the financial data prediction instruction. The reference time series includes financial data observation values ​​corresponding to multiple time points, such as financial data observation values ​​corresponding to each hour of a day. Then the trend of financial data after these multiple time points is predicted by the method provided in the present application.

[0095] Figure 2 This is a flow chart of the first embodiment of the data prediction method of this application. Figure 2 Taking the execution subject as a terminal as an example, the data prediction method includes the following steps S10 to S40:

[0096] Step S10, in response to the data prediction instruction, extracting a reference time series from the data prediction instruction, the reference time series including observation values ​​corresponding to a plurality of time points, and the plurality of observation values ​​including extreme values.

[0097] The data prediction instruction is a command issued by the user to the terminal to predict the future trend of specific data. The instruction contains key information about the prediction task, such as the data range to be predicted, the time span of the prediction, etc., and is the trigger signal for the entire prediction process.

[0098] The reference time series is a series of data sets extracted from the data forecasting instructions and arranged in chronological order, where each data corresponds to a specific time point. These data reflect the value of a variable at different times, and contain extreme values, i.e., the maximum and minimum values ​​in the sequence, as well as data within a specific range near the maximum and minimum values. These extreme values ​​are of great significance for understanding the fluctuations and trends of the data.

[0099] The observation value is the specific data corresponding to each time point in the reference time series. It is the actual measurement or recording result of the research object at a specific moment, and is the basis for constructing the time series and conducting subsequent analysis. The extreme value is the maximum value, minimum value, and data within a specific range near the maximum value and minimum value in the reference time series, representing the extreme state of the data within a certain period of time. For example, in the stock price time series, the peak and valley values ​​of the stock price and the observation values ​​within a specific range near the peak and valley values ​​are extreme values. They reflect the extreme conditions of stock price fluctuations and play a key role in predicting future price trends.

[0100] Step S20, determining the characteristic vectors of multiple observations.

[0101] The feature vector is a vector obtained by transforming or encoding the observation value. It maps each observation value into a multidimensional vector space, so that the model can better capture the characteristics and patterns of the data. The feature vector not only contains the information of the observation value itself, but also incorporates the location information of the data, the relationship with other observation values, and the indication information indicating whether the corresponding observation value is an extreme value, etc., so that the model can process and analyze it.

[0102] Step S30, through the attention layer in the Transformer model, determine the attention weight between each feature vector and all feature vectors based on the weight control parameter, and update each feature vector based on the determined attention weight. The weight control parameter is used to increase the attention weight between each feature vector and the extreme feature vector.

[0103] The Transformer model is a deep learning model based on the self-attention mechanism, which is widely used in natural language processing, time series prediction and other fields. It effectively captures the long-distance dependencies between elements in the sequence through the self-attention mechanism and performs well in processing sequence data. The attention layer is the core component in the Transformer model, which is responsible for calculating the attention weights between each input element and all other elements, thereby dynamically focusing on different parts of the input sequence.

[0104] Attention weight is a value used to measure the degree of association between each feature vector and other feature vectors in the attention layer. The larger the attention weight, the closer the relationship between the two feature vectors, and the model will pay more attention to it when processing.

[0105] The weight control parameter is a parameter used to adjust the attention weight calculation process. Its function is to increase the attention weight between each eigenvector and the eigenvector of the extreme value. In this way, the model will pay more attention to the information related to the extreme value when calculating the attention, which helps to better capture the key features and patterns in the data.

[0106] Exemplarily, the attention weight between each eigenvector and all eigenvectors is determined based on the weight control parameter, so as to increase the attention weight between each eigenvector and the eigenvector of the extreme value. The implementation method is as follows: before calculating the attention weight between each eigenvector and the eigenvector of the extreme value, an additional weight control parameter is assigned to the eigenvector of the extreme value to increase the attention weight between each eigenvector and the eigenvector of the extreme value. Specifically, before calculating the attention weight between each eigenvector and the eigenvector of the extreme value, the eigenvector of the extreme value is multiplied by the weight control parameter. The weight control parameter does not need to be assigned to other eigenvectors except the eigenvector of the extreme value.

[0107] Alternatively, the attention weight between each eigenvector and all eigenvectors is determined based on the weight control parameter, so as to increase the attention weight between each eigenvector and the eigenvector of the extreme value. The implementation method can also be: after preliminarily calculating the attention weight between each eigenvector and the eigenvector of the extreme value, the attention weight is added with the weight control parameter to increase the attention weight between each eigenvector and the eigenvector of the extreme value. However, the attention weight between each eigenvector and other eigenvectors except the eigenvector of the extreme value does not need to be added with the weight control parameter.

[0108] Step S40, predicting observation values ​​after multiple time points based on the updated feature vector through the prediction layer in the Transformer model.

[0109] The prediction layer is the part of the Transformer model that comes after the attention layer. It uses a specific algorithm or function to predict the observed value at a future time point based on the feature vector updated by the attention layer. The prediction layer contains some fully connected layers or other neural network structures to map the feature vector to the predicted numerical space.

[0110] Optionally, the weight control parameter is a learnable parameter. That is, after predicting the observation values ​​after multiple time points based on the updated feature vector through the prediction layer in the Transformer model, the model prediction error is determined based on the prediction result of the Transformer model; based on the model prediction error, the Transformer model is trained to reduce the model prediction error, wherein training the Transformer model includes adjusting the weight control parameter.

[0111] Model prediction error is used to measure the degree of deviation between the model prediction results and the actual observed values. The calculation methods of model prediction error include mean square error, mean absolute error, etc. Model prediction error reflects the accuracy of the model in the prediction task. The smaller the error, the more accurate the model prediction.

[0112] Training a Transformer model refers to the process of continuously adjusting the model parameters to make the model's prediction results closer to the actual values. In this application, not only the conventional weight parameters of the Transformer model need to be adjusted, but also the weight control parameters need to be adjusted. During the training process, if the model prediction error is large, for example, the model prediction error is greater than the model target error, the weight control parameters are increased to make the model pay more attention to extreme value information; if the model prediction error is small, for example, less than the model target error, the weight control parameters are reduced to avoid the model from focusing too much on extreme values ​​and ignoring other information.

[0113] In order to further improve the prediction accuracy of the model at extreme values, the extreme value prediction error can also be determined based on the prediction results of the Transformer model. Based on the extreme value prediction error, the Transformer model is trained to reduce the extreme value prediction error. For example, during the training process, if the extreme value prediction error is large, such as the extreme value prediction error is greater than the extreme value target error, the weight control parameter is increased to make the model pay more attention to the extreme value information; if the extreme value prediction error is small, such as less than the extreme value target error, the weight control parameter is reduced to prevent the model from paying too much attention to the extreme value and ignoring other information.

[0114] Adjusting the weight control parameters can dynamically optimize the model's attention to extreme value information. In the early stages of training, the model may not accurately grasp the characteristics and rules of extreme values, resulting in large prediction errors. By adjusting the weight control parameters and increasing the attention weight of the feature vector of extreme values, the model can better learn the pattern of extreme values ​​and improve the prediction accuracy of extreme values. As the training progresses, when the model has a good prediction effect on extreme values, the model adjusts the weight control parameters to reduce the attention weight of the feature vector of extreme values, which can avoid the model from overfitting extreme value data, thereby ensuring the generalization ability of the model over the entire time series. In addition, different time series data have different characteristics, such as the distribution of extreme values, the degree of data fluctuation, etc. By adjusting the weight control parameters based on the model prediction error, the model can adapt to different data characteristics. For data where extreme values ​​have a greater impact on the prediction results, the model can increase the weight control parameters to highlight the role of extreme value information; for data where extreme values ​​are relatively less important, the model can reduce the weight control parameters to consider all data information more evenly. This adaptive ability enables the model to maintain good prediction performance when processing various types of time series data.

[0115] The data prediction scheme provided by the present application enables the model to better learn the characteristics and laws of extreme values ​​by enhancing the model's attention to the eigenvectors of extreme values, thereby improving the prediction accuracy at extreme values ​​such as peaks and troughs. Among them, in response to the data prediction instruction, a reference time series is extracted from the data prediction instruction, and the reference time series includes observations corresponding to multiple time points, and the multiple observations include extreme values. Then, the eigenvectors of multiple observations are determined. Through the attention layer in the Transformer model, the attention weight between each eigenvector and all eigenvectors is determined based on the weight control parameter. Since the weight control parameter can increase the attention weight between each eigenvector and the eigenvector of the extreme value, when the eigenvector is updated based on the attention weight, the model can pay more attention to the eigenvector corresponding to the extreme value. That is, the model can allocate more attention to capture the complex patterns and dynamic laws behind the extreme value during the learning process. After updating each eigenvector based on the determined attention weight, the updated eigenvector contains more information about the extreme value and the association information between the extreme value and other observations. Therefore, the prediction layer can more accurately predict the observations after multiple time points based on the updated eigenvector, especially can improve the prediction accuracy at the extreme value.

[0116] Based on the above first embodiment, a second embodiment of the present application is proposed. For the same or similar contents as the first embodiment, please refer to the above introduction and will not be described in detail later. Figure 3 In the second embodiment, the above step S30 includes steps S301 to S303:

[0117] Step S301, determine the similarity between each feature vector and all feature vectors through the attention layer in the Transformer model.

[0118] Similarity is used to measure the similarity between two feature vectors. In the attention mechanism of the Transformer model, the calculation of similarity can help the model determine the closeness of the association between different feature vectors. Similarity can be calculated by cosine similarity, dot product similarity, etc. The higher the similarity between two feature vectors, the closer the directions of the two feature vectors in the feature space are, which means that they may have a closer semantic or data association.

[0119] Optionally, the similarity between each feature vector and all feature vectors is determined through an attention layer, including: converting each feature vector into a query vector and a key vector respectively through an attention layer; performing a dot product operation on each query vector and all key vectors to obtain a similarity matrix, wherein each element (i, j) in the similarity matrix represents the similarity between the i-th query vector and the j-th key vector, where i and j are both natural numbers greater than 0.

[0120] The query vector is a vector obtained by linearly transforming the original feature vector. In the attention mechanism, the query vector is used to match the key vector to determine the degree of attention paid by the feature vector at the current position to the feature vectors at other positions. The query vector can be regarded as a "query tool" to search for relevant information in the entire sequence.

[0121] The key vector is a vector obtained by linearly transforming the original feature vector. The key vector is used in conjunction with the query vector. The key vector at each position represents a certain identifier of the feature at that position. By performing a dot product operation with the query vector, it can reflect the correlation between feature vectors at different positions.

[0122] In the attention mechanism, the dot product operation can quickly calculate the similarity between the query vector and the key vector. The similarity matrix is ​​a matrix obtained by performing the dot product operation on the query vector and the key vector. The size of the similarity matrix is ​​N×N, where N is the number of feature vectors, and it shows the degree of correlation between the feature vectors at each position in the sequence.

[0123] In the embodiment of the present application, the dot product operation is a very efficient matrix operation, which can quickly calculate the similarities between all query vectors and key vectors, avoiding the tedious process of calculating one by one, and greatly improving the calculation efficiency.

[0124] Step S302, through the attention layer, the similarity between each feature vector and all feature vectors is added to the weight control parameter between each feature vector and all feature vectors to obtain the attention weight between each feature vector and all feature vectors, wherein the weight control parameter between each feature vector and the extreme value feature vector is greater than the weight control parameter between each feature vector and other feature vectors.

[0125] The attention weight reflects the importance of other feature vectors to it when updating a certain feature vector. Through the attention weight, the model can dynamically pay attention to different parts of the sequence. In this application, the attention weight is obtained by adding the similarity between the feature vectors to the weight control parameter, which determines the fusion ratio of each feature vector to other feature vectors during the update process. The weight control parameter between each feature vector and the feature vector of the extreme value is greater than the weight control parameter between each feature vector and the feature vector of the extreme value, which can enhance the model's attention to the extreme value feature vector, allowing the model to pay more attention to the information related to the extreme value when calculating the attention weight, thereby improving the ability to capture and process the extreme value.

[0126] Optionally, through the attention layer, the similarity between each feature vector and all feature vectors is added to the weight control parameter between each feature vector and all feature vectors to obtain the attention weight between each feature vector and all feature vectors, including: obtaining a weight control matrix, the weight control matrix is ​​a binary matrix of the same size as the similarity matrix, each element (i, j) in the weight control matrix represents the weight control parameter between the i-th query vector and the j-th key vector, and when the j-th key vector corresponds to an extreme value, the element (i, j) is a natural number other than 0, otherwise the element (i, j) is 0; through the attention layer, the similarity matrix and the weight control matrix are added and converted into a probability distribution to obtain an attention matrix, each element (i, j) in the attention matrix represents the attention weight between the i-th query vector and the j-th key vector.

[0127] The weight control matrix is ​​a binary matrix of the same size as the similarity matrix. That is, the elements in the matrix can only take two values. In this scheme, the elements are either 0 or natural numbers other than 0, such as 1. This matrix has a simple structure and can clearly identify which positions need weight control. In this matrix, when the jth key vector corresponds to an extreme value, the element (i, j) is a non-zero natural number, indicating that the attention to the eigenvector of this extreme value should be strengthened when calculating the attention weight; if the jth key vector does not correspond to an extreme value, the element (i, j) is 0, which means that the attention weight of the eigenvector at this position is not adjusted additionally.

[0128] Adding the similarity matrix and the weight control matrix and converting them into a probability distribution to obtain the attention matrix means converting the matrix elements into a numerical distribution that satisfies the probabilistic property, that is, the sum of the elements in each row of the matrix is ​​1, and the value of each element is between 0 and 1. In the attention mechanism, converting the matrix into a probability distribution helps determine the relative importance of each key vector to the query vector.

[0129] The attention matrix is ​​obtained by adding the similarity matrix and the weight control matrix and converting them into a probability distribution. Each element (i, j) in the matrix represents the attention weight between the i-th query vector and the j-th key vector, reflecting the degree to which the feature vector at the j-th position should be paid attention to when calculating the feature vector at the i-th position.

[0130] Exemplarily, before adding the similarity matrix to the weight control matrix, the weight control matrix is ​​multiplied by the target learning parameter. The target learning parameter is a learnable parameter, that is, a parameter that will be adjusted as the model is trained. If a fixed weight control matrix is ​​used directly, the model may over-rely on extreme value information, resulting in good performance on training data but poor performance on new data, that is, overfitting. By introducing a learnable target learning parameter, the model can automatically adjust the degree of attention to extreme values ​​during training, avoid excessive reliance on extreme values, and thus improve the generalization ability of the model on different data sets.

[0131] Exemplarily, the attention matrix is ​​calculated by the following formula (1).

[0132] Formula (1):

[0133] Among them, q i is the query vector, k j is the key vector, d k is the dimension of the key vector, γ is the target learning parameter, PeakValleyMask ij is the weight control matrix, w ij is the attention matrix.

[0134] In the embodiment of the present application, the weight control matrix can accurately mark the key vectors corresponding to the extreme values. By adding the weight control matrix on the basis of the similarity matrix, the model can specifically enhance the attention to extreme value information when calculating the attention weight, and maintain normal attention to non-extreme value information. In this way, the model can not only highlight the importance of extreme value information, but also take into account the information of other parts of the sequence, ensuring the flexibility and comprehensiveness of the model when processing the entire time series data.

[0135] Step S303, through the attention layer, update each feature vector based on the determined attention weight.

[0136] Optionally, in addition to converting each feature vector into a query vector and a key vector, the above steps also need to convert each feature vector into a value vector. After obtaining the attention matrix, multiply the attention matrix by the value vector matrix to obtain an updated feature vector matrix. Among them, each row in the value vector matrix is ​​a value vector. Each row in the updated feature vector matrix is ​​an updated feature vector, which fuses the feature vectors at various positions in the sequence, and depending on the attention weights, the fusion ratio of feature vectors at different positions is also different.

[0137] One thing that needs to be explained is that there are multiple attention layers in the Transformer model. After each attention layer converts the query vector, key vector, and value vector of the input data, it can calculate the attention weight between the query vector and the key vector through the above method, and update the feature vector based on the attention weight.

[0138] In an embodiment of the present application, by setting a larger weight control parameter between each eigenvector and the eigenvector of the extreme value, the model will pay more attention to the extreme value eigenvector when calculating the attention weight. This enables the model to integrate more extreme value information when updating the eigenvector based on the attention weight, and thus better learn and capture the complex patterns and laws behind the extreme value, and improve the prediction accuracy of the extreme value. And the attention weight is calculated by combining the similarity between the eigenvectors and the weight control parameters, which not only takes into account the intrinsic correlation of the eigenvector itself, but also introduces additional prior information through the weight control parameters, which enhances the adaptability and flexibility of the model and improves the generalization ability of the model.

[0139] Based on the above first embodiment of the present application, a third embodiment of the present application is proposed. The same or similar contents as the first embodiment can be referred to the above introduction, and will not be repeated in the following. Figure 4 In the third embodiment, step S20 includes step S201 and step S202.

[0140] Step S201, through the embedding layer in the Transformer model, generate content embedding vectors and position codes of multiple observation values, and add indication information in the position code of each observation value, the indication information is used to indicate that the corresponding observation value is an extreme value or a background value, and the background value is an observation value other than the extreme value among the multiple observation values.

[0141] In the Transformer model, the role of the embedding layer is to convert discrete input data, that is, observations in the reference time series, into continuous vector representations. For multiple observations in the time series, the embedding layer maps each observation into a high-dimensional vector space. This vector is called a content embedding vector, which can capture the semantic and numerical relationships between observations, making it easier for subsequent models to process.

[0142] Position encoding is used to add the position information of each observation in the sequence. It is a vector with the same dimension as the content embedding vector. By combining position encoding with the content embedding vector, the model can understand the order of observations in the time series, thereby better capturing the temporal dependencies in the sequence.

[0143] Indicative information is additional information added to the positional encoding of each observation, which is used to clearly indicate whether the corresponding observation is an extreme value or a background value. Extreme values ​​are the maximum value, minimum value, and observations within a specific range near the maximum value and minimum value in the reference time series, while background values ​​are observations other than extreme values. Indicative information can help the model pay more attention to the key information contained in extreme values ​​when processing the series.

[0144] Exemplarily, the implementation method of adding the indication information in the position code of each observation value is: adding a binary flag bit in the position code of each observation value. If the binary flag bit is 1, it indicates that the corresponding observation value is an extreme value, and if the binary flag bit is 0, it indicates that the corresponding observation value is a background value.

[0145] Step S202: concatenate the position code of each observation value with the content embedding vector through the embedding layer to obtain the feature vector of each observation value.

[0146] The feature vector obtained by concatenating the position encoding of each observation value with the content embedding vector integrates the content information and position information of the observation value, as well as the information whether the observation value is an extreme value or a background value. The subsequent model can understand the semantics of extreme values ​​in the global context, such as whether the extreme value is a key turning point in trend changes or a trigger point for abnormal signals, thereby helping the model learn hidden patterns and contextual associations, which helps the model make accurate predictions based on this information.

[0147] This embodiment is only one implementation method of generating feature vectors of multiple observation values ​​through the embedding layer in the Transformer model and fusing the indication information into the feature vector. In addition to the above method, other methods can also be used to fuse the indication information into the feature vector, and this application does not impose any restrictions on this.

[0148] In the embodiment of the present application, considering that in time series data, extreme values ​​often represent key trend changes, abnormal events or important turning points, which have a significant impact on the prediction results. By integrating the indicative information into the feature vector, the model can clearly distinguish between extreme values ​​and background values, so that more attention can be paid to the information contained in the extreme values ​​during feature extraction. In the prediction stage, since the feature vector that integrates the indicative information can provide the model with richer and more accurate information, the model can make more reasonable and accurate predictions, especially improve the prediction accuracy at the extreme values.

[0149] Moreover, the position encoding, content embedding vector and indicator information are combined to form a feature vector, so that the model can simultaneously utilize multiple aspects of information such as the content, location and special attributes of the observation value. This comprehensive feature expression can more comprehensively describe the characteristics of the time series, allowing the model to learn more complex patterns and relationships, thereby improving the model's prediction performance and generalization ability.

[0150] Based on the above first embodiment of the present application, a fourth embodiment of the present application is proposed. The same or similar contents as the first embodiment can be referred to the above introduction, and will not be repeated in the following. Figure 5 In the fourth embodiment, steps S501 to S503 are included between step S10 and step S20.

[0151] Step S501, placing the sliding window at the starting position of the reference time series so that the sliding window covers part of the observations in the reference time series.

[0152] Sliding window is a tool for local analysis on time series data. It is essentially a "window" of fixed length. In this application, this window slides on the reference time series, covering a portion of the continuous observations in the reference time series each time. By moving this window, different local segments of the reference time series can be processed in turn.

[0153] The starting position of the sliding window is the position where the sliding window is initially placed on the reference time series, that is, the position of the first observation, so that the sliding window can start analyzing from the beginning of the reference time series.

[0154] Step S502, determining extreme values ​​in the sliding window, where the extreme values ​​in the sliding window include the maximum value, the minimum value, and the target number of observations around the maximum value and the minimum value in the sliding window.

[0155] The extreme value in the sliding window is a special value in the part of the observations covered by the sliding window. It not only includes the maximum and minimum values ​​in the window, but also includes a certain number of observations around the maximum and minimum values. The observations around the maximum and minimum values ​​are also included in the extreme value category because these adjacent values ​​may be closely related to the maximum and minimum values, which is of great significance for analyzing the local characteristics of the time series.

[0156] The target number is used to specify how many observations are selected around the maximum and minimum values ​​as part of the extreme value. For example, if the target number is set to 2, then when determining the extreme value in the sliding window, in addition to the maximum and minimum values, 2 observations before and after the maximum value and 2 observations before and after the minimum value will be selected as extreme values.

[0157] Step S503, after adjusting the position of the sliding window, re-determine the extreme values ​​in the sliding window until all the extreme values ​​in the reference time series are determined.

[0158] Adjusting the position of the sliding window means moving the sliding window forward along the reference time series by a certain step size. After each move, the window will cover different observation value fragments in the reference time series, so that different parts of the reference time series can be analyzed. The moving step size can be set according to specific needs. For example, the moving step size is 2, which means that the window moves forward by two observation value positions each time.

[0159] In the embodiment of the present application, the time series is locally scanned with the help of a sliding window, and the extreme values ​​can be accurately found in each local segment. And the observations near the maximum and minimum values ​​are also regarded as extreme values, which can avoid missing some important information related to extreme values ​​and improve the accuracy of subsequent extreme value predictions. Compared with the global analysis of the entire time series to identify extreme values, the sliding window method only needs to process the local data within the window, which greatly reduces the computational complexity.

[0160] Based on the above first embodiment of the present application, a fifth embodiment of the present application is proposed. For the same or similar contents as the first embodiment, please refer to the above introduction, and no further description will be given later. Figure 6 In the fifth embodiment, steps S601 to S603 are included between step S10 and step S20.

[0161] Step S601, based on multiple observations in a reference time series, determine the change rate corresponding to each observation.

[0162] In the reference time series, the rate of change of the observations measures the speed and direction of the change between the values ​​of adjacent observations, and can be calculated by subtracting the previous observation from the next observation. A positive rate of change indicates an upward trend in the observations, while a negative rate of change indicates a downward trend.

[0163] Step S602, screening observations whose corresponding change rate signs are different from the change rate signs corresponding to the previous observations.

[0164] The sign of the rate of change refers to the positive or negative nature of the rate of change. A positive sign indicates that the observed value at the current time point is increasing compared to the previous time point, and a negative sign indicates that the observed value at the current time point is decreasing compared to the previous time point. A change in sign means that the trend of the reference time series has turned.

[0165] For an observation in the reference time series, the previous observation is the observation that is immediately adjacent to and before the observation in time order. Observations whose corresponding rate of change has a different sign from the rate of change corresponding to the previous observation are turning points in the reference time series. These turning points are the maximum or minimum values ​​in the reference time series.

[0166] Exemplarily, the observations may also be screened in combination with the second-order derivative of the reference time series. Specifically, after screening out the observations whose corresponding change rate has a different sign from the change rate corresponding to the previous observation, the second-order derivative corresponding to the screened observations is further determined. If the second-order derivative corresponding to the observation is positive, it means that the observation is the minimum value. If the second-order derivative corresponding to the observation is negative, it means that the observation is the maximum value.

[0167] Step S603: determine the screened observation value and a target number of observation values ​​around the observation value as extreme values.

[0168] Among them, the target number is used to determine how many observations around the filtered observations are selected as part of the extreme value.

[0169] In the embodiment of the present application, by comparing the signs of the rates of change of adjacent observations, the turning points in the reference time series can be accurately identified. Since the trend of the reference time series changes at the extreme values, these turning points often correspond to the extreme values. Therefore, the method of determining the extreme values ​​by screening out these turning points is simple and effective. In addition, the observations of the target number around the screened turning points are also determined as extreme values, which can avoid missing some important information related to the extreme values ​​and provide richer information for subsequent analysis and prediction.

[0170] Based on the above first embodiment of the present application, a sixth embodiment of the present application is proposed. For the same or similar contents as the first embodiment, please refer to the above introduction, and no further description will be given later. Figure 7 In the sixth embodiment, after step S40, steps S701 to S704 are also included.

[0171] Step S701, obtaining observation values ​​corresponding to multiple predicted time points after multiple time points, wherein the observation values ​​corresponding to the multiple predicted time points include extreme values ​​and background values ​​other than the extreme values.

[0172] The predicted time points are the time points whose values ​​are expected to be predicted by the model. These time points are after multiple time points in the reference time series and are the moments that need to be predicted in the future. The observed values ​​corresponding to the predicted time points are the actual values ​​corresponding to these predicted time points.

[0173] Background values ​​are the values ​​in the time series other than extreme values. They constitute the main part of the time series and reflect the fluctuations and changes of the target variable under normal circumstances. Compared with extreme values, their changes are relatively stable.

[0174] Step S702, based on the observation values ​​corresponding to the multiple prediction time points and the prediction values ​​corresponding to the multiple prediction time points in the model prediction results, determine the extreme value prediction error and the background value prediction error.

[0175] The extreme value prediction error is a measure of the accuracy of the model's prediction of the extreme value, which is determined by calculating the difference between the observed value and the predicted value of the extreme value. For example, the extreme value prediction error is the difference between the observed value and the predicted value of the extreme value or the mean square error.

[0176] The background value prediction error is an indicator to measure the accuracy of the model's prediction of the background value, which is determined by calculating the difference between the observed value and the predicted value of the background value. For example, the background value prediction error is the difference between the observed value and the predicted value of the background value or the mean square error.

[0177] Exemplarily, for the observed values ​​and predicted values ​​corresponding to multiple prediction time points, they are first distinguished according to extreme values ​​and background values, that is, the observed values ​​and predicted values ​​corresponding to all extreme values, and the observed values ​​and predicted values ​​corresponding to all background values ​​are determined. Then, based on the observed values ​​and predicted values ​​corresponding to the extreme values, the extreme value prediction error is determined, and based on the observed values ​​and predicted values ​​corresponding to the background values, the background value prediction error is determined.

[0178] Step S703, performing weighted summation on the extreme value prediction error and the background value prediction error to obtain the model prediction error.

[0179] Optionally, a weighted sum of the extreme value prediction error and the background value prediction error is performed to obtain a model prediction error, including: based on the size of the extreme value prediction error, the weight of the extreme value prediction error used in the previous round of training is updated, and the size of the extreme value prediction error is positively correlated with the size of the weight of the extreme value prediction error; based on the updated weight, a weighted sum of the extreme value prediction error and the background value prediction error is performed to obtain the model prediction error.

[0180] The weight of the extreme value prediction error is a coefficient assigned to the extreme value prediction error when calculating the model prediction error. It reflects the importance of the extreme value prediction error in the overall model prediction error. The larger the weight, the greater the impact of the extreme value prediction error on the result when calculating the model prediction error.

[0181] The previous round of training refers to the previous round of training process before the current training step. In model training, multiple rounds of iterative training will be performed, and each round of training will update the model parameters. The previous round of training refers to the training operation of the previous iteration.

[0182] The size of the extreme value prediction error is positively correlated with the size of the extreme value prediction error weight, which means that the larger the extreme value prediction error, the larger the extreme value prediction error weight; conversely, the smaller the extreme value prediction error, the smaller the extreme value prediction error weight. The larger the extreme value prediction error, the greater the difficulty of predicting the extreme value. The greater the weight of the extreme value prediction error, the more the model pays attention to the extreme value part, thereby improving the prediction accuracy of the extreme value. Conversely, the smaller the extreme value prediction error, the smaller the extreme value prediction error, the smaller the weight of the extreme value prediction error, which can avoid the model from paying too much attention to the extreme value part, resulting in affecting the prediction accuracy of the background value.

[0183] Exemplarily, a linear update method is used to update the weight of the extreme value prediction error used in the previous round of training. That is, a fixed update step size D is set, and the weight is adjusted linearly according to the change of the extreme value prediction error. Specifically, assume that the extreme value prediction error of the current round is E2, and the extreme value prediction error of the previous round is E1. If E2 is greater than E1, the weight W1 of the extreme value prediction error of the previous round is increased by D to obtain the updated weight. If E2 is less than E1, the weight W1 is reduced by D to obtain the updated weight.

[0184] Exemplarily, the weight of the extreme value prediction error used in the previous round training is dynamically adjusted according to the change ratio of the extreme value prediction error. Still taking the above data as an example, the ratio of the extreme value prediction error E2 of the current round to the extreme value prediction error E1 of the previous round is added to the extreme value prediction error weight W1 of the previous round to obtain the updated weight.

[0185] Step S704: training the Transformer model based on the model prediction error.

[0186] The reference time series contains extreme values ​​and background values. In traditional training, the model may not learn enough extreme values ​​because of the scarcity of extreme values. The embodiment of the present application can reasonably allocate attention to extreme values ​​and background values ​​during model training by dynamically adjusting the weight of the extreme value prediction error. When the extreme value prediction error is large, the weight of the extreme value prediction error is increased to increase attention to the extreme value. When the extreme value prediction error is small, the weight of the extreme value prediction error is reduced, and more attention is put back on the background value to ensure that the model fully learns the entire time series. In this way, the model can not only accurately predict the background value, but also make stable predictions of the extreme value.

[0187] One point that needs to be explained is that the present application improves the embedding layer in the Transformer model so that the feature vector it converts contains information indicating whether the corresponding observation is an extreme value. In addition, the attention layer is improved so that it automatically adjusts the degree of attention to extreme values ​​during feature extraction through weight control parameters. In addition, the loss function used for model training is improved so that it distinguishes between the prediction loss of extreme values ​​and the prediction loss of background values, and adaptively adjusts the weight of the prediction loss of extreme values ​​according to the prediction loss of extreme values, thereby adjusting the model's attention to extreme values. The present application improves the Transformer model from multiple dimensions, thereby improving the prediction accuracy of the model, especially the prediction accuracy at extreme values.

[0188] Another point that needs to be explained is that the above examples are only used to understand the present application and do not constitute a limitation on the data prediction method of the present application. More simple transformations based on this technical concept are all within the scope of protection of the present application.

[0189] This application also provides a data prediction device, please refer to Figure 8 , the data prediction device comprises:

[0190] The instruction response module 10 is used to respond to the data prediction instruction and extract a reference time series from the data prediction instruction, where the reference time series includes observation values ​​corresponding to multiple time points, and the multiple observation values ​​include extreme values;

[0191] A vector determination module 20, used to determine the characteristic vectors of a plurality of observations;

[0192] A vector updating module 30 is used to determine the attention weight between each feature vector and all feature vectors based on a weight control parameter through an attention layer in a transformer model, and to update each feature vector based on the determined attention weight, wherein the weight control parameter is used to increase the attention weight between each feature vector and the feature vector of the extreme value;

[0193] The data prediction module 40 is used to predict the observation values ​​after multiple time points based on the updated feature vector through the prediction layer in the Transformer model.

[0194] Optionally, the vector updating module 30 includes:

[0195] A similarity determination unit, used to determine the similarity between each feature vector and all feature vectors through an attention layer;

[0196] An attention determination unit, used for adding the similarity between each feature vector and all feature vectors to the weight control parameter between each feature vector and all feature vectors through the attention layer to obtain the attention weight between each feature vector and all feature vectors;

[0197] Among them, the weight control parameter between each eigenvector and the extreme value eigenvector is greater than the weight control parameter between each eigenvector and other eigenvectors.

[0198] Optionally, the similarity determination unit is used to convert each feature vector into a query vector and a key vector respectively through an attention layer; perform a dot product operation on each query vector and all key vectors to obtain a similarity matrix, wherein each element (i, j) in the similarity matrix represents the similarity between the i-th query vector and the j-th key vector, where i and j are both natural numbers greater than 0.

[0199] Optionally, the attention determination unit is used to obtain a weight control matrix, which is a binary matrix of the same size as the similarity matrix, each element (i, j) in the weight control matrix represents a weight control parameter between the i-th query vector and the j-th key vector, and when the j-th key vector corresponds to an extreme value, the element (i, j) is a natural number other than 0, otherwise the element (i, j) is 0; through the attention layer, the similarity matrix and the weight control matrix are added and converted into a probability distribution to obtain an attention matrix, wherein each element (i, j) in the attention matrix represents the attention weight between the i-th query vector and the j-th key vector.

[0200] Optionally, the device further comprises:

[0201] The parameter adjustment module is used to determine the model prediction error based on the prediction result of the Transformer model; based on the model prediction error, the Transformer model is trained to reduce the model prediction error, wherein the training of the Transformer model includes adjusting the weight control parameters.

[0202] Optionally, the vector determination module 20 is used to generate feature vectors of multiple observations through an embedding layer in a Transformer model, and fuse the indication information into the feature vectors;

[0203] The indication information is used to indicate that the corresponding observation value is an extreme value or a background value, and the background value is an observation value other than the extreme value among multiple observation values.

[0204] Optionally, the vector determination module 20 is used to generate content embedding vectors and position codes of multiple observation values ​​through an embedding layer, and add indication information to the position code of each observation value; the position code of each observation value is concatenated with the content embedding vector to obtain a feature vector of each observation value.

[0205] Optionally, the device further comprises:

[0206] The first extreme value determination module is used to place the sliding window at the starting position of the reference time series so that the sliding window covers part of the observations in the reference time series; determine the extreme values ​​in the sliding window, the extreme values ​​in the sliding window include the maximum value, the minimum value and the target number of observations around the maximum value and the minimum value in the sliding window; after adjusting the position of the sliding window, re-determine the extreme values ​​in the sliding window until all the extreme values ​​in the reference time series are determined.

[0207] Optionally, the device further comprises:

[0208] The second extreme value determination module is used to determine the change rate corresponding to each observation value based on multiple observation values ​​in the reference time series; screen observation values ​​whose corresponding change rate has a different sign from the change rate corresponding to the previous observation value; and determine the screened observation value and a target number of observation values ​​around the observation value as extreme values.

[0209] Optionally, the device further includes a model training module, which includes:

[0210] An observation value acquisition unit, used to acquire observation values ​​corresponding to multiple predicted time points after multiple time points, wherein the observation values ​​corresponding to the multiple predicted time points include extreme values ​​and background values ​​other than the extreme values;

[0211] A first error determination unit, configured to determine an extreme value prediction error and a background value prediction error based on observation values ​​corresponding to a plurality of prediction time points and prediction values ​​corresponding to a plurality of prediction time points in a model prediction result;

[0212] A second error determination unit is used to perform weighted summation of the extreme value prediction error and the background value prediction error to obtain a model prediction error;

[0213] The model training unit is used to train the Transformer model based on the model prediction error.

[0214] Optionally, the second error determination unit is used to update the weight of the extreme value prediction error used in the previous round of training based on the size of the extreme value prediction error, and the size of the extreme value prediction error is positively correlated with the size of the weight of the extreme value prediction error; based on the updated weight, the extreme value prediction error and the background value prediction error are weightedly summed to obtain the model prediction error.

[0215] The data prediction device provided by the present application adopts the data prediction method in the above embodiment, which can solve the technical problem of low prediction accuracy of the Transformer model at extreme values ​​in the related art. Compared with the prior art, the beneficial effects of the data prediction device provided by the present application are the same as the beneficial effects of the data prediction method provided by the above embodiment, and the other technical features in the data prediction device are the same as the features disclosed in the above embodiment method, which will not be repeated here.

[0216] The present application provides a data prediction device, which includes: at least one processor; and a memory that is communicatively connected to the at least one processor; wherein the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the data prediction method in the above-mentioned embodiment one.

[0217] Reference below Fig. 9 , which shows a schematic diagram of the structure of a data prediction device suitable for implementing an embodiment of the present application. The data prediction device in the embodiment of the present application may include but is not limited to mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Descriptions), PMPs (Portable Media Players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Fig. 9 The data prediction device shown is merely an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.

[0218] like Fig. 9As shown, the data prediction device may include a processing device 1001 (e.g., a central processing unit, a graphics processor, etc.), which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM: Read Only Memory) 1002 or a program loaded from a storage device 1003 to a random access memory (RAM: Random Access Memory) 1004. In RAM1004, various programs and data required for the operation of the data prediction device are also stored. The processing device 1001, ROM1002, and RAM1004 are connected to each other through a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Generally, the following systems can be connected to the I / O interface 1006: an input device 1007 including, for example, a touch screen, a touch pad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output device 1008 including, for example, a liquid crystal display (LCD: Liquid Crystal Display), a speaker, a vibrator, etc.; a storage device 1003 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1009. The communication device 1009 can allow the data prediction device to communicate with other devices wirelessly or by wire to exchange data. Although the figure shows a data prediction device with various systems, it should be understood that it is not required to implement or have all the systems shown. More or fewer systems can be implemented or have alternatively.

[0219] In particular, according to the embodiments disclosed in the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, the embodiments disclosed in the present application include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network through a communication device, or installed from a storage device 1003, or installed from a ROM 1002. When the computer program is executed by the processing device 1001, the above-mentioned functions defined in the method of the embodiment disclosed in the present application are executed.

[0220] The data prediction device provided by the present application adopts the data prediction method in the above embodiment, which can solve the technical problem of low prediction accuracy of the Transformer model at extreme values ​​in the related art. Compared with the prior art, the beneficial effects of the data prediction device provided by the present application are the same as the beneficial effects of the data prediction method provided by the above embodiment, and the other technical features in the data prediction device are the same as the features disclosed in the method of the previous embodiment, which will not be repeated here.

[0221] It should be understood that the various parts disclosed in this application can be implemented by hardware, software, firmware or a combination thereof. In the description of the above embodiments, specific features, structures, materials or characteristics can be combined in any one or more embodiments or examples in a suitable manner.

[0222] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art who is familiar with the present technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.

[0223] The present application provides a computer-readable storage medium having computer-readable program instructions (ie, computer programs) stored thereon, wherein the computer-readable program instructions are used to execute the data prediction method in the above-mentioned embodiment.

[0224] The computer-readable storage medium provided in the present application may be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, systems or devices, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in combination with an instruction execution system, system or device. The program code contained on the computer-readable storage medium may be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination of the above.

[0225] The computer-readable storage medium may be included in the data prediction device; or may exist independently without being assembled into the data prediction device.

[0226] The above-mentioned computer-readable storage medium carries one or more programs. When the above-mentioned one or more programs are executed by the data prediction device, the data prediction device: responds to the data prediction instruction, extracts a reference time series from the data prediction instruction, and the reference time series includes observations corresponding to multiple time points, and the multiple observations include extreme values; determines the feature vectors of the multiple observations; determines the attention weight between each feature vector and all feature vectors based on the weight control parameter through the attention layer in the Transformer model, and updates each feature vector based on the determined attention weight, and the weight control parameter is used to increase the attention weight between each feature vector and the feature vector of the extreme value; through the prediction layer in the Transformer model, based on the updated feature vector, predicts the observations after multiple time points.

[0227] Computer program code for performing the operations of the present application may be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, C++, and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0228] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present application. In this regard, each square box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two square boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0229] The modules involved in the embodiments described in this application may be implemented by software or hardware, wherein the name of the module does not constitute a limitation on the unit itself in some cases.

[0230] The readable storage medium provided in this application is a computer-readable storage medium, which stores computer-readable program instructions (i.e., computer programs) for executing the above-mentioned data prediction method, and can solve the technical problem of low prediction accuracy of the Transformer model at extreme values ​​in the related art. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as the beneficial effects of the data prediction method provided in the above-mentioned embodiment, and will not be repeated here.

[0231] The present application also provides a computer program product, including a computer program, which implements the steps of the above-mentioned data prediction method when executed by a processor.

[0232] The computer program product provided by this application can solve the technical problem of low prediction accuracy of the Transformer model at extreme values ​​in the related art. Compared with the prior art, the beneficial effects of the computer program product provided by this application are the same as the beneficial effects of the data prediction method provided by the above embodiment, which will not be repeated here.

[0233] The above descriptions are only some embodiments of the present application, and are not intended to limit the patent scope of the present application. All equivalent structural changes made using the contents of the present application specification and drawings under the technical concept of the present application, or direct / indirect applications in other related technical fields are included in the patent protection scope of the present application.

Claims

1. A data prediction method, characterized in that: The method comprises: In response to the data prediction instruction, extracting a reference time series from the data prediction instruction, the reference time series including observation values ​​corresponding to a plurality of time points, and the plurality of observation values ​​including an extreme value; determining a feature vector of the plurality of observations; Determine, through an attention layer in a transformer model, an attention weight between each feature vector and all feature vectors based on a weight control parameter, and update each feature vector based on the determined attention weight, wherein the weight control parameter is used to increase the attention weight between each feature vector and the feature vector of the extreme value; The prediction layer in the Transformer model predicts the observation values ​​after the multiple time points based on the updated feature vector.

2. The method according to claim 1, characterized in that The attention layer in the transformer model determines the attention weight between each feature vector and all feature vectors based on the weight control parameter, including: Determine the similarity between each feature vector and all feature vectors through the attention layer; Through the attention layer, the similarity between each feature vector and all feature vectors is added to the weight control parameter between each feature vector and all feature vectors to obtain the attention weight between each feature vector and all feature vectors; The weight control parameter between each eigenvector and the eigenvector of the extreme value is greater than the weight control parameter between each eigenvector and other eigenvectors.

3. The method according to claim 2, characterized in that Determining the similarity between each feature vector and all feature vectors through the attention layer includes: Through the attention layer, each feature vector is converted into a query vector and a key vector respectively; A dot product operation is performed on each query vector and all key vectors to obtain a similarity matrix, wherein each element (i, j) in the similarity matrix represents the similarity between the i-th query vector and the j-th key vector, wherein i and j are both natural numbers greater than 0.

4. The method according to claim 3, characterized in that The attention layer adds the similarity between each feature vector and all feature vectors to the weight control parameter between each feature vector and all feature vectors to obtain the attention weight between each feature vector and all feature vectors, including: Obtain a weight control matrix, where the weight control matrix is ​​a binary matrix of the same size as the similarity matrix, each element (i, j) in the weight control matrix represents a weight control parameter between the i-th query vector and the j-th key vector, and when the j-th key vector corresponds to an extreme value, the element (i, j) is a natural number other than 0, otherwise the element (i, j) is 0; Through the attention layer, the similarity matrix and the weight control matrix are added and converted into a probability distribution to obtain an attention matrix, wherein each element (i, j) in the attention matrix represents the attention weight between the i-th query vector and the j-th key vector.

5. The method according to claim 1, characterized in that After predicting the observation values ​​after the plurality of time points based on the updated feature vectors through the prediction layer in the Transformer model, the method further comprises: Determining a model prediction error based on a prediction result of the Transformer model; Based on the model prediction error, the Transformer model is trained to reduce the model prediction error, wherein training the Transformer model includes adjusting the weight control parameter.

6. The method according to claim 1, characterized in that The step of determining the feature vectors of the plurality of observations comprises: Generate feature vectors of the plurality of observations through an embedding layer in the Transformer model, and fuse the indication information into the feature vectors; The indication information is used to indicate that the corresponding observation value is the extreme value or the background value, and the background value is the observation value other than the extreme value among the multiple observation values.

7. A data prediction device, characterized in that: The device comprises: An instruction response module, configured to respond to a data prediction instruction and extract a reference time series from the data prediction instruction, wherein the reference time series includes observation values ​​corresponding to a plurality of time points, and the plurality of observation values ​​include extreme values; A vector determination module, used to determine the characteristic vectors of the plurality of observations; A vector update module, for determining an attention weight between each feature vector and all feature vectors based on a weight control parameter through an attention layer in a transformer model, and updating each feature vector based on the determined attention weight, wherein the weight control parameter is used to increase the attention weight between each feature vector and the feature vector of the extreme value; A data prediction module is used to predict the observation values ​​after the multiple time points based on the updated feature vectors through the prediction layer in the Transformer model.

8. A data prediction device, characterized in that: The device comprises: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the data prediction method according to any one of claims 1 to 6.

9. A storage medium, characterized in that: The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, the steps of the data prediction method according to any one of claims 1 to 6 are implemented.

10. A computer program product, characterized in that The computer program product comprises a computer program, and when the computer program is executed by a processor, the steps of the data prediction method according to any one of claims 1 to 6 are implemented.

Citation Information

Cited By

  • A fault-tolerant control method for vehicle steer-by-wire that integrates embodied large model prediction technology

    CN122667106A