Multi-head parallel time sequence prediction method and device, equipment and storage medium
Through the multi-head parallel time series prediction method, combined with the gating attention mechanism and the weighted loss mechanism, the problem of insufficient error accumulation and trend fluctuation capture capabilities in the long-term prediction of the timing prediction model in the prior art is solved, and higher prediction accuracy and model performance are achieved.
Patent Information
- Application Number
- CN202510107751.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-23
- Publication Date
- 2025-05-09
AI Technical Summary
The existing timing prediction models have reduced prediction accuracy due to error accumulation in long-term prediction, and they lack the ability to capture long-term trends and short-term fluctuations, making it difficult to meet high-precision needs.
The multi-head parallel time series prediction method is adopted to determine the target context vector through the gating attention mechanism, decompose the prediction task and determine the initial weight, and use the weighting loss mechanism to perform parallel prediction, and dynamically adjust the weight to improve the prediction results.
It effectively avoids the accuracy reduction caused by error accumulation in long-term prediction, improves the performance, stability, flexibility and adaptability of the timing prediction model, and can better cope with complex and high volatility timing data prediction tasks.
Smart Images

Figure CN119961680A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of time series prediction, and in particular to a multi-head parallel time series prediction method, device, equipment and storage medium. Background Art
[0002] With the rapid development of artificial intelligence technology, time series data prediction has been widely used in many fields.
[0003] However, current time series prediction solutions usually rely on autoregressive models or block batch regression models for prediction. By using the prediction results of the previous time step as input to generate future predictions, as the prediction length increases, the error will gradually accumulate, leading to a decrease in the stability and accuracy of the model. The block batch regression method predicts future time points through block parallel calculations. This method usually faces the problem of the model's insufficient ability to capture long-term trends and short-term fluctuations. Both of them often show poor prediction results when dealing with long-term, high-volatility, and nonlinear complex time series data. Especially in complex tasks that need to capture long-term trends and short-term fluctuations at the same time, the performance of these models is difficult to meet high-precision requirements. Summary of the invention
[0004] In view of this, the purpose of the present invention is to provide a multi-head parallel time series prediction method, device, equipment and storage medium, which can effectively avoid the decline in prediction accuracy due to error accumulation in long-term prediction, and improve the performance, stability, flexibility and adaptability of the time series prediction model, so as to better cope with complex and highly volatile time series data prediction tasks. The specific scheme is as follows:
[0005] In a first aspect, the present application provides a multi-head parallel time series prediction method, comprising:
[0006] Collecting time series data, and determining a corresponding training set by performing feature alignment on the time series data;
[0007] Determine a target context vector for prediction corresponding to the time series data in the training set based on a gated attention mechanism in a preset time series prediction model to obtain a corresponding prediction task;
[0008] Decomposing the task based on the number of target prediction heads corresponding to the prediction task to obtain a plurality of subtasks, and determining the initial weights corresponding to the target prediction heads respectively through the prediction window length information respectively corresponding to the target prediction heads in the preset time series prediction model;
[0009] Parallel prediction is performed through a preset weighted loss mechanism, each target prediction head and the prediction window corresponding to each target prediction head, the initial weight and the subtask to determine the target prediction result corresponding to the prediction task, so as to complete the model training operation and trigger the multi-head parallel time series prediction operation based on the obtained trained model.
[0010] Optionally, the collecting time series data and determining a corresponding training set by aligning features of the time series data includes:
[0011] Collect time series data through different channels, and classify and summarize the time series data based on field categories to obtain summary results;
[0012] Based on the summary results, data cleaning and data denoising are performed on the time series data belonging to different field categories respectively to complete corresponding preprocessing operations and obtain preprocessing results;
[0013] Feature alignment is performed based on the preprocessing results to determine a training set.
[0014] Optionally, determining a target context vector for prediction corresponding to the time series data in the training set based on a gated attention mechanism in a preset time series prediction model to obtain a corresponding prediction task includes:
[0015] After the time series data in the training set is input into a preset time series prediction model, a global time-dependent feature corresponding to the time series data in the training set is extracted based on a gated attention mechanism in the preset time series prediction model;
[0016] The target context vector for prediction is determined through the feature aggregation mechanism and the global time-dependent feature to obtain the corresponding prediction task.
[0017] Optionally, the task decomposition based on the number of target prediction heads corresponding to the prediction task to obtain a plurality of subtasks includes:
[0018] Determine the corresponding number of target prediction heads based on the prediction time step length of the prediction task;
[0019] The prediction task is decomposed according to the number of target prediction heads to obtain a plurality of subtasks.
[0020] Optionally, determining the initial weights corresponding to the target prediction heads respectively by using the prediction window length information respectively corresponding to the target prediction heads in the preset time series prediction model includes:
[0021] Acquire prediction window length information and time difference information corresponding to each target prediction head respectively; the time difference information is the time difference between the prediction window of the target prediction head and the current time point in the prediction task;
[0022] The corresponding initial weight is determined based on the prediction window length information and the time difference information respectively corresponding to each of the target prediction heads.
[0023] Optionally, the performing parallel prediction by using a preset weighted loss mechanism, each of the target prediction heads and the prediction windows respectively corresponding to each of the target prediction heads, the initial weights and the subtasks includes:
[0024] Determining the subtasks corresponding to the target prediction heads based on the prediction windows corresponding to the target prediction heads;
[0025] Performing time series predictions of different time periods through each of the target prediction heads and the corresponding subtasks to obtain sub-prediction results corresponding to each of the subtasks;
[0026] The initial weights are dynamically adjusted based on a preset weighted loss mechanism and each of the sub-prediction results, and a target prediction result corresponding to the prediction task is determined according to the obtained adjusted weights and each of the sub-prediction results.
[0027] Optionally, the dynamically adjusting the initial weight based on a preset weighted loss mechanism and each of the sub-prediction results includes:
[0028] The initial weight is dynamically adjusted based on each of the sub-prediction results, the actual value corresponding to each of the sub-prediction results, the loss function corresponding to each of the target prediction heads, and the importance information to obtain the corresponding adjusted weight.
[0029] In a second aspect, the present application provides a multi-head parallel time series prediction device, comprising:
[0030] A data collection module, used to collect time series data and determine a corresponding training set by performing feature alignment on the time series data;
[0031] A task determination module, used to determine a target context vector for prediction corresponding to the time series data in the training set based on a gated attention mechanism in a preset time series prediction model, so as to obtain a corresponding prediction task;
[0032] A task decomposition module, used to decompose the task based on the number of target prediction heads corresponding to the prediction task to obtain multiple subtasks, and determine the initial weights corresponding to each target prediction head respectively through the prediction window length information corresponding to each target prediction head in the preset time series prediction model;
[0033] The multi-head parallel prediction module is used to perform parallel prediction through a preset weighted loss mechanism, each target prediction head and the prediction window corresponding to each target prediction head, the initial weight and the subtask to determine the target prediction result corresponding to the prediction task, so as to complete the model training operation and trigger the multi-head parallel time series prediction operation based on the obtained trained model.
[0034] In a third aspect, the present application provides an electronic device, including:
[0035] Memory, used to store computer programs;
[0036] The processor is used to execute the computer program to implement the steps of the aforementioned multi-head parallel time series prediction method.
[0037] In a fourth aspect, the present application provides a computer-readable storage medium for storing a computer program, which, when executed by a processor, implements the steps of the aforementioned multi-head parallel time series prediction method.
[0038] It can be seen that in the present application, time series data is collected, and the corresponding training set is determined by feature alignment of the time series data; the target context vector for prediction corresponding to the time series data in the training set is determined based on the gated attention mechanism in the preset time series prediction model to obtain the corresponding prediction task; task decomposition is performed based on the number of target prediction heads corresponding to the prediction task to obtain multiple subtasks, and the initial weights corresponding to each target prediction head are determined through the prediction window length information corresponding to each target prediction head in the preset time series prediction model; parallel prediction is performed through the preset weighted loss mechanism, each target prediction head and the prediction window corresponding to each target prediction head, the initial weight and the subtask to determine the target prediction result corresponding to the prediction task, so as to complete the model training operation, and trigger the multi-head parallel time series prediction operation based on the obtained trained model. That is, in this application, first collect time series data and determine the training set, then use the gated attention mechanism in the preset time series prediction model to determine the prediction task, then decompose the task, and determine the initial weight of the corresponding target prediction head, then use each target prediction head to perform parallel predictions of different prediction windows, and use the preset weighted loss mechanism to determine the result corresponding to the prediction task, so as to complete the model training operation, and trigger the multi-head parallel time series prediction operation based on the obtained trained model. In this way, the decline in prediction accuracy due to error accumulation in long-term predictions can be effectively avoided, and the performance, stability, flexibility and adaptability of the time series prediction model can be improved, so that it can better cope with complex and highly volatile time series data prediction tasks. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying creative work.
[0040] Figure 1 A flow chart of a multi-head parallel time series prediction method provided in this application;
[0041] Figure 2 A specific multi-head parallel time series prediction method process architecture diagram provided in this application;
[0042] Figure 3 A schematic diagram of the structure of a multi-head parallel time series prediction device provided in this application;
[0043] Figure 4 A structural diagram of an electronic device provided for this application. DETAILED DESCRIPTION
[0044] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0045] In the current time series prediction scheme, predictions are usually based on autoregressive models or block batch regression models. In the former, as the prediction length increases, the error will gradually accumulate, which will lead to a decrease in the stability and accuracy of the model. The latter usually faces the model's inability to capture long-term trends and short-term fluctuations. Both often show poor prediction results when dealing with long-term, high-volatility and nonlinear complex time series data, especially in complex tasks that need to capture long-term trends and short-term fluctuations at the same time. The performance of these models is difficult to meet high-precision requirements. To this end, the present application provides a multi-head parallel time series prediction scheme, which can effectively avoid the decrease in prediction accuracy due to error accumulation in long-term predictions, and improve the performance, stability, flexibility and adaptability of the time series prediction model, so that it can better cope with complex and highly volatile time series data prediction tasks.
[0046] See also Figure 1 As shown, the embodiment of the present invention discloses a multi-head parallel time series prediction method, including:
[0047] Step S11, collect time series data, and determine the corresponding training set by aligning features of the time series data.
[0048] Specifically, in this embodiment, time series data from various fields including but not limited to energy, environment, finance, etc. are collected through corresponding channels, and the data is analyzed according to the field category so that the model can learn the short-term change characteristics of different fields. That is, first, time series data is collected through different channels, and the time series data is classified and summarized based on the field category to obtain a summary result. Then, based on the summary result, preprocessing operations such as data cleaning and data denoising are performed on the time series data belonging to different field categories to obtain the preprocessing results. Afterwards, feature alignment is performed based on the preprocessing results to determine the training set. In this way, high-quality data input can be provided for subsequent model training.
[0049] Step S12: determine the target context vector for prediction corresponding to the time series data in the training set based on the gated attention mechanism in the preset time series prediction model to obtain the corresponding prediction task.
[0050] In this embodiment, when the pre-built preset time series prediction model is trained based on the determined training set, the gated attention mechanism in the model is first used to preliminarily capture the long-term and short-term dependencies in the time series, capture the patterns, trends and complex relationships in the data, and aggregate the output features of the encoder to generate a context vector for prediction. That is, after the time series data in the training set is input into the preset time series prediction model, the global time dependency features corresponding to the time series data in the training set are extracted based on the gated attention mechanism in the preset time series prediction model; the target context vector for prediction is determined through the feature aggregation mechanism and the global time dependency features to obtain the corresponding prediction task.
[0051] It should be understood that, regarding the gated attention mechanism, the gated attention mechanism can be used in this embodiment to adjust the attention weight so that the preset time series prediction model can dynamically select different focus points according to the current input time series data. In other words, this mechanism enables the model to flexibly pay attention to short-term fluctuations or long-term trends, avoiding excessive attention to unimportant information. The gated weight calculation formula is shown in the following formula (1), which acts on each query (i.e. Query) and key (i.e. Key) combination to generate a gating weight .
[0052] (1);
[0053] in, is the sigmoid activation function, and are the weights and biases obtained from training, represents the concatenation of query and key, is the value output by the gating mechanism, which is between 0 and 1. In this embodiment, the gating weight is combined with the standard attention weight through formula (2) to obtain the weighted attention matrix The final output As shown in formula (3), the gating mechanism controls the attention weight of each time step, allowing the model to adaptively adjust the attention to different time steps when processing time series data according to the current prediction target and the characteristics of the time step, so as to better capture short-term fluctuations and long-term trends.
[0054] (2);
[0055] (3);
[0056] In the formula, is part of the attention mechanism, (Query, query), (Key, key) and (Vaule, value) are three matrices representing query, key, and value vectors, respectively, generated by applying different weight matrices to the same input embedding; represents transpose; for The dimension of the matrix plays a role in scaling.
[0057] Step S13, performing task decomposition based on the number of target prediction heads corresponding to the prediction task to obtain multiple subtasks, and determining the initial weights corresponding to each target prediction head through the prediction window length information corresponding to each target prediction head in the preset time series prediction model.
[0058] In this embodiment, it should be pointed out that the prediction output part in the preset time series prediction model adopts a multi-head parallel mechanism, and controls the number of prediction heads according to actual needs such as the prediction time step length of the task and the importance of short-term capture, so as to achieve dynamic controllability. Then, based on the number of prediction heads, the prediction task is decomposed and assigned to multiple independent prediction heads to predict data in different time periods in parallel, fundamentally avoiding the problem of gradual error accumulation in traditional autoregressive methods. Among them, regarding task decomposition, first determine the corresponding target number of prediction heads based on the prediction time step length of the prediction task, and then decompose the prediction task according to the target number of prediction heads to obtain multiple subtasks. Subtasks are then assigned, and each subtask is independently responsible for by a prediction head.
[0059] At the same time, in this embodiment, it is also necessary to obtain the prediction window length information and time difference information corresponding to each target prediction head; the time difference information is the time difference between the prediction window of the target prediction head and the current time point in the prediction task; based on the prediction window length information and time difference information corresponding to each target prediction head, the corresponding initial weight is determined. Among them, each prediction head is independent of each other, and each prediction head is responsible for predicting time windows of different lengths. For example, the prediction window of the first prediction head (that is, the closest to the current time point) is the shortest, and the windows of subsequent prediction heads increase successively. The windows of each prediction head do not overlap, thereby covering the long and short-term prediction lengths, thereby providing better prediction results for the short-term prediction of the model, while avoiding the error accumulation problem in the traditional autoregressive model and ensuring the controllability of the prediction length. As shown in formula (4), assuming that the prediction window of the last prediction head is The length is , then the prediction window length information from the last prediction head to the first prediction head is 4 to achieve a decrease, is the prediction window length attenuation value between adjacent prediction heads. If the total prediction length is set to , a total of If there are prediction heads, then as shown in formula (5), the sum of the total prediction lengths of all prediction heads should be equal to , Indicates The prediction window length of the prediction head.
[0060] (4);
[0061] (5);
[0062] Step S14, performing parallel prediction through a preset weighted loss mechanism, each of the target prediction heads and the prediction windows corresponding to each of the target prediction heads, the initial weights and the subtasks to determine the target prediction result corresponding to the prediction task, so as to complete the model training operation, and trigger a multi-head parallel time series prediction operation based on the obtained trained model.
[0063] In this embodiment, when performing multi-head parallel prediction, the subtasks corresponding to each target prediction head are determined based on the prediction windows corresponding to each target prediction head; time series predictions of different time periods are performed through each target prediction head and the corresponding subtask to obtain sub-prediction results corresponding to each subtask; the initial weights are dynamically adjusted based on the preset weighted loss mechanism and each sub-prediction result, and the target prediction result corresponding to the prediction task is determined based on the adjusted weight and each sub-prediction result. That is to say, in this embodiment, when making predictions based on the preset time series prediction model, the preset weighted loss mechanism in the model is used to add dynamic loss weights to each prediction head. For example, the closer the prediction head is to the current time point, the greater the loss weight, thereby achieving the improvement of the accuracy of the model for short-term predictions, and providing long-term change trends within an acceptable range.
[0064] Furthermore, regarding the application of the preset weighted loss mechanism, the initial weights are dynamically adjusted based on each sub-prediction result, the actual value corresponding to each sub-prediction result, the loss function corresponding to each target prediction head, and the importance information to obtain the corresponding adjusted weights. Specifically, there is a weighted loss function in the mechanism, and the loss function of each prediction head in the weighted loss function can be weighted according to the length of its predicted time window and its importance. As shown in formula (6), Represents the target number of prediction heads for parallel prediction; It is The weight of the prediction head, represents the sum of the weight values of all dynamic prediction heads; as shown in formula (7), where is a decay coefficient whose value range is (0, 1) and is used to control the rate of weight decay. Each prediction head has a separate weight, which is mainly based on the order of the prediction heads. The closer the prediction window is to the current time point, the higher the weight of the prediction head, thereby improving the model's ability to capture short-term fluctuations; It is The prediction loss of a prediction head is shown in formula (8). It is Prediction head to time point The predicted value of It's time point The true value of It is The prediction time window that a prediction head is responsible for.
[0065] (6);
[0066] (7);
[0067] (8).
[0068] In this way, the model is further trained and fine-tuned based on the loss function, making the model better at capturing short-term fluctuations and providing relatively long-term trend forecasts. The trained model can remove the reduction in prediction accuracy caused by cumulative errors. In addition, the dynamic adjustment of the number of prediction heads can make the prediction length of the model customizable, realizing a wider range of application scenarios.
[0069] It can be seen that in an embodiment of the present application, time series data is collected, and a corresponding training set is determined by feature alignment of the time series data; a target context vector for prediction corresponding to the time series data in the training set is determined based on a gated attention mechanism in a preset time series prediction model to obtain a corresponding prediction task; task decomposition is performed based on the number of target prediction heads corresponding to the prediction task to obtain a plurality of subtasks, and the initial weights corresponding to each target prediction head are determined through prediction window length information corresponding to each target prediction head in the preset time series prediction model; parallel prediction is performed through a preset weighted loss mechanism, each target prediction head and a prediction window corresponding to each target prediction head, the initial weights and the subtasks to determine the target prediction result corresponding to the prediction task, so as to complete the model training operation, and trigger a multi-head parallel time series prediction operation based on the obtained trained model. That is, in this application, first collect time series data and determine the training set, then use the gated attention mechanism in the preset time series prediction model to determine the prediction task, then decompose the task, and determine the initial weight of the corresponding target prediction head, then use each target prediction head to perform parallel predictions of different prediction windows, and use the preset weighted loss mechanism to determine the result corresponding to the prediction task, so as to complete the model training operation, and trigger the multi-head parallel time series prediction operation based on the obtained trained model. In this way, the decline in prediction accuracy due to error accumulation in long-term predictions can be effectively avoided, and the performance, stability, flexibility and adaptability of the time series prediction model can be improved, so that it can better cope with complex and highly volatile time series data prediction tasks.
[0070] Combine the following Figure 2 The process architecture diagram disclosed in the specification specifically illustrates the technical solution of the embodiment of the present application.
[0071] like Figure 2As shown, taking the model implementation process of load monitoring prediction as an example, in a specific implementation, the load time series data of the actual application scenario is monitored and collected through the data collection and processing module, and the data is cleaned, denoised and feature aligned according to the time frequency classification. Then, the time series prediction model is built based on the transformer structure of the model, and the gated attention mechanism is added as the encoding feature aggregation module on the original basis to further improve the model's ability to capture long and short dependencies, and the calculation efficiency is increased by adding a multi-head decoder module to improve the prediction efficiency of the model; a multi-head parallel mechanism is adopted in the prediction output part of the model, and the number of prediction heads is controlled according to actual needs such as prediction length and short-term capture importance, so as to achieve dynamic control. Then, based on the collection of a large amount of data, the built time series prediction model is trained to allow the model to learn the characteristic structure and intrinsic trend of the data. This process adopts a new prediction loss function (i.e., the weighted loss function in the model), adds a certain weight to the loss of each prediction head, and implements a weight attenuation mechanism based on the window position predicted by the prediction head, so that the model pays more attention to the recent fluctuation trend. This is because the failure of short-term change prediction in the actual load prediction and control system will have a greater impact. Finally, after multiple rounds of weighted loss training, the model can achieve fast reasoning for small sample applications and complete load monitoring and timing prediction tasks.
[0072] See also Figure 3 As shown, the embodiment of the present application also discloses a multi-head parallel time series prediction device, including:
[0073] A data collection module 11 is used to collect time series data and determine a corresponding training set by performing feature alignment on the time series data;
[0074] A task determination module 12 is used to determine a target context vector for prediction corresponding to the time series data in the training set based on a gated attention mechanism in a preset time series prediction model to obtain a corresponding prediction task;
[0075] A task decomposition module 13 is used to decompose the task based on the number of target prediction heads corresponding to the prediction task to obtain multiple subtasks, and determine the initial weights corresponding to the target prediction heads respectively through the prediction window length information corresponding to the target prediction heads in the preset time series prediction model;
[0076] The multi-head parallel prediction module 14 is used to perform parallel prediction through a preset weighted loss mechanism, each target prediction head and the prediction window corresponding to each target prediction head, the initial weight and the subtask to determine the target prediction result corresponding to the prediction task, so as to complete the model training operation and trigger the multi-head parallel time series prediction operation based on the obtained trained model.
[0077] Among them, for more specific working processes of the above-mentioned modules, please refer to the corresponding contents disclosed in the aforementioned embodiments, which will not be repeated here.
[0078] It can be seen that in this application, first, the time series data is collected and the training set is determined, and then the prediction task is determined by using the gated attention mechanism in the preset time series prediction model, and then the task is decomposed and the initial weight of the corresponding target prediction head is determined. Then, each target prediction head is used to perform parallel predictions of different prediction windows, and the preset weighted loss mechanism is used to determine the result corresponding to the prediction task, so as to complete the model training operation and trigger the multi-head parallel time series prediction operation based on the obtained trained model. In this way, the decline in prediction accuracy due to error accumulation in long-term predictions can be effectively avoided, and the performance, stability, flexibility and adaptability of the time series prediction model can be improved, so that it can better cope with complex and highly volatile time series data prediction tasks.
[0079] In some specific embodiments, the data collection module 11 can be specifically used to collect time series data through different channels, and classify and summarize the time series data based on field categories to obtain summary results; based on the summary results, the time series data belonging to different field categories are respectively cleaned and denoised to complete corresponding preprocessing operations and obtain preprocessing results; feature alignment is performed based on the preprocessing results to determine the training set.
[0080] In some specific embodiments, the task determination module 12 can be specifically used to extract the global time-dependent features corresponding to the time series data in the training set based on the gated attention mechanism in the preset time series prediction model after the time series data in the training set is input into the preset time series prediction model; determine the target context vector for prediction through the feature aggregation mechanism and the global time-dependent features to obtain the corresponding prediction task.
[0081] In some specific embodiments, the task decomposition module 13 can be specifically used to determine the corresponding target prediction head number based on the prediction time step length of the prediction task; and decompose the prediction task according to the target prediction head number to obtain multiple subtasks.
[0082] In some specific embodiments, the task decomposition module 13 can be specifically used to obtain prediction window length information and time difference information corresponding to each target prediction head; the time difference information is the time difference between the prediction window of the target prediction head and the current time point in the prediction task; based on the prediction window length information and the time difference information corresponding to each target prediction head, the corresponding initial weight is determined.
[0083] In some specific embodiments, the multi-head parallel prediction module 14 can be specifically used to determine the subtasks corresponding to each target prediction head based on the prediction windows corresponding to each target prediction head; perform time series predictions for different time periods through each target prediction head and the corresponding subtask to obtain sub-prediction results corresponding to each subtask; dynamically adjust the initial weights based on a preset weighted loss mechanism and each sub-prediction result, and determine the target prediction result corresponding to the prediction task based on the adjusted weights and each sub-prediction result.
[0084] In some specific embodiments, the multi-head parallel prediction module 14 can be used to dynamically adjust the initial weights based on each of the sub-prediction results, the actual values corresponding to each of the sub-prediction results, the loss functions corresponding to each of the target prediction heads, and the importance information to obtain the corresponding adjusted weights.
[0085] Furthermore, the present application also discloses an electronic device. Figure 4 This is a structural diagram of an electronic device 20 according to an exemplary embodiment. The content in the diagram cannot be regarded as any limitation on the scope of use of the present application.
[0086] Figure 4 A schematic diagram of the structure of an electronic device 20 provided in an embodiment of the present application. The electronic device 20 may specifically include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. The memory 22 is used to store a computer program, which is loaded and executed by the processor 21 to implement the relevant steps in the multi-head parallel time series prediction method disclosed in any of the aforementioned embodiments. In addition, the electronic device 20 in this embodiment may specifically be an electronic computer.
[0087] In this embodiment, the power supply 23 is used to provide working voltage for each hardware device on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and the external device, and the communication protocol it follows is any communication protocol that can be applied to the technical solution of the present application, and is not specifically limited here; the input and output interface 25 is used to obtain external input data or output data to the outside world, and its specific interface type can be selected according to specific application needs and is not specifically limited here.
[0088] In addition, the memory 22, as a carrier for storing resources, can be a read-only memory, a random access memory, a disk or an optical disk, etc. The resources stored thereon can include an operating system 221, a computer program 222, etc., and the storage method can be temporary storage or permanent storage.
[0089] The operating system 221 is used to manage and control the hardware devices and computer program 222 on the electronic device 20, and can be Windows Server, Netware, Unix, Linux, etc. In addition to including a computer program that can be used to complete the multi-head parallel time series prediction method performed by the electronic device 20 disclosed in any of the aforementioned embodiments, the computer program 222 can further include a computer program that can be used to complete other specific tasks.
[0090] Furthermore, the present application also discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, the multi-head parallel time series prediction method disclosed above is implemented. The specific steps of the method can refer to the corresponding contents disclosed in the above embodiments, and will not be repeated here.
[0091] In this specification, each embodiment is described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the embodiments can be referred to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part.
[0092] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in the above description according to function. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0093] The steps of the method or algorithm described in conjunction with the embodiments disclosed herein may be implemented directly using hardware, a software module executed by a processor, or a combination of the two. The software module may be placed in a random access memory (RAM), a memory, a read-only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.
[0094] Finally, it should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the presence of other identical elements in the process, method, article or device including the elements.
[0095] The technical solution provided by the present application is introduced in detail above. Specific examples are used in this article to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea. At the same time, for general technicians in this field, according to the idea of the present application, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on the present application.
Claims
1. A multi-head parallel time series prediction method, characterized in that: include: Collecting time series data, and determining a corresponding training set by performing feature alignment on the time series data; Determine a target context vector for prediction corresponding to the time series data in the training set based on a gated attention mechanism in a preset time series prediction model to obtain a corresponding prediction task; Decomposing the task based on the number of target prediction heads corresponding to the prediction task to obtain a plurality of subtasks, and determining the initial weights corresponding to the target prediction heads respectively through the prediction window length information respectively corresponding to the target prediction heads in the preset time series prediction model; Parallel prediction is performed through a preset weighted loss mechanism, each target prediction head and the prediction window corresponding to each target prediction head, the initial weight and the subtask to determine the target prediction result corresponding to the prediction task, so as to complete the model training operation and trigger the multi-head parallel time series prediction operation based on the obtained trained model.
2. The multi-head parallel time series prediction method according to claim 1, characterized in that: The collecting of time series data and determining a corresponding training set by aligning features of the time series data includes: Collect time series data through different channels, and classify and summarize the time series data based on field categories to obtain summary results; Based on the summary results, data cleaning and data denoising are performed on the time series data belonging to different field categories respectively to complete corresponding preprocessing operations and obtain preprocessing results; Feature alignment is performed based on the preprocessing results to determine a training set.
3. The multi-head parallel time series prediction method according to claim 1, characterized in that: The determining, based on the gated attention mechanism in the preset time series prediction model, a target context vector for prediction corresponding to the time series data in the training set to obtain a corresponding prediction task includes: After the time series data in the training set is input into a preset time series prediction model, a global time-dependent feature corresponding to the time series data in the training set is extracted based on a gated attention mechanism in the preset time series prediction model; The target context vector for prediction is determined through the feature aggregation mechanism and the global time-dependent feature to obtain the corresponding prediction task.
4. The multi-head parallel time series prediction method according to claim 1, characterized in that: The task is decomposed based on the number of target prediction heads corresponding to the prediction task to obtain a plurality of subtasks, including: Determine the corresponding number of target prediction heads based on the prediction time step length of the prediction task; The prediction task is decomposed according to the number of target prediction heads to obtain a plurality of subtasks.
5. The multi-head parallel time series prediction method according to claim 1, characterized in that: The determining of the initial weights respectively corresponding to the target prediction heads by using the prediction window length information respectively corresponding to the target prediction heads in the preset time series prediction model comprises: Acquire prediction window length information and time difference information corresponding to each target prediction head respectively; the time difference information is the time difference between the prediction window of the target prediction head and the current time point in the prediction task; The corresponding initial weight is determined based on the prediction window length information and the time difference information respectively corresponding to each of the target prediction heads.
6. The multi-head parallel time series prediction method according to any one of claims 1 to 5, characterized in that: The performing parallel prediction by using a preset weighted loss mechanism, each of the target prediction heads and the prediction windows respectively corresponding to each of the target prediction heads, the initial weights and the subtasks includes: Determining the subtasks corresponding to the target prediction heads based on the prediction windows corresponding to the target prediction heads; Performing time series predictions of different time periods through each of the target prediction heads and the corresponding subtasks to obtain sub-prediction results corresponding to each of the subtasks; The initial weights are dynamically adjusted based on a preset weighted loss mechanism and each of the sub-prediction results, and a target prediction result corresponding to the prediction task is determined according to the obtained adjusted weights and each of the sub-prediction results.
7. The multi-head parallel time series prediction method according to claim 6, characterized in that: The dynamically adjusting the initial weight based on the preset weighted loss mechanism and each of the sub-prediction results includes: The initial weight is dynamically adjusted based on each of the sub-prediction results, the actual value corresponding to each of the sub-prediction results, the loss function corresponding to each of the target prediction heads, and the importance information to obtain the corresponding adjusted weight.
8. A multi-head parallel time series prediction device, characterized in that: include: A data collection module, used to collect time series data and determine a corresponding training set by performing feature alignment on the time series data; A task determination module, used to determine a target context vector for prediction corresponding to the time series data in the training set based on a gated attention mechanism in a preset time series prediction model, so as to obtain a corresponding prediction task; A task decomposition module, used to decompose the task based on the number of target prediction heads corresponding to the prediction task to obtain multiple subtasks, and determine the initial weights corresponding to each target prediction head respectively through the prediction window length information corresponding to each target prediction head in the preset time series prediction model; The multi-head parallel prediction module is used to perform parallel prediction through a preset weighted loss mechanism, each target prediction head and the prediction window corresponding to each target prediction head, the initial weight and the subtask to determine the target prediction result corresponding to the prediction task, so as to complete the model training operation and trigger the multi-head parallel time series prediction operation based on the obtained trained model.
9. An electronic device, characterized in that: include: Memory, used to store computer programs; A processor, configured to execute the computer program to implement the multi-head parallel time series prediction method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: Used to store a computer program, which, when executed by a processor, implements the multi-head parallel time series prediction method as described in any one of claims 1 to 7.
Citation Information
Cited By
Traffic prediction method and device, electronic equipment and storage medium
CN121486216A