Prediction model construction method and apparatus, and prediction model calling method and apparatus
By interacting information from different dimensions and suppressing noise in the existing technology in the construction of the multivariate time series prediction model, the problem of inaccurate multivariate time series prediction in the prior art is solved, and higher prediction accuracy is achieved.
Patent Information
- Application Number
- PCT/CN2024/108811
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-13
- Filing Date
- 2024-07-31
- Publication Date
- 2025-05-22
AI Technical Summary
The prior art fails to effectively suppress noise when predicting multiple time series, resulting in inaccurate predictions.
In the process of constructing the prediction model, information interaction is carried out from different dimensions, information content in the multivariate time series is enriched, and during the information interaction, noise generated between dimensions and noise inherent in the dimension itself is suppressed.
More accurate prediction of multiple time series is achieved, providing a reliable data basis for making reasonable decisions based on accurate prediction results.
Smart Images

Figure CN2024108811_22052025_PF_FP_ABST
Abstract
Description
A method and device for constructing and calling a prediction model
[0001] This application claims priority to the Chinese patent application filed with the China Patent Office on November 13, 2023, with application number 202311514572.7 and invention name “A method and device for constructing and calling a prediction model”, the entire contents of which are incorporated by reference into this application. Technical Field
[0002] The present application relates to the field of computer technology, and in particular to a method for constructing a prediction model, a method for calling a prediction model, and a device. Background Art
[0003] Multivariate time series (also called multivariate time series or multi-feature time series) refer to data composed of multiple features collected in chronological order. For example, if device A includes temperature sensor 1, vibration sensor 10, and each sensor collects data once every hour for 24 hours, the resulting multivariate time series can be represented as a 24x10 matrix A. In matrix A, aij refers to the data collected by sensor j at sampling time i on device A, where i = 1, 2, ..., 24, and j = 1, 2, ..., 10. The "elements" in a multivariate time series are often correlated, making multivariate time series prediction crucial for practical production and life. Multivariate time series prediction involves studying the temporal variations of multiple related variables to predict their future behavior. The more accurate the prediction, the more rational decisions can be made based on it. Therefore, achieving accurate multivariate time series prediction is a worthy research topic.
[0004] Current predictions for multivariate time series fail to address noise suppression, resulting in inaccurate predictions. Even when noise interference is considered and suppressed, such as with the Dlinear model, it is still necessary to separate the different variables and construct separate prediction models. Such prediction models cannot accurately predict multivariate time series. Furthermore, they only suppress the noise caused by interactions between different variables by modeling them separately, but cannot suppress the noise inherent in each variable, resulting in inaccurate predictions.
[0005] Based on this, there is an urgent need to provide a prediction model that can accurately predict the situation of multiple "elements" of multivariate time series in the future, so that it is possible to make reasonable decisions based on accurate prediction results.
[0006] Summary of the Invention
[0007] Based on this, the present application provides a method and device for constructing and calling a prediction model. The constructed prediction model can accurately predict multivariate time series and provide a reliable data basis for reasonable decision-making.
[0008] In a first aspect, the present application provides a method for constructing a prediction model, wherein the prediction model in the method is used to predict a multivariate time series. The construction process of the prediction model may include: on the one hand, information interaction from different dimensions to enrich the information content in the multivariate time series; on the other hand, in the process of information interaction, suppressing the noise generated between the dimensions and the noise brought by the dimensions themselves, so as to ensure that the prediction accuracy of the constructed prediction model is higher. As an example, the construction method of the prediction model may include: first, obtaining multiple first training samples, each of the multiple first training samples includes a first multivariate time series and a second multivariate time series, and the second multivariate time series is a multivariate time series collected after the first multivariate time series in the first training sample to which it belongs; then, based on the first multivariate time series, the second multivariate time series, an initial model and an information bottleneck constraint, obtaining a prediction model, the initial model including a model for information interaction in at least one dimension, the information bottleneck constraint can at least be used to suppress the noise in the first multivariate time series, and the prediction model is a trained model for predicting a multivariate time series after a known multivariate time series. In this way, during the prediction model training process, sufficient information interaction is carried out on the multivariate time series in the training samples, and various noises are suppressed based on the information bottleneck constraint conditions, so that the constructed prediction model can predict the multivariate time series more accurately.
[0009] In one possible implementation, the information bottleneck constraint can also be used to enhance the learning of useful information in multivariate time series, thereby enabling the further construction of a prediction model with more accurate prediction effects, making it possible to make more accurate predictions of multivariate time series based on the prediction model.
[0010] In one possible implementation, taking at least one dimension including a first dimension as an example, the process of obtaining a prediction model based on the first multivariate time series, the second multivariate time series, the initial model, and the information bottleneck constraint during the construction of the prediction model is illustratively described. The process of obtaining the prediction model based on the first multivariate time series, the second multivariate time series, the initial model, and the information bottleneck constraint may include updating the initial model N times in the first dimension, where N is an integer greater than or equal to 1.
[0011] As an example, since the initial model is updated N times in the first dimension, the updating process of the initial model in the first dimension is similar each time. The updating process of the initial model in the first dimension may include: first, inputting a first multivariate time series into the initial model and outputting a third multivariate time series; then, obtaining first mutual information based on a first intermediate result of the first multivariate time series, wherein the first intermediate result is a high-dimensional mapping of the first multivariate time series in the first dimension, the first mutual information being used to indicate the impact of noise from dimensions other than the first dimension on the prediction process, and the first mutual information being used to indicate useful information learned from dimensions other than the first dimension; obtaining second mutual information based on the second multivariate time series and the first intermediate result, the second mutual information being used to indicate the connection between representation and prediction; then, determining a first loss based on the first mutual information and the second mutual information, and updating the initial model based on the first loss to obtain an updated initial model; and updating the updated initial model based on the second and third multivariate time series. Alternatively, the initial model may be first updated based on the second and third multivariate time series to obtain an updated initial model; and then, updating the updated initial model based on the first and second mutual information. In this implementation, after performing N updates to the initial model in the first dimension, the initial model after the Nth update to the initial model in the first dimension can be used as the constructed prediction model. In this way, the impact of noise from other features on the training process can be suppressed, thereby obtaining a more accurate prediction model.
[0012] For example, the initial model may include a first encoding module and a first decoding module that exchange information in a first dimension, with the output of the first encoding module serving as the input to the first decoding module. Inputting the first multivariate time series into the initial model and outputting a third multivariate time series may include: inputting the first multivariate time series into the first encoding module in the initial model, with the output of the first encoding module serving as the first intermediate result; and inputting the first intermediate result into the first decoding module, with the first decoding module outputting the third multivariate time series.
[0013] In another possible implementation, taking the example of at least one dimension including a second dimension in addition to the first dimension, the process of obtaining a prediction model based on the first multivariate time series, the second multivariate time series, the initial model, and the information bottleneck constraint during the construction of the prediction model is exemplified. The process of obtaining the prediction model based on the first multivariate time series, the second multivariate time series, the initial model, and the information bottleneck constraint may include N updates to the initial model in the first dimension and M updates to the initial model in the second dimension, where M is an integer greater than or equal to 1. Regarding the N updates to the initial model in the first dimension, see the above implementation.
[0014] As an example, since the initial model is updated M times in the second dimension, the updating process of the initial model in the second dimension is similar each time. The updating process of the initial model in the second dimension can include: first, inputting the first multivariate time series into the initial model and outputting a fourth multivariate time series; then, obtaining third mutual information based on the first multivariate time series and the output second intermediate result, wherein the second intermediate result is a high-dimensional mapping of the first multivariate time series in the second dimension, and the third mutual information is used to indicate the impact of noise from various dimensions on the prediction process, wherein the dimensions include the second dimension; obtaining fourth mutual information based on the second multivariate time series and the second intermediate result, and the fourth mutual information is used to indicate the connection between representation and prediction; then, determining a second loss based on the third mutual information and the fourth mutual information, and updating the initial model based on the second loss to obtain an updated initial model; and updating the updated initial model based on the second multivariate time series and the fourth multivariate time series. Alternatively, the initial model can be first updated based on the second multivariate time series and the fourth multivariate time series to obtain an updated initial model; and then updating the updated initial model based on the third mutual information and the fourth mutual information. In this implementation, after performing N updates to the initial model in the first dimension and M updates to the initial model in the second dimension, the final updated initial model serves as the constructed prediction model. This approach can suppress the impact of noise in each dimension on the training process, achieving more comprehensive noise suppression and ultimately a more accurate prediction model.
[0015] For example, the initial model may further include a second encoding module and a second decoding module for information exchange in the second dimension, with the output of the second encoding module serving as the input of the second decoding module. Then, inputting the first multivariate time series into the initial model and outputting a fourth multivariate time series may include: inputting the first multivariate time series into the second encoding module in the initial model, with the output of the second encoding module serving as the aforementioned second intermediate result; and inputting the second intermediate result into the second decoding module, with the second decoding module outputting the fourth multivariate time series.
[0016] In another possible implementation, taking at least one dimension including a second dimension as an example, the process of obtaining a prediction model based on the first multivariate time series, the second multivariate time series, the initial model, and the information bottleneck constraint during the construction of the prediction model is exemplified. The process of obtaining the prediction model based on the first multivariate time series, the second multivariate time series, the initial model, and the information bottleneck constraint may include L updates to the initial model in the second dimension, where L is an integer greater than or equal to 1. The L updates to the initial model in the second dimension refer to the above implementation and are not further described here.
[0017] As an example, in the step of obtaining a prediction model based on the first multivariate time series, the second multivariate time series, the initial model, and the information bottleneck constraint in the process of constructing the prediction model, the update of the initial model in the second dimension and the update of the initial model in the first dimension can be performed according to a preset rule, and the initial model after the previous update is used as the initial model in the next update operation. Wherein, the preset rules may include, but are not limited to: the update of the initial model in the second dimension and the update of the initial model in the first dimension can be performed alternately, and the alternation frequency can be 1 or any other positive integer. Taking the alternation frequency equal to 1 as an example, after each update of the initial model in the second dimension is performed, the update of the initial model in the first dimension can be performed once; or, after performing M updates to the initial model in the second dimension, the update of the initial model in the first dimension can be performed N times; or, after performing N updates to the initial model in the first dimension, the update of the initial model in the second dimension can be performed M times.
[0018] As an example, the first dimension may be a variable dimension, a feature dimension, or a channel dimension, and the second dimension may be a time dimension or a frequency dimension.
[0019] In one possible implementation, after the prediction model is constructed, the method may further include: obtaining a sixth multivariate time series based on the fifth multivariate time series and the prediction model. The fifth multivariate time series is the multivariate time series to be predicted, and the sixth multivariate time series is the prediction result of the prediction model for a multivariate time series that may appear after the fifth multivariate time series. If the prediction performance of the prediction model is sufficiently good, then the multivariate time series collected after the fifth multivariate time series is sufficiently similar to the sixth multivariate time series.
[0020] As an example, the fifth multivariate time series and the first multivariate time series are used as inputs of the prediction model in the prediction stage and the training stage respectively. The scale of the fifth multivariate time series can be the same as that of the first multivariate time series. If the prediction is made directly based on the constructed prediction model, the scale of the sixth multivariate time series is the same as that of the second multivariate time series. If there are specific requirements for the scale of the output multivariate time series, then the constructed prediction model can be adjusted. In this case, the scale of the sixth multivariate time series can also be different from that of the second multivariate time series.
[0021] In the second aspect, the present application also provides a method for calling a prediction model, in which, after a first multivariate time series to be predicted is generated, a prediction model constructed based on the method provided in an embodiment of the present application can be obtained, and the prediction model is used to predict the first multivariate time series to obtain a second multivariate time series. As an example, the method may include: first, obtaining a prediction model, wherein the prediction model is a trained model for predicting a multivariate time series, and the training process of the prediction model is constrained by an information bottleneck constraint, and the information bottleneck constraint is used to suppress the multivariate time series in the training sample; then, based on the prediction model and the first multivariate time series, a second multivariate time series is obtained, and the second multivariate time series is the prediction result of the prediction model for the multivariate time series that appears after the first multivariate time series. In this way, since the prediction model can fully interact with the first multivariate time series and has the function of suppressing various noises, the prediction result (i.e., the second multivariate time series) obtained by using the prediction model for the first multivariate time series is accurate and close to the real multivariate time series that appears after the first multivariate time series, providing a reliable basis for making reasonable decisions based on the prediction results.
[0022] In one possible implementation, the information bottleneck constraint can also be used to enhance the learning of useful information in multivariate time series, thereby making it possible to make more accurate predictions of multivariate time series based on the prediction model.
[0023] It should be noted that for the relevant description of the second aspect, please refer to the corresponding description of the first aspect. For example, for the specific implementation method and construction process of the prediction model in the second aspect, please refer to the corresponding description of the first aspect.
[0024] In a third aspect, the present application also provides a device for constructing a prediction model, which may include: an acquisition unit and a training unit. The acquisition unit is used to acquire multiple first training samples, each of the multiple first training samples includes a first multivariate time series and a second multivariate time series, and the second multivariate time series is a multivariate time series collected after the first multivariate time series in the first training sample to which it belongs; the training unit is used to obtain a prediction model based on the first multivariate time series, the second multivariate time series, an initial model, and an information bottleneck constraint. The initial model includes a model for information interaction in at least one dimension, and the information bottleneck constraint is used to suppress noise in the first multivariate time series. The prediction model is a trained model for predicting a multivariate time series after a known multivariate time series.
[0025] In one possible implementation, the information bottleneck constraint can also be used to enhance the learning of useful information in multivariate time series.
[0026] In a possible implementation, the at least one dimension includes a first dimension, and the training unit is specifically used to: perform N updates on the initial model in the first dimension, where N is an integer greater than or equal to 1.
[0027] As an example, the training unit performs N updates on the initial model in the first dimension, and the updating process of the initial model in the first dimension includes: inputting the first multivariate time series into the initial model and outputting a third multivariate time series, the initial model including a first encoding module and a first decoding module for information interaction in the first dimension, and the output of the first encoding module serves as the input of the first decoding module; obtaining first mutual information based on the first multivariate time series and the first intermediate result output by the first encoding module, the first intermediate result being a high-dimensional mapping of the first multivariate time series in the first dimension, the first mutual information being used to indicate the impact of noise from dimensions other than the first dimension on the prediction process, and the first mutual information being also used to indicate useful information learned from dimensions other than the first dimension; obtaining second mutual information based on the second multivariate time series and the first intermediate result, the second mutual information being used to indicate the connection between representation and prediction; updating the initial model based on the first mutual information and the second mutual information to obtain an updated initial model; and updating the updated initial model based on the first multivariate time series and the third multivariate time series.
[0028] In another possible implementation, the at least one dimension also includes a second dimension, and the training unit is further used to: perform M updates on the initial model in the second dimension, where M is an integer greater than or equal to 1.
[0029] As an example, the training unit performs M updates on the initial model in the second dimension, and the updating process of the initial model in the second dimension includes: inputting the first multivariate time series into the initial model and outputting a fourth multivariate time series, the initial model also including a second encoding module and a second decoding module for information interaction in the second dimension, and the output of the second encoding module serves as the input of the second decoding module; obtaining third mutual information based on the first multivariate time series and the second intermediate result output by the second encoding module, the second intermediate result being a high-dimensional mapping of the first multivariate time series in the second dimension, and the third mutual information being used to indicate the impact of noise from each dimension on the prediction process, the dimensions including the second dimension; obtaining fourth mutual information based on the second multivariate time series and the second intermediate result, the fourth mutual information being used to indicate the connection between representation and prediction; updating the initial model based on the third mutual information and the fourth mutual information to obtain an updated initial model; and updating the updated initial model based on the first multivariate time series and the fourth multivariate time series.
[0030] As an example, the updating of the initial model in the second dimension and the updating of the initial model in the first dimension are performed alternately according to preset rules, and the initial model after the previous update is used as the initial model in the next update operation.
[0031] As an example, the first dimension is a feature dimension or a channel dimension, and the second dimension is a time dimension or a frequency dimension.
[0032] In another possible implementation, the at least one dimension includes a second dimension, and the training unit is specifically used to: perform L updates on the initial model in the second dimension, where L is an integer greater than or equal to 1.
[0033] In one possible implementation, the apparatus may further include a prediction unit configured to obtain a sixth multivariate time series based on the fifth multivariate time series and the prediction model, where the sixth multivariate time series is a prediction result of the prediction model for a multivariate time series that occurs after the fifth multivariate time series.
[0034] As an example, the fifth multivariate time series has the same scale as the first multivariate time series, and the sixth multivariate time series has the same scale as or a different scale than the second multivariate time series.
[0035] It should be noted that for the relevant description of the device for constructing the prediction model of the third aspect, please refer to the corresponding description of the first aspect.
[0036] In a fourth aspect, the present application further provides a device for calling a prediction model, which may include: an acquisition unit and a prediction unit. The acquisition unit is used to acquire a prediction model, wherein the prediction model is a trained model for predicting a multivariate time series, and the training process of the prediction model is constrained by an information bottleneck constraint, and the information bottleneck constraint is used to suppress the noise of the multivariate time series in the training sample; the prediction unit is used to obtain a second multivariate time series based on the prediction model and the first multivariate time series, wherein the second multivariate time series is the prediction result of the prediction model for the multivariate time series that appears after the first multivariate time series.
[0037] In one possible implementation, the information bottleneck constraint can also be used to enhance the learning of useful information in multivariate time series.
[0038] It should be noted that for the relevant description of the calling device of the prediction model of the fourth aspect, please refer to the corresponding description of the second aspect.
[0039] In a fifth aspect, the present application provides a communication device, comprising a processor and a memory; the processor is used to execute instructions stored in the memory so that the communication device implements the methods corresponding to the first aspect, the second aspect and their possible implementation methods.
[0040] In a sixth aspect, the present application further provides a storage medium storing instructions, which, when executed on a processor, implement methods corresponding to the first aspect, the second aspect, and possible implementations thereof.
[0041] In a seventh aspect, the present application also provides a program product, which includes a program. When the program runs on a processor, it implements the methods corresponding to the first aspect, the second aspect and their possible implementation methods.
[0042] In an eighth aspect, the present application provides a chip comprising a memory and a processor, wherein the memory is used to store instructions, and the processor is used to call and execute the instructions from the memory to implement the methods corresponding to the first aspect, the second aspect and their possible implementation methods. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] FIG1 is a flow chart of a method 100 for constructing a prediction model in an embodiment of the present application;
[0044] FIG2 is a schematic diagram of a framework for information interaction of an initial model in an embodiment of the present application;
[0045] FIG3 is a schematic diagram of initial model training corresponding to Channel Attention in an embodiment of the present application;
[0046] FIG4 is a schematic diagram of initial model training corresponding to Temporal Attention in an embodiment of the present application;
[0047] FIG5 is a flow chart of a method 200 for calling a prediction model in an embodiment of the present application;
[0048] FIG6 is a schematic structural diagram of a communication device 600 according to an embodiment of the present application;
[0049] FIG7 is a schematic structural diagram of a communication device 700 according to an embodiment of the present application;
[0050] FIG8 is a schematic structural diagram of a communication device 800 in an embodiment of the present application. DETAILED DESCRIPTION
[0051] Multivariate time series forecasting has widespread application in many fields, such as finance, economics, meteorology, and transportation. By obtaining future forecasts for multiple variables involved in a known multivariate time series, decision makers can understand the future trends of these variables, serving as a crucial reference for making appropriate decisions. For example, in finance, investors need to predict the future trends of multiple variables, such as stock prices, exchange rates, and interest rates, to make investment decisions. In meteorology, forecasters need to predict the future trends of multiple variables, such as temperature, humidity, and rainfall, to provide accurate weather forecasts. In transportation, planners need to predict the future trends of multiple variables, such as traffic flow and congestion, to formulate reasonable transportation plans. Furthermore, in grid-connected industrial and commercial entities, industrial entities need to predict load power consumption to determine the amount of electricity they can buy and sell from the grid to ensure that their production is met while maximizing their profits. Therefore, accurate multivariate time series forecasting is crucial in practical applications.
[0052] Current forecasts for multivariate time series are not accurate enough.
[0053] For example, the Frequency Enhanced Decomposed Transformer (FEDFormer) has the following main principles: relying on the attention mechanism to capture point-wise relationships, and utilizing the time-to-frequency domain transformation to transform operations such as enhancement and attention to the frequency domain, it can optimize prediction performance to a certain extent and reduce the computational complexity of the model. However, FEDFormer has the following problems when predicting multivariate time series: First, it does not consider suppressing noise in multivariate time series, resulting in inaccurate prediction results; Second, to reduce computational complexity, it converts time domain data with larger vector lengths into frequency domain data with smaller vector lengths before performing operations in the frequency domain. The sampling in the frequency domain is random, resulting in a loss of information content in the multivariate time series and inaccurate prediction results; Third, FEDFormer itself has a certain degree of complexity and requires training with a large amount of data to achieve a certain level of prediction accuracy.
[0054] For another example, the main principle of Dlinear can be understood as: using a multilayer perceptron (MLP) structure for modeling to capture point-to-point relationships. In specific implementation, a multivariate time series is data collected at T moments with n features, where n and T are both integers greater than 1. Dlinear splits the multivariate time series in the feature dimension and divides the multivariate time series into n "unit time series," where n is the number of "elements" in the multivariate time series. Each "unit time series" only indicates the data of one feature at multiple moments. For each "unit time series," a separate model is established. Each model only needs to consider using the MLP structure to capture the mapping from T input moments to K future output moments, where K is an integer greater than or equal to 1 and can be equal to or different from T. Although Dlinear predicts multivariate time series by modeling different features separately, it can prevent the prediction of the "unit time series" corresponding to each feature from introducing noise from other features. However, there are the following problems: Problem 1: Separate modeling in the feature dimension does not introduce any information from other features, resulting in the prediction of each feature lacking the contribution of other features that are associated with it, and the prediction result is likely to be inaccurate; Problem 2: Although it is guaranteed not to introduce noise from other features, the "unit time series" corresponding to each feature will inevitably have its own noise. This method does not consider suppressing the noise of the feature itself in the "unit time series" corresponding to each feature.
[0055] Moreover, whether it is FEDFormer or Dlinear, modeling is only performed in the time dimension, not in the feature dimension.
[0056] Based on this, an embodiment of the present application provides a method for constructing a prediction model for predicting multivariate time series. During the modeling process of the prediction model, information interaction is carried out from different dimensions to enrich the information content in the multivariate time series. Moreover, in the process of information interaction, the noise generated between dimensions and the noise brought by the dimensions themselves are suppressed to obtain a prediction model with higher prediction accuracy.
[0057] In one possible implementation, an embodiment of the present application provides a method for constructing a prediction model, in which, first, a plurality of first training samples are obtained, each of the plurality of first training samples includes a first multivariate time series and a second multivariate time series, and the second multivariate time series is a multivariate time series collected after the first multivariate time series in the first training sample to which it belongs; then, based on the first multivariate time series, the second multivariate time series, an initial model, and an information bottleneck constraint, a prediction model is obtained, the initial model including a model for information interaction in at least one dimension, the information bottleneck constraint is used to suppress noise in the first multivariate time series, and the prediction model is a trained model for predicting a multivariate time series after a known multivariate time series. In this way, during the prediction model training process, sufficient information interaction is carried out on the multivariate time series in the training samples, and various noises are suppressed based on the information bottleneck constraint, so that the constructed prediction model can more accurately predict the multivariate time series.
[0058] It should be noted that the subject implementing the method for constructing the prediction model may be the device for constructing the prediction model provided in the embodiment of the present application, and the device for constructing the prediction model may be carried in an electronic device or a functional module of the electronic device.
[0059] In another possible implementation, an embodiment of the present application further provides a method for calling a prediction model. In this method, after a first multivariate time series to be predicted is generated, a prediction model constructed based on the method provided in the embodiment of the present application can be obtained, and the prediction model is used to predict the first multivariate time series to obtain a second multivariate time series. Because the prediction model can fully exchange information with the first multivariate time series and has the function of suppressing various noises, the prediction result (i.e., the second multivariate time series) obtained for the first multivariate time series using the prediction model is accurate and close to the actual multivariate time series that appears after the first multivariate time series, providing a reliable basis for making reasonable decisions based on the prediction results.
[0060] It should be noted that the subject implementing the prediction model calling method may be the prediction model calling device provided in the embodiments of this application, and the prediction model calling device may be hosted in an electronic device or a functional module of the electronic device. The subject implementing the prediction model calling method may be the same as or different from the subject implementing the prediction model construction method.
[0061] Before introducing the method provided in the embodiments of the present application, the concepts that may be involved in the embodiments of the present application are first explained.
[0062] The attention mechanism generally refers to a mechanism that extracts partial information from all available information for information perception. This mechanism originates from research on human vision. Due to information processing bottlenecks in human cognition, humans selectively focus on a portion of all information while ignoring other visible information. This is why the attention mechanism emerged. For neural networks, the attention mechanism can be understood as the layer-by-layer aggregation and updating of information. For multivariate time series prediction, the attention mechanism updates information at a specific location using information from other locations.
[0063] Multi-view (MV) refers to learning from multiple perspectives, allowing the model to better understand things from multiple aspects, thereby improving model performance. The understanding of multi-view can include the following two aspects: First, multi-view refers to multiple sources. For example, when recognizing a person, faces, fingerprints, etc. can be used as input from different sources; second, multi-view refers to multiple feature subsets. For example, an image can be represented by different features such as color and text. For the prediction of multivariate time series, multi-view can be understood as modeling from different perspectives of the data. For example, the prediction model of a multivariate time series can exchange information in the time dimension and the feature (also called meta-, variable, or channel) dimension.
[0064] Mutual information is a measure of information in information theory. It can be understood as the amount of information contained in a random variable about another random variable, or it can be understood as the uncertainty of a random variable reduced by knowing another random variable.
[0065] The Information Bottleneck (IB) is an information-theoretic approach used to explain and analyze the complexity of data. This theory posits that information in data can be divided into two components: useful information and useless information. Useful information refers to the portion of the data that can be used to predict or explain the data, while useless information provides no predictive or explanatory value. The core idea of the information bottleneck is to extract key information from the data (key information can also be understood as as much useful information as possible) by minimizing the difference between useful and useless information. This difference is a bottleneck that limits data compression and prediction capabilities, hence the term "information bottleneck."
[0066] The following is an introduction to a method for constructing a prediction model provided in an embodiment of the present application.
[0067] Figure 1 is a flow chart of a method 100 for constructing a prediction model according to an embodiment of the present application. The method 100 can be applied to a device for constructing a prediction model, such as the device 600 for constructing a prediction model shown in Figure 6 below.
[0068] As shown in FIG1 , the method 100 may include, for example, the following steps S101 to S102 :
[0069] S101 , obtaining a plurality of first training samples, wherein each of the plurality of first training samples comprises a first multivariate time series and a second multivariate time series, wherein the second multivariate time series is a multivariate time series collected after the first multivariate time series in the first training sample to which it belongs.
[0070] As an example, S101 may include: obtaining historical data of a preset time length (such as 30 days), which historical data may be data collected and saved according to a preset time granularity (such as 15 minutes) within the preset time length; dividing the historical data in the time dimension according to a preset step size (such as 192) to obtain multiple first training samples, each first training sample may include a first multivariate time series and a second multivariate time series, the first multivariate time series may be the first T (such as 96) multivariate time series in the first training sample, and the second multivariate time series is the multivariate time series in the first training sample other than the first multivariate time series.
[0071] Taking the time granularity of historical data as 15 minutes (i.e., data of each feature is collected every 15 minutes) and the preset time length as 30 days as an example, the historical data includes feature data of 96*30=2880 time points in total; assuming that the constructed prediction model uses the multivariate time series of one day (i.e., 96 time points) to predict the multivariate time series of the next day, then the input and output of the prediction model require feature data of 96*2=192 consecutive time points. Therefore, when dividing the feature data of 2880 time points in the historical data, 192 is used as the preset step size to obtain 15 first training samples, each of which includes feature data of 192 consecutive time points; in each first training sample, the feature data of the first 96 time points constitute the first multivariate time series, which is used as the input when training the prediction model using the first training sample, and the feature data of the last 96 time points constitute the second multivariate time series, which is used as the theoretical output corresponding to the input first multivariate time series when training the prediction model using the first training sample, which can be understood as the label of the training data in the usual training sample.
[0072] It can be seen that the first training sample obtained in S101 provides a data basis for implementing the training step in S102.
[0073] S102, obtaining a prediction model based on the first multivariate time series, the second multivariate time series, the initial model and the information bottleneck constraint, wherein the initial model includes a model for information interaction in at least one dimension, the information bottleneck constraint is used to suppress noise in the first multivariate time series, and the prediction model is a trained model for predicting a multivariate time series after a known multivariate time series.
[0074] As an example, S102 may include: S11, training the initial model based on the i-th first training sample to obtain an updated initial model, where i = 1, 2, ..., A, and A is the number of first training samples; S12, determining whether the training stop condition is met, and if so, executing S14; otherwise, executing S13; S13, i = i + 1 and returning to execute S11; S14, ending the training process and determining the updated initial model as the prediction model in S102. In this way, by training the initial model multiple times and continuously updating the initial model, the prediction performance of the initial model is continuously improved. The initial model obtained when the training stop condition is met is recorded as the prediction model, thereby achieving accurate prediction of multivariate time series. The training stop condition may include, but is not limited to: i equals A, or the prediction accuracy of the initial model is greater than a preconfigured accuracy threshold.
[0075] In an embodiment of the present application, in order to enable the constructed prediction model to interact with the input multivariate time series, so as to improve the richness of the data input to the prediction model and better explore the correlation between features, the architecture of the initial model is a model architecture that includes at least one dimension for information interaction. The architecture of the initial model may include, but is not limited to: Multi-View Attention architecture or MLP architecture. The following description will be made by taking the architecture of the initial model as an example of a Multi-View Attention architecture. The at least one dimension may include a time dimension, a frequency dimension, a variable dimension, a feature dimension or a channel dimension. The following description will be made by taking the three cases where at least one dimension is the first dimension, the second dimension or the first dimension + the second dimension, the first dimension is the variable dimension, the feature dimension or the channel dimension, and the second dimension is the time dimension or the frequency dimension as an example.
[0076] For information interaction in the initial model, see Figure 2. In Figure 2, the multivariate time series is a 6*3 sequence, including feature data for 3 features at 6 time points. 6 can be understood as the sequence length, and 3 can be understood as the number of features. When the first multivariate time series passes through the Multi-View Attention architecture, information interaction occurs from different perspectives. Possible approaches, for example, may include: information interaction from the time dimension (Temporal Attention in Figure 2) and / or information interaction from the feature dimension (Channel Attention in Figure 2). Temporal Attention can be understood as information interaction using data from different time points of the same feature. For example, in Figure 2, for a column of features represented by black, the data at each time point can be exchanged with the data from all or part of the other time points of the feature to obtain updated feature data at that time point. The number of time points after information interaction is not limited. Channel Attention can be understood as using different feature data at the same (or different) time points to interact with each other. For example, in Figure 2, a feature data in the first row can be interacted with all or part of the other feature data in the first row, and / or with all or part of the other feature data in each row to obtain the updated feature data of the feature. The number of features after information interaction is not limited.
[0077] The role of information interaction can be understood as: discovering the correlation between each data in the multivariate time series, so that the amount of information that can be extracted is rich enough during information extraction, so that it can be more accurately mapped to downstream tasks (that is, more accurate predictions for the future).
[0078] The Multi-View Attention architecture includes N times of Channel Attention and M times of Temporal Attention. Therefore, the execution order of Channel Attention and Temporal Attention can be flexibly designed based on requirements. For example, Channel Attention and Temporal Attention can be executed alternately, and the alternation can be regular or irregular. When the alternation is regular, the alternation frequency can be 1 or an integer greater than 1. For another example, Channel Attention and Temporal Attention can be executed sequentially, with N times of Channel Attention followed by M times of Temporal Attention, or M times of Temporal Attention followed by N times of Channel Attention. For another example, Channel Attention and Temporal Attention can be executed simultaneously without interfering with each other. After executing N times of Channel Attention and M times of Temporal Attention, the initial models obtained from each are fused. Figure 2 takes the Multi-View Attention architecture as an example, which includes N Channel Attentions and M Temporal Attentions, and in which Channel Attention and Temporal Attention are performed alternately with an alternating frequency of 1. Both M and N are integers greater than or equal to 1, and the values of M and N can be the same or different.
[0079] During the process of integrating interactive information across different dimensions of the first multivariate time series input into the initial model, an information bottleneck constraint is designed as a constraint on the initial model training process to suppress noise. This constraint also corresponds to different constraints for different dimensions. Optionally, this constraint can also be used to enhance the learning of useful information. The following example illustrates this constraint, which can suppress noise and enhance the learning of useful information.
[0080] In one possible implementation, if the at least one dimension includes a first dimension, then S102 may include N updates to the initial model in the first dimension. In the scenario shown in FIG2 , this can be understood as S102 including N executions of Channel Attention. Each update of the initial model in the first dimension is similar. The following uses a single update of the initial model in the first dimension as an example to describe the implementation of S102 in this implementation.
[0081] As an example, the updating process of the initial model in the first dimension in S102 may include: S21, inputting the first multivariate time series into the initial model and outputting a third multivariate time series; S22, obtaining first mutual information based on the first multivariate time series and the first intermediate result, wherein the first intermediate result is a high-dimensional mapping of the first multivariate time series in the first dimension, the first mutual information being used to indicate the impact of noise from dimensions other than the first dimension on the prediction process, and the first mutual information being used to indicate useful information learned from dimensions other than the first dimension; S23, obtaining second mutual information based on the second multivariate time series and the first intermediate result, wherein the second mutual information is used to indicate the connection between representation and prediction; S24, determining a first loss based on the first mutual information and the second mutual information, and updating the initial model based on the first loss. Optionally, the process may further include: S25, updating the initial model based on the second multivariate time series and the third multivariate time series. The "initial model" updated in S25 may be based on the "initial model" obtained after the update in S24.
[0082] It is understood that the initial model may include a first encoding module and a first decoding module that perform information exchange in the first dimension, with the output of the first encoding module serving as the input to the first decoding module. For example, the aforementioned first intermediate result serves as both the input to the first encoding module and the output of the first decoding module. Both the first encoding module and the first decoding module are capable of performing information exchange corresponding to the first dimension. For example, the process of the first encoding module performing high-dimensional mapping on the input first multivariate time series to obtain the first intermediate result involves information exchange in the first dimension.
[0083] It should be noted that updating the initial model based on the second multivariate time series and the third multivariate time series in S25 may, for example, be updating the initial model based on the difference between the second multivariate time series and the third multivariate time series.
[0084] It should be noted that the execution order of S22 and S23 is not limited in the embodiment of the present application. The execution order of S24 and S25 is not limited in the embodiment of the present application.
[0085] For example, assuming the first dimension is the variable dimension, feature dimension, or channel dimension, as shown in Figure 3, the initial model 10 may include a first encoding module 11 and a first decoding module (also called a first prediction module) 12, which interact with information in the first dimension. The first training sample includes a first multivariate time series X and a second multivariate time series Y. The first multivariate time series X is input into the first encoding module 11, which outputs a first intermediate result Zi. This first intermediate result Zi is input into the first decoding module 12, which outputs a third multivariate time series Yi'. In Figure 3, Xi represents the i-th feature in the first multivariate time series X, Xj represents features other than the i-th feature in the first multivariate time series X, and Zi represents the representation of Xi (also called a high-dimensional mapping). In Channel Attention, the role of the information bottleneck constraint 1 corresponding to the first dimension can be understood as: on the one hand, suppressing the impact of noise from other features on the training process; on the other hand, increasing the useful information learned from these other features. For example, suppressing noise and increasing the useful information learned can be achieved through the mutual information between Zi and Xj (i.e., the first mutual information described above). The mutual information between Zi and Yi (i.e., the second mutual information) is used to indicate the connection between representation and prediction. The information bottleneck constraint 1 can be expressed, for example, by the following formula (1): L IB =I(Zi; Yi)-β*I(Zi; Xj)...Formula (1)
[0086] Among them, I(Zi;Yi) can correspond to the second mutual information in S23, I(Zi;Xj) can correspond to the first mutual information in S22, β is a hyperparameter that can be flexibly set based on actual needs, L IB This is the basis for updating the initial model in S24 in this Channel Attention, and can be understood as the first loss determined based on the first mutual information and the second mutual information.
[0087] In this example, after executing Channel Attention N times, the initial model updated by S24 or S25 in the Nth Channel Attention can be used as the prediction model in S102.
[0088] As another example, if the at least one dimension includes a second dimension in addition to the first dimension, then S102 may include N updates to the initial model in the first dimension and M updates to the initial model in the second dimension. Corresponding to the scenario in Figure 2, it can be understood that S102 includes executing N Channel Attentions and M Temporal Attentions. The process of updating the initial model in the first dimension each time can be referred to S21 to S25 above. The process of updating the initial model in the second dimension each time is similar. The following takes the process of updating the initial model in the second dimension once as an example to introduce the implementation of S102 in this implementation method.
[0089] In which, the updating process of the initial model in the second dimension in S102 may include: S31, inputting the first multivariate time series into the initial model and outputting a fourth multivariate time series; S32, obtaining third mutual information based on the first multivariate time series and the second intermediate result, wherein the second intermediate result is a high-dimensional mapping of the first multivariate time series in the second dimension, and the third mutual information is used to indicate the impact of noise from each dimension on the prediction process, wherein the dimensions include the second dimension; S33, obtaining fourth mutual information based on the second multivariate time series and the second intermediate result, wherein the fourth mutual information is used to indicate the connection between representation and prediction; S34, determining a second loss based on the third mutual information and the fourth mutual information, and updating the initial model based on the second loss. Optionally, the process may also include: S35, updating the initial model based on the second multivariate time series and the fourth multivariate time series.
[0090] It is understandable that the initial model may further include a second encoding module and a second decoding module for information exchange in the second dimension, with the output of the second encoding module serving as the input to the second decoding module. For example, the aforementioned second intermediate result is both the output of the second encoding module and the input to the second decoding module. Both the second encoding module and the second decoding module can implement information exchange corresponding to the second dimension. For example, the process of the second encoding module performing high-dimensional mapping on the input first multivariate time series to obtain the second intermediate result involves information exchange in the second dimension.
[0091] It should be noted that, in S35 , the initial model is updated based on the second multivariate time series and the fourth multivariate time series. For example, the initial model may be updated based on the difference between the second multivariate time series and the fourth multivariate time series.
[0092] It should be noted that the execution order of S32 and S33 is not limited in the embodiment of the present application. The execution order of S34 and S35 is not limited in the embodiment of the present application.
[0093] It should be noted that the difference between S33 and S23 is that S23 suppresses the impact of noise from other features on the training process based on the first loss, but S33 can suppress the impact of noise from each feature on the training process based on the second loss. Each feature, including this feature and other features, can suppress noise more comprehensively, so that the prediction model can perform more comprehensive noise suppression on the multivariate time series to be inferred, making it possible for the prediction model to make more accurate predictions.
[0094] For example, assuming the first dimension is the variable dimension, feature dimension, or channel dimension, and the second dimension is the time dimension or frequency dimension, as shown in Figure 4, the initial model 10 may further include a second encoding module 21 and a second decoding module (also called a second prediction module) 22 for information exchange in the second dimension. The first training sample includes a first multivariate time series X and a second multivariate time series Y. The first multivariate time series X is input into the second encoding module 21, which outputs a second intermediate result Z. The second intermediate result Z is input into the second decoding module 22, which outputs a fourth multivariate time series Y'. In Figure 4, Z represents X. In Temporal Attention, the role of the information bottleneck constraint 2 corresponding to the second dimension can be understood as: first, suppressing the impact of noise from other features on the training process; second, suppressing the impact of noise from the feature itself on the training process. For example, suppressing noise from other features and suppressing noise from the feature itself can be achieved by reducing the mutual information between Z and X (i.e., the third mutual information mentioned above). The mutual information between Z and Y (i.e., the fourth mutual information) indicates the connection between representation and prediction. The information bottleneck constraint 2 can be expressed, for example, by the following formula (2): L IB =I(Z;Y)-β*I(Z;X)...Formula (2)
[0095] Among them, I(Z; Y) can correspond to the fourth mutual information in S33, I(Z; X) can correspond to the third mutual information in S32, β is a hyperparameter that can be flexibly set based on actual needs, L IB This is the basis for executing S34 to update the initial model in this Temporal Attention, which can be understood as the second loss determined based on the third mutual information and the fourth mutual information.
[0096] In this example, after executing N Channel Attentions and M Temporal Attentions, the updated initial model obtained from the last Attention can be used as the prediction model in S102. In this example, the update of the initial model in the second dimension and the update of the initial model in the first dimension can be performed according to preset rules. Regardless of whether the initial model is updated in the second dimension or the first dimension, the initial model after the previous update can be used as the initial model in the next update operation.
[0097] In another possible implementation, if the at least one dimension includes a second dimension, then S102 may include L updates to the initial model in the second dimension. In the scenario shown in FIG2 , this can be understood as S102 including L executions of Temporal Attention. Each update of the initial model in the second dimension is similar, and the update of the initial model in the first dimension can be described, for example, in S31 to S35 above. L may be an integer greater than or equal to 1, and L and N may be the same or different.
[0098] In this implementation, after executing Temporal Attention L times, the updated initial model in the Lth Temporal Attention can be used as the prediction model in S102.
[0099] In some implementations, after the prediction model construction device completes S101 to S102, the obtained prediction model can be integrated into a device or server that can be loaded by the user, so that the prediction model calling device can load the prediction model from the device or server when there is a prediction task for a multivariate time series, and use the prediction model to implement the prediction of the multivariate time series.
[0100] As an example, after S102, method 100 may further include: S103, based on the fifth multivariate time series and the prediction model, obtaining a sixth multivariate time series, wherein the sixth multivariate time series is a prediction result of the prediction model for a multivariate time series that occurs after the fifth multivariate time series. The fifth multivariate time series may have the same scale as the first multivariate time series, and the sixth multivariate time series may have the same scale as or different scale from the second multivariate time series.
[0101] Typically, when prediction is performed directly based on the prediction model constructed in S102 , the scale of the sixth multivariate time series is the same as that of the second multivariate time series.
[0102] Taking the prediction task of multivariate time series in the scenario of "load prediction" as an example, when constructing a high-precision prediction model in method 100, historical data can come from at least one data set, including the load data of a certain site for one year in the at least one data set, with a time granularity of 15 minutes, and also including multiple numerical meteorological characteristics of the site for one year that may affect the load. The training process refers to S101 to S102 in method 100 to obtain a load prediction model. When using this load prediction model, it can be based on the load data X at T historical time points (one time point occurs every 15 minutes) 1:T ∈R T×D , predict the load situation X at the next K time points T+1:T+K ∈R K×D , where D is the number of features involved in the multivariate time series. If T is 96 in the first multivariate time series and K is 96 in the second multivariate time series during training, then when using this load forecasting model, the historical multivariate time series of T = 96 time points (96 * 15 minutes = one day) can be used to predict the load situation for the next day. For example, the multivariate time series collected every 15 minutes between 3:00 PM today and 3:00 PM yesterday is used as the input of the load forecasting model. The forecast results output by the load forecasting model can indicate the load situation from 3:00 PM today to 3:00 PM tomorrow.
[0103] When a prediction result of a specific scale is required, and this scale differs from the scale of the second multivariate time series, the prediction model structure can be adjusted to achieve the prediction result of the specific scale. The prediction model can include Multi-View Attention and MLP. The MLP is used to perform simple adjustments to the data output by Multi-View Attention. Adjusting the MLP structure can achieve the effect of adjusting the prediction result to the specific scale.
[0104] The prediction effect of the prediction model constructed in the embodiment of the present application (also referred to as MTS-VIB) was compared with other models capable of predicting multivariate time series. It was verified that in multiple cross-validations conducted on the same data set, the prediction performance of the prediction model constructed in the embodiment of the present application was better than that of other models.
[0105] The prediction model constructed in the embodiments of this application is applicable to various fields involving multivariate time series forecasting, such as influenza epidemics, traffic, weather, electricity, and transformer oil temperature. It has been verified that the prediction model constructed in the embodiments of this application has the best prediction performance among various multivariate time series prediction models based on the open source datasets ILI, Traffic, Weather, Electricity, ETTm2, ETTm1, and ETTh2.
[0106] As can be seen, in method 100, through multi-perspective information interaction and information bottleneck constraints, a prediction model is trained. This prediction model can extract richer useful information and suppress various noises, thereby enabling high-precision prediction of multivariate time series. Furthermore, this prediction model is not limited to the selection of the initial model and has strong generalization capabilities.
[0107] The following introduces a method for calling a prediction model provided in an embodiment of the present application.
[0108] Figure 5 is a flow chart of a method 200 for calling a prediction model provided in an embodiment of the present application. The method 200 can be applied to a device for calling a prediction model, such as the device 700 for calling a prediction model shown in Figure 7 below.
[0109] As shown in FIG5 , the method 200 may include, for example, the following steps S201 to S202 :
[0110] S201, obtaining a prediction model, where the prediction model is a trained model for predicting a multivariate time series. The prediction model is trained under an information bottleneck constraint, which is used to suppress noise in the multivariate time series in the training sample.
[0111] S202 : Obtain a second multivariate time series based on the prediction model and the first multivariate time series, where the second multivariate time series is a prediction result of the prediction model for a multivariate time series that occurs after the first multivariate time series.
[0112] Optionally, the information bottleneck constraint can also be used to enhance the learning of useful information in the first multivariate time series.
[0113] It should be noted that the prediction model can refer to the relevant description of the prediction model in method 100. S201 and S202 can refer to the relevant description of S103. The first multivariate time series can correspond to the fifth multivariate time series in S103, and the second multivariate time series can correspond to the sixth multivariate time series in S103.
[0114] It can be seen that in method 200, since the prediction model can fully interact with the first multivariate time series and has the function of enhancing the learning of useful information in the first multivariate time series and suppressing various noises, the prediction result (i.e., the second multivariate time series) obtained for the first multivariate time series using the prediction model is accurate and close to the real multivariate time series that appears after the first multivariate time series, providing a reliable basis for users to make reasonable decisions based on the prediction results.
[0115] Accordingly, the embodiment of the present application further provides a communication device 600 (also referred to as a prediction model construction device 600), as shown in Figure 6. The communication device 600 may include: an acquisition unit 601 and a training unit 602.
[0116] An acquisition unit 601 is configured to acquire multiple first training samples, each of which includes a first multivariate time series and a second multivariate time series, where the second multivariate time series is a multivariate time series acquired after the first multivariate time series in the corresponding first training sample. Acquisition unit 601 may execute S101 shown in FIG1 .
[0117] A training unit 602 is configured to obtain a prediction model based on the first multivariate time series, the second multivariate time series, an initial model, and an information bottleneck constraint, wherein the initial model includes a model for information interaction in at least one dimension, the information bottleneck constraint is configured to suppress noise in the first multivariate time series, and the prediction model is a trained model for predicting a multivariate time series subsequent to a known multivariate time series. The training unit 602 may execute S102 shown in FIG1 .
[0118] In a possible implementation, the information bottleneck constraint is further used to enhance the learning of useful information in the first multivariate time series.
[0119] In a possible implementation, the at least one dimension includes a first dimension, and the training unit 602 is specifically configured to perform N updates on the initial model in the first dimension, where N is an integer greater than or equal to 1.
[0120] As an example, the training unit 602 performs N updates on the initial model in the first dimension, and the updating process of the initial model in the first dimension includes: inputting the first multivariate time series into the initial model and outputting a third multivariate time series, the initial model including a first encoding module and a first decoding module for information interaction in the first dimension, and the output of the first encoding module serves as the input of the first decoding module; obtaining first mutual information based on the first multivariate time series and the first intermediate result output by the first encoding module, the first intermediate result being a high-dimensional mapping of the first multivariate time series in the first dimension, the first mutual information being used to indicate the impact of noise from dimensions other than the first dimension on the prediction process, and the first mutual information being also used to indicate useful information learned from dimensions other than the first dimension; obtaining second mutual information based on the second multivariate time series and the first intermediate result, the second mutual information being used to indicate the connection between representation and prediction; updating the initial model based on the first mutual information and the second mutual information to obtain an updated initial model; and updating the updated initial model based on the first multivariate time series and the third multivariate time series.
[0121] In another possible implementation, the at least one dimension also includes a second dimension, and the training unit 602 is further used to: perform M updates on the initial model in the second dimension, where M is an integer greater than or equal to 1.
[0122] As an example, the training unit 602 performs M updates on the initial model in the second dimension, and the updating process of the initial model in the second dimension once includes: inputting the first multivariate time series into the initial model and outputting a fourth multivariate time series, the initial model also including a second encoding module and a second decoding module for information interaction in the second dimension, and the output of the second encoding module serves as the input of the second decoding module; obtaining third mutual information based on the first multivariate time series and the second intermediate result output by the second encoding module, the second intermediate result being a high-dimensional mapping of the first multivariate time series in the second dimension, and the third mutual information being used to indicate the impact of noise from each dimension on the prediction process, the dimensions including the second dimension; obtaining fourth mutual information based on the second multivariate time series and the second intermediate result, the fourth mutual information being used to indicate the connection between representation and prediction; updating the initial model based on the third mutual information and the fourth mutual information to obtain an updated initial model; and updating the updated initial model based on the first multivariate time series and the fourth multivariate time series.
[0123] As an example, the updating of the initial model in the second dimension and the updating of the initial model in the first dimension are performed alternately according to preset rules, and the initial model after the previous update is used as the initial model in the next update operation.
[0124] As an example, the first dimension is a feature dimension or a channel dimension, and the second dimension is a time dimension or a frequency dimension.
[0125] In another possible implementation, the at least one dimension includes a second dimension, and the training module is specifically used to: perform L updates on the initial model in the second dimension, where L is an integer greater than or equal to 1.
[0126] In one possible implementation, the apparatus 600 may further include a prediction unit configured to obtain a sixth multivariate time series based on the fifth multivariate time series and the prediction model, where the sixth multivariate time series is a prediction result of the prediction model for a multivariate time series that occurs after the fifth multivariate time series.
[0127] As an example, the fifth multivariate time series has the same scale as the first multivariate time series, and the sixth multivariate time series has the same scale as or a different scale than the second multivariate time series.
[0128] It should be noted that various specific implementation modes of the communication device 600 can refer to the relevant introduction of the method 100 corresponding to Figure 1, and will not be repeated in this embodiment.
[0129] Accordingly, the embodiment of the present application further provides a communication device 700 (also referred to as a prediction model calling device 700), as shown in Figure 7. The communication device 700 may include: an acquisition unit 701 and a prediction unit 702.
[0130] Acquisition unit 701 is configured to acquire a prediction model. The prediction model is a trained model for predicting a multivariate time series. The prediction model is trained under an information bottleneck constraint, which is used to suppress noise in the multivariate time series in the training samples. Acquisition unit 701 may execute S201 shown in FIG. 5 .
[0131] Prediction unit 702 is configured to obtain a second multivariate time series based on the prediction model and the first multivariate time series, where the second multivariate time series is a prediction result of the prediction model for a multivariate time series that occurs after the first multivariate time series. Prediction unit 702 may execute S202 shown in FIG5 .
[0132] Optionally, the information bottleneck constraint can also be used to enhance the learning of useful information in the first multivariate time series.
[0133] It should be noted that various specific implementation modes of the communication device 700 can be found in the relevant introduction of the method 200 corresponding to FIG5 , and will not be described in detail in this embodiment.
[0134] Referring to Figure 8 , an embodiment of the present application provides a communication device 800 . The communication device 800 may be the execution entity of any of the aforementioned embodiments. The communication device 800 may implement the functions of the aforementioned embodiments. The communication device 800 includes at least one processor 801 , a bus system 802 , a memory 803 , and at least one communication interface 804 .
[0135] The communication device 800 is a hardware device that can be used to implement the functional modules in the communication device 600 shown in Figure 6. For example, those skilled in the art can imagine that the acquisition unit 601 and the training unit 602 in the communication device 600 shown in Figure 6 can be implemented by the at least one processor 801 calling the code in the memory 803. For another example, those skilled in the art can imagine that the acquisition unit 701 and the prediction unit 702 in the communication device 700 shown in Figure 7 can be implemented by the at least one processor 801 calling the code in the memory 803.
[0136] Optionally, the communication device 800 may be a network device or a control entity implementing an embodiment of the present application.
[0137] Optionally, the processor 801 may be a general-purpose central processing unit (CPU), a network processor (NP), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits for controlling the execution of the program of the present application.
[0138] The bus system 802 may include a path for transmitting information between the components.
[0139] The communication interface 804 is used to communicate with other devices or communication networks.
[0140] The above-mentioned memory 803 can be a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM) or other types of dynamic storage devices that can store information and instructions, or an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compressed optical disc, laser disc, optical disc, digital versatile disc, Blu-ray disc, etc.), a magnetic disk storage medium or other magnetic storage device, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited to this. The memory can exist independently and be connected to the processor through a bus. The memory can also be integrated with the processor.
[0141] The memory 803 is used to store application code for executing the solution of the present application, and the execution is controlled by the processor 801. The processor 801 is used to execute the application code stored in the memory 803, thereby realizing the functions of the method of the present application.
[0142] In a specific implementation, as an embodiment, the processor 801 may include one or more CPUs, such as CPU0 and CPU1 in FIG8 .
[0143] In a specific implementation, as an embodiment, the communication device 800 may include multiple processors, such as processor 801 and processor 807 in Figure 8. Each of these processors may be a single-core (single-CPU) processor or a multi-core (multi-CPU) processor. The processor here may refer to one or more devices, circuits, and / or processing cores for processing data (e.g., computer program instructions).
[0144] It should be understood that the communication devices or communication equipment in the various product forms mentioned above respectively have any functions implemented by the execution subject in the above method embodiments, which will not be described in detail here.
[0145] The present application also provides a chip including a processor and an interface circuit, wherein the interface circuit is configured to receive instructions and transmit them to the processor; the processor, which may be, for example, a specific implementation of the message processing device in the present application embodiment, may be configured to execute the above-described method 100 or method 200. The processor is coupled to a memory configured to store programs or instructions. When the programs or instructions are executed by the processor, the chip system implements the method in any of the above-described method embodiments.
[0146] Optionally, there may be one or more processors in the chip system. The processor may be implemented in hardware or software. When implemented in hardware, the processor may be a logic circuit, an integrated circuit, etc. When implemented in software, the processor may be a general-purpose processor implemented by reading software code stored in a memory.
[0147] Optionally, the memory in the chip system may be one or more memories. The memory may be integrated with the processor or may be provided separately from the processor, which is not limited in this application. For example, the memory may be a non-transient processor, such as a read-only memory (ROM), which may be integrated with the processor on the same chip or provided on different chips. This application does not specifically limit the type of memory or the configuration of the memory and the processor.
[0148] Exemplarily, the chip system can be a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a system on chip (SoC), a central processor unit (CPU), a network processor (NP), a digital signal processor (DSP), a microcontroller unit (MCU), a programmable logic device (PLD) or other integrated chips.
[0149] In addition, an embodiment of the present application further provides a storage medium, in which program code or instructions are stored. When the storage medium is run on a processor, the processor executes a method in any one of the implementation modes of the above embodiments.
[0150] In addition, an embodiment of the present application also provides a program product, which, when executed on a processor, enables the processor to execute any one of the aforementioned methods 100 or 200.
[0151] It should be understood that "determining B based on A" mentioned in the embodiments of the present application does not mean determining B only based on A, but B can also be determined based on A and / or other information.
[0152] It should be understood that the network architecture and business scenarios described in the embodiments of the present application are intended to more clearly illustrate the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided in the embodiments of the present application. Ordinary technicians in this field can know that with the evolution of network architecture and the emergence of new business scenarios, the technical solutions provided in the embodiments of the present application are also applicable to similar technical problems.
[0153] In this application, ordinal numbers such as "1", "2", "3", "first", "second" and "third" are used to distinguish multiple objects and are not used to limit the order of multiple objects.
[0154] “A and / or B” mentioned in this application should be understood to include the following situations: only A, only B, or both A and B.
[0155] Through the description of the above embodiments, it can be known that those skilled in the art can clearly understand that all or part of the steps in the above embodiment methods can be implemented by means of software plus a general hardware platform. Based on this understanding, the technical solution of the present application can be embodied in the form of a software product, which can be stored in a storage medium, such as a read-only memory (ROM) / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network communication device such as a router) to execute the methods described in each embodiment or certain parts of the embodiments of the present application.
[0156] Each embodiment in this specification is described in a progressive manner. The same or similar parts between the embodiments can be referred to each other. Each embodiment focuses on the differences from other embodiments. In particular, for system embodiments and device embodiments, since they are basically similar to method embodiments, the description is relatively simple. For relevant parts, refer to the partial description of the method embodiment. The device and system embodiments described above are merely schematic. The modules described as separate components may or may not be physically separated, and the components displayed as modules may or may not be physical modules, that is, they may be located in one place or distributed on multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without making any creative effort.
[0157] The above description is only a preferred embodiment of the present application and is not intended to limit the scope of protection of the present application. It should be noted that those skilled in the art may make several improvements and modifications without departing from the scope of protection of the present application, and such improvements and modifications should also be considered as within the scope of protection of the present application.
Claims
1. A method for constructing a prediction model, characterized in that: include: Acquire a plurality of first training samples, wherein each of the plurality of first training samples comprises a first multivariate time series and a second multivariate time series, wherein the second multivariate time series is a multivariate time series collected after the first multivariate time series in the first training sample to which it belongs; A prediction model is obtained based on the first multivariate time series, the second multivariate time series, an initial model and an information bottleneck constraint, wherein the initial model includes a model for information interaction in at least one dimension, the information bottleneck constraint is used to suppress noise in the first multivariate time series, and the prediction model is a trained model for predicting a multivariate time series after a known multivariate time series.
2. The method according to claim 1, characterized in that The information bottleneck constraint is also used to enhance the learning of useful information in the first multivariate time series.
3. The method according to claim 1 or 2, characterized in that: The at least one dimension includes a first dimension, and obtaining a prediction model based on the first multivariate time series, the second multivariate time series, an initial model, and an information bottleneck constraint condition includes: The initial model is updated N times in the first dimension, where N is an integer greater than or equal to 1.
4. The method according to claim 3, characterized in that In the N times of updating the initial model in the first dimension, the updating process of the initial model in the first dimension once includes: Inputting the first multivariate time series into the initial model and outputting a third multivariate time series, wherein the initial model includes a first encoding module and a first decoding module for information interaction in the first dimension, and the output of the first encoding module serves as the input of the first decoding module; Obtaining first mutual information according to the first multivariate time series and the first intermediate result, where the first intermediate result is a high-dimensional mapping of the first multivariate time series in the first dimension, the first mutual information is used to indicate the influence of noise from dimensions other than the first dimension on the prediction process, and the first mutual information is also used to indicate useful information learned from dimensions other than the first dimension; Obtaining second mutual information according to the second multivariate time series and the first intermediate result, wherein the second mutual information is used to indicate a connection between the representation and the prediction; Based on the first mutual information and the second mutual information, updating the initial model to obtain an updated initial model; The updated initial model is updated based on the second multivariate time series and the third multivariate time series.
5. The method according to claim 3 or 4, characterized in that: The at least one dimension further includes a second dimension, and the obtaining of a prediction model based on the first multivariate time series, the second multivariate time series, the initial model and the information bottleneck constraint condition further includes: The initial model is updated in the second dimension M times, where M is an integer greater than or equal to 1.
6. The method according to claim 5, characterized in that In the M times of updating the initial model in the second dimension, the updating process of the initial model in the second dimension once includes: Input the first multivariate time series into the initial model and output a fourth multivariate time series; Obtaining third mutual information according to the first multivariate time series and the second intermediate result, wherein the second intermediate result is a high-dimensional mapping of the first multivariate time series in the second dimension, and the third mutual information is used to indicate the influence of noise from each dimension on the prediction process, wherein the dimensions include the second dimension; Obtaining fourth mutual information according to the second multivariate time series and the second intermediate result, wherein the fourth mutual information is used to indicate a connection between the representation and the prediction; Based on the third mutual information and the fourth mutual information, updating the initial model to obtain an updated initial model; The updated initial model is updated based on the second multivariate time series and the fourth multivariate time series.
7. The method according to claim 5 or 6, characterized in that: The updating of the initial model in the second dimension and the updating of the initial model in the first dimension are performed according to a preset rule, and the initial model after the previous update is used as the initial model in the next update operation.
8. The method according to any one of claims 5 to 7, characterized in that: The first dimension is a variable dimension, a feature dimension or a channel dimension, and the second dimension is a time dimension or a frequency dimension.
9. The method according to claim 1 or 2, characterized in that: The at least one dimension includes a second dimension, and obtaining a prediction model based on the first multivariate time series, the second multivariate time series, an initial model, and an information bottleneck constraint condition includes: The initial model is updated L times in the second dimension, where L is an integer greater than or equal to 1.
10. The method according to any one of claims 1 to 9, characterized in that: The method further comprises: Based on the fifth multivariate time series and the prediction model, a sixth multivariate time series is obtained, and the sixth multivariate time series is a prediction result of the prediction model for the multivariate time series that appears after the fifth multivariate time series.
11. The method according to claim 10, characterized in that The fifth multivariate time series has the same scale as the first multivariate time series, and the sixth multivariate time series has the same scale as or a different scale than the second multivariate time series.
12. A method for calling a prediction model, characterized in that: include: Obtaining a prediction model, where the prediction model is a trained model for predicting a multivariate time series, and the prediction model is trained under an information bottleneck constraint, where the information bottleneck constraint is used to suppress noise in the multivariate time series in the training sample; Based on the prediction model and the first multivariate time series, a second multivariate time series is obtained, where the second multivariate time series is a prediction result of the prediction model for a multivariate time series that occurs after the first multivariate time series.
13. A device for constructing a prediction model, characterized in that: include: An acquisition unit, configured to acquire a plurality of first training samples, wherein each of the plurality of first training samples comprises a first multivariate time series and a second multivariate time series, wherein the second multivariate time series is a multivariate time series collected after the first multivariate time series in the first training sample to which it belongs; A training unit is used to obtain a prediction model based on the first multivariate time series, the second multivariate time series, an initial model and an information bottleneck constraint, wherein the initial model includes a model for information interaction in at least one dimension, the information bottleneck constraint is used to suppress noise in the first multivariate time series, and the prediction model is a trained model for predicting a multivariate time series after a known multivariate time series.
14. The device according to claim 13, characterized in that The information bottleneck constraint is also used to enhance the learning of useful information in the first multivariate time series.
15. The device according to claim 13 or 14, characterized in that The at least one dimension includes a first dimension, and the training unit is specifically used to: perform N updates on the initial model in the first dimension, where N is an integer greater than or equal to 1.
16. The device according to claim 15, characterized in that The training unit performs N updates on the initial model in the first dimension, and a process of updating the initial model in the first dimension once, comprising: Input the first multivariate time series into the initial model and output a third multivariate time series; Obtaining first mutual information according to the first multivariate time series and the first intermediate result, where the first intermediate result is a high-dimensional mapping of the first multivariate time series in the first dimension, the first mutual information is used to indicate the influence of noise from dimensions other than the first dimension on the prediction process, and the first mutual information is also used to indicate useful information learned from dimensions other than the first dimension; Obtaining second mutual information according to the second multivariate time series and the first intermediate result, wherein the second mutual information is used to indicate a connection between the representation and the prediction; Based on the first mutual information and the second mutual information, updating the initial model to obtain an updated initial model; The updated initial model is updated based on the first multivariate time series and the third multivariate time series.
17. The device according to claim 15 or 16, characterized in that The at least one dimension also includes a second dimension, and the training unit is further used to: perform M updates on the initial model in the second dimension, where M is an integer greater than or equal to 1.
18. The device according to claim 17, characterized in that The training unit performs M updates of the initial model in the second dimension, and a process of updating the initial model in the second dimension once, comprising: Input the first multivariate time series into the initial model and output a fourth multivariate time series; Obtaining third mutual information according to the first multivariate time series and the second intermediate result, wherein the second intermediate result is a high-dimensional mapping of the first multivariate time series in the second dimension, and the third mutual information is used to indicate the influence of noise from each dimension on the prediction process, wherein the dimensions include the second dimension; Obtaining fourth mutual information according to the second multivariate time series and the second intermediate result, wherein the fourth mutual information is used to indicate a connection between the representation and the prediction; Based on the third mutual information and the fourth mutual information, updating the initial model to obtain an updated initial model; The updated initial model is updated based on the first multivariate time series and the fourth multivariate time series.
19. The device according to claim 17 or 18, characterized in that The updating of the initial model in the second dimension and the updating of the initial model in the first dimension are performed alternately according to a preset rule, and the initial model after the previous updating is used as the initial model in the next updating operation.
20. The device according to any one of claims 17 to 19, characterized in that The first dimension is a feature dimension or a channel dimension, and the second dimension is a time dimension or a frequency dimension.
21. The device according to claim 13 or 14, characterized in that The at least one dimension includes a second dimension, and the training unit is specifically used to: perform L updates on the initial model in the second dimension, where L is an integer greater than or equal to 1.
22. The device according to any one of claims 13 to 21, characterized in that: The device also includes: The prediction unit is used to obtain a sixth multivariate time series based on the fifth multivariate time series and the prediction model, wherein the sixth multivariate time series is a prediction result of the prediction model for the multivariate time series that appears after the fifth multivariate time series.
23. The device according to claim 22, characterized in that The fifth multivariate time series has the same scale as the first multivariate time series, and the sixth multivariate time series has the same scale as or a different scale than the second multivariate time series.
24. A prediction model calling device, characterized in that: include: An acquisition unit, used for acquiring a prediction model, wherein the prediction model is a trained model for predicting a multivariate time series, wherein the prediction model is trained under an information bottleneck constraint, and the information bottleneck constraint is used for suppressing noise of the multivariate time series in the training sample; The prediction unit is used to obtain a second multivariate time series based on the prediction model and the first multivariate time series, where the second multivariate time series is the prediction result of the prediction model for the multivariate time series that appears after the first multivariate time series.
25. A communication device, characterized in that: The communication device includes a memory and a processor; The memory is used to store instructions; The processor is used to execute the instructions in the memory and perform the method according to any one of claims 1 to 12.
26. A storage medium, characterized in that The storage medium includes instructions, and when the instructions are executed on a processor, the processor is caused to execute the method according to any one of claims 1 to 12.
Citation Information
Patent Citations
Method and device for constructing and calling prediction model
CN119989605A
Electric equipment fault forecasting method based on multi-dimension time sequence
CN103996077A
Brain activity multi-dimensional time series signal decoding method and device based on self-learning
CN106264460A
Web Service Qos prediction method based on multivariate time series
CN106357437A
Adaptive time series data prediction method
CN111651444A