Time sequence prediction method and device for hierarchical service

By obtaining and encoding multi-level time series in hierarchical business, building joint probability distribution and sampling, the problem of insufficient utilization of hierarchical correlation in hierarchical business time series prediction is solved, and more accurate time series prediction is achieved.

CN120146255APending Publication Date: 2025-06-13ALIPAY (HANGZHOU) INFORMATION TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510162438.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2022-06-02
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

It is difficult for the prior art to make full use of the timing information of the hierarchical structure for time series prediction, especially in hierarchical business scenarios. How to effectively consider the correlation between each level is a challenge.

Method used

By obtaining the historical time series of each business entity in a multi-level hierarchy, encoding is performed to obtain the encoded vector, constructing a joint probability distribution, and sampling based on this distribution to generate a prediction sequence that satisfies the consistency constraint.

Benefits of technology

This method can make full use of time series information at each level, consider the correlation between levels, thereby improving the accuracy of time series prediction and information fusion effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120146255A_ABST
    Figure CN120146255A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a time sequence prediction method and device for hierarchical services, which are used for predicting a time sequence formed by service volumes of service subjects of a single service on multiple levels, and a single level corresponds to at least one service subject. According to one implementation mode, after historical time sequences in one-to-one correspondence with business subjects in multiple levels are obtained, the historical time sequences can be coded to obtain corresponding coding vectors, then joint probability distribution met by the coding vectors is constructed, and the joint probability distribution is obtained. And further, sampling is carried out on each business main body according to joint probability distribution to obtain a sampling sequence, and each prediction sequence for each business main body is determined. The method can improve the accuracy of the prediction result.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of the invention patent application "Time Series Prediction Method and Device for Hierarchical Services" with the application number 202210621721.9 and filed on June 2, 2022. Technical Field

[0002] One or more embodiments of this specification relate to the field of computer technology, and in particular to a time series prediction method and device for hierarchical services. Background Art

[0003] A time series (or dynamic series) refers to a series formed by arranging the values of the same statistical indicator in the order of their occurrence time. The main purpose of time series analysis is to predict the future based on existing historical data. Time series prediction can be applied to various scenarios, such as time series prediction of the passenger flow in a commodity supermarket, time series prediction of funds in financial services, prediction of the computing resource traffic required in cloud computing, logistics demand, power consumption prediction in smart grids, and so on. The prediction results can, for example, serve commercial decisions. With the development of artificial intelligence, machine learning models can be used for time series analysis.

[0004] A hierarchical time series is a special time series that can describe related services from multiple levels, making the time series have a hierarchical structure. For example, the time series of the sales volume of a large supermarket chain can have a hierarchical structure according to its geographical location hierarchy, such as being divided into the following multiple levels: the national total sales volume time series; the sales volume time series of each province and city; the sales volume time series of each store; and so on. How to make full use of the time series information of the hierarchical structure is an important issue in time series prediction. Summary of the Invention

[0005] One or more embodiments of this specification describe a time series prediction method and device for hierarchical services to solve one or more problems mentioned in the background art.

[0006] According to a first aspect, there is provided a time series prediction method for hierarchical services, which is used to predict the time series of the business volume composition of business entities at multiple hierarchical levels on a predetermined business indicator, where a single hierarchical level corresponds to at least one business entity; the method includes: obtaining respective historical time series corresponding one by one to each business entity in the multiple hierarchical levels, where a single historical time series includes respective business volumes corresponding to a single business entity in a plurality of predetermined time periods arranged in sequence; encoding each historical time series to obtain respective encoded vectors; determining the joint probability distribution satisfied by each encoded vector; based on respective sampling sequences obtained by sampling each business entity according to the joint probability distribution, determining respective prediction sequences for each business entity, where a single prediction sequence describes the respective business volumes of a single business entity in subsequent multiple predetermined time periods, and each prediction sequence satisfies a consistency constraint, and the consistency constraint includes: for any subsequent time period, the business volume of a single business entity at a higher hierarchical level is consistent with the sum of the business volumes of the respective business entities corresponding to it at a lower hierarchical level.

[0007] In one embodiment, the consistency constraint is described by a pre-determined optimization matrix; the determining respective prediction sequences for each business entity based on respective sampling sequences obtained by sampling each business entity according to the joint probability distribution includes: using the sampling matrix obtained by arranging the respective sampling sequences in sequence as a preliminary prediction result; adjusting the preliminary prediction result using the optimization matrix to make it satisfy the consistency constraint, so as to obtain respective prediction sequences.

[0008] In a further embodiment, the optimization matrix is determined based on the hierarchical structure of each business entity in the multiple hierarchical levels, and is formed by splicing the identity matrix for describing the business entities at the lowest hierarchical level and the negative matrix of the sum matrix for describing the business entities at other hierarchical levels, or splicing the negative identity matrix and the sum matrix.

[0009] In a still further embodiment, the business entities at the lowest hierarchical level correspond one by one to the rows / columns of the corresponding identity matrix, and a single business entity at each other hierarchical level corresponds to a single row / column in the sum matrix, and this single row / column is formed by superimposing the respective rows / columns of the respective business entities corresponding to this single business entity at the lowest hierarchical level in the corresponding identity matrix.

[0010] In another embodiment, the consistency constraint is defined by the product of the optimization matrix and the sampling matrix being 0. Adjusting the preliminary prediction result by using the optimization matrix to satisfy the consistency constraint, so as to obtain each prediction sequence includes: under the consistency constraint, solving a time series matrix with the smallest difference from the preliminary prediction result as the predicted prediction matrix; determining each prediction sequence according to each row / column in the prediction matrix.

[0011] In one embodiment, the difference between the time series matrix and the preliminary prediction result is measured by the norm of the difference between the time series matrix and the sampling matrix.

[0012] In one embodiment, the encoding of each historical time series is implemented by one of the following neural networks: a single recurrent neural network; an unfolded recurrent neural network composed of multiple recurrent neural networks corresponding to each business entity one by one.

[0013] In one embodiment, the joint probability distribution is a multivariate Gaussian distribution, and the multivariate Gaussian distribution corresponds to a first mean matrix and a first covariance matrix. A single sampling sequence is determined by the following method: sampling a first vector in the standardized multivariate Gaussian distribution, where the mean matrix of the standardized multivariate Gaussian distribution is a 0 matrix and the covariance matrix is an identity matrix; determining the single sampling sequence based on the superposition result of adding the first mean matrix to the product of the square root of the first covariance matrix and the first vector.

[0014] According to a second aspect, there is provided a time series prediction device for hierarchical services, configured to predict a time series of the business volume composition of business entities at multiple hierarchical levels on a predetermined business indicator, where a single hierarchical level corresponds to at least one business entity; the device includes:

[0015] An acquisition unit, configured to acquire each historical time series corresponding to each business entity in the multiple hierarchical levels, and a single historical time series corresponds to each business volume respectively corresponding to a single business entity in a plurality of predetermined time periods arranged in sequence;

[0016] An encoding unit, configured to encode each historical time series to obtain corresponding encoded vectors;

[0017] A mapping unit, configured to determine the joint probability distribution satisfied by each encoded vector;

[0018] A prediction unit configured to determine respective prediction sequences for respective business entities based on respective sampling sequences obtained by the sampling module sampling respective business entities according to the joint probability distribution, wherein a single prediction sequence describes respective business volumes of a single business entity in multiple subsequent predetermined time periods, and the respective prediction sequences satisfy a consistency constraint, and the consistency constraint includes: for any subsequent time period, the business volume of a single business entity at a higher level is consistent with the sum of the business volumes of at least one business entity corresponding thereto at a lower level.

[0019] According to a third aspect, there is provided a computer-readable storage medium having a computer program stored thereon, which when executed on a computer causes the computer to execute the method of the first aspect.

[0020] According to a fourth aspect, there is provided a computing device including a memory and a processor, characterized in that the memory stores executable code, and when the processor executes the executable code, the method of the first aspect is implemented.

[0021] Through the method and device provided by the embodiments of the present specification, in the time series prediction process for hierarchical services, the time series of each level are mapped into high-order vectors and a joint probability distribution is constructed, so as to consider the correlation of the time series of each level, and resampling is performed according to this joint probability distribution to obtain corresponding predicted time series. This prediction method makes full use of the correlation relationship of service data at different levels to achieve information fusion, thereby improving the accuracy of the prediction result. Description of the Drawings

[0022] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can also obtain other drawings based on these drawings without creative efforts.

[0023] Figure 1 Showing a specific architecture diagram of a hierarchical time series problem;

[0024] Figure 2 Showing a flowchart of an improved time series prediction method for hierarchical services;

[0025] Figure 3 Showing a schematic diagram of a historical time series processing architecture according to a specific example;

[0026] Figure 4 Showing a schematic block diagram of a time series prediction device for hierarchical services according to an embodiment. Detailed Embodiments

[0027] The technical solution provided in this specification will be described below with reference to the accompanying drawings.

[0028] Figure 1 A specific architecture diagram showing hierarchical timing issues is as follows Figure 1 As shown, in this specific application scenario, three business levels are involved. The lowest business level is the level corresponding to specific business entities (such as business entities D, E, F, G, etc.), the higher level is the level corresponding to business entities with certain comprehensive meanings (such as business entities B, C, etc.), and the even higher level can be the level corresponding to business entities with stronger comprehensiveness (such as business entity A, etc.). Here, a business entity can be an entity that can independently distinguish and count corresponding businesses, such as a store, a category of goods, goods or stores in a region, and so on. Figure 1 shows a hierarchical business problem with a three-level structure. In practice, the hierarchical business structure can have more (such as 8 levels) or fewer levels (such as 2 levels), which is not limited here.

[0029] A single business entity can correspond to a certain business volume for the corresponding business indicators within a single time period. The business volumes in multiple time periods constitute a time series, Figure 1 abbreviated as "time series" in this specification. For example, business entity A (such as the stores of a large supermarket chain nationwide) corresponds to time series y A , t (such as the sales volume business indicator describing the sales business), business entity B (such as the stores of a large supermarket chain in North China) corresponds to time series y B , t , business entity D (such as the stores of a large supermarket chain in South China) corresponds to time series y D , t , and so on. Among them, each time period of the sampled time series can be continuous or discontinuous, which is not limited in this specification. For example, in one embodiment, a single time period is one week, and the data for multiple consecutive weeks constitutes a time series. In another embodiment, a single time period is one day on Monday, and the data for multiple Mondays constitutes a time series, but the time between two adjacent Mondays is discontinuous.

[0030] Among them, according to different specific business scenarios, the business entities at each level are different, and the corresponding business indicators are also different. For example, in the shopping platform scenario, the lowest level can correspond to each merchant, and the corresponding business indicator is, for example, sales volume, and the corresponding time period can be one week, one day, one month, etc. A higher level can correspond to a business entity of a small category, such as merchants in categories like kitchen appliances, newborn clothes, home improvement materials, etc., and an even higher level can correspond to business entities of larger categories, such as merchants in categories like electrical appliances, clothing, building materials, etc. Similarly, in the investment and financial management scenario, the highest level can correspond to various investment channels, such as insurance, funds, fixed deposits, stocks, etc., a higher level can correspond to each investment object, such as insurance companies, fund units, banks, stock main categories, etc., and the lowest level corresponds to various specific investment methods, such as insurance types, fund types, bank fixed deposit product types, stock main bodies, etc. The business indicators described by the time series are investment amount, investment ratio, number of investment users, etc. In other business scenarios, the business entities at each level can also be in other forms, which will not be elaborated one by one here.

[0031] According to the meaning of hierarchical business, the time series of business indicators corresponding to business entities at adjacent levels have the following consistency relationship: within a single time period, the business volume corresponding to a single business entity at a higher level is the sum of the business volumes of the business entities it corresponds to at a lower level. Further, when the time periods corresponding to the time series are the same, the time series corresponding to a single business entity at a higher level can be the sum of the time series of the business entities it corresponds to at a lower level. Referring Figure 1 as shown, there is: y A,t = y B,t + y C,t ; y B,t = y D,t + y E,t ; y C,t = y F,t + y G,t .

[0032] In the hierarchical time series prediction process as Figure 1 shown, the candidate time series can be predicted based on the historical time series of each business entity. For example, based on the historical time series corresponding to the sales volume of each business entity in the past 3 months, predict the time series corresponding to the sales volume of each business entity in the next 30 days, etc.

[0033] To this end, the time series of each business entity can be predicted separately and then harmonized uniformly. For example, the time series corresponding to the sales volume in the next 30 days can be predicted separately and independently according to the historical time series corresponding to the sales volume of business entities A, B, C, D, E, F, and G in the past 3 months. After that, the prediction results are harmonized uniformly so that the final prediction result meets the consistency constraint. For example, the time series predicted for A is the sum of the time series predicted for B and C, the time series predicted for B is the sum of the time series predicted for D and E, and the time series predicted for C is the sum of the time series predicted for F and G.

[0034] In fact, there may be a correlation between the time series of each business entity. For example, when the sales volume of the mother and baby product category continues to rise, the sales volume of a single mother and baby store with a relatively low historical sales volume will also continue to rise in the future. Another example is that when the investment ratios of insurance, funds, and stocks continue to rise, the fixed deposit investment ratio decreases. Therefore, this specification provides a time series prediction scheme for hierarchical services, which mines the correlation between the time series of each business entity by establishing a joint probability distribution between the historical time series of each business entity, so as to more accurately predict the time series of each business entity in the future.

[0035] The following refers to Figure 2 a specific example shown to describe the technical concept of this specification.

[0036] As Figure 2 shown, a time series prediction process for hierarchical services in an embodiment is shown. The execution subject of this process can be a computer, device, or server with certain computing capabilities. This process can be assisted by a machine learning model. The corresponding machine learning model can be called a prediction model according to its function for convenience of description. For the sake of description, in the Figure 2 embodiment shown, as an example, it can be assumed that the business entity hierarchy includes at least a first level (such as the level where B and C are located in Figure 1 ) and a second level (such as the level where D, E, F, and G are located in Figure 1 ). If it is assumed that any business entity at the first level corresponds to N second business entities at the second level, then the current time series prediction process for hierarchical services can predict the time series of each business entity, including these N + 1 business entities, on a predetermined business indicator.

[0037] As Figure 2As shown in the figure, the time series prediction process for hierarchical services may include the following steps: Step 201, obtain respective historical time series corresponding to each business entity in multiple hierarchical levels. A single historical time series corresponds to respective business volumes of a single business entity in multiple sequentially arranged predetermined time periods. Step 202, encode each historical time series to obtain respective encoded vectors. Step 203, determine the joint probability distribution satisfied by each encoded vector. Step 204, based on the sampling sequences obtained by sampling each business entity according to the joint probability distribution respectively, determine respective prediction sequences for each business entity. Among them, a single prediction sequence describes the respective business volumes of a single business entity in subsequent multiple predetermined time periods. Each prediction sequence satisfies a consistency constraint. The consistency constraint includes: for any subsequent time period, the business volume of a single business entity at a higher hierarchical level is consistent with the sum of the business volumes of the respective business entities corresponding to it at a lower hierarchical level.

[0038] First, in step 201, obtain respective historical time series corresponding to each business entity in the multiple hierarchical levels of hierarchical services. It can be understood that hierarchical services can correspond to multiple hierarchical levels, such as Figure 1 the three hierarchical levels in. Each hierarchical level has at least one business entity. Each business entity at a lower hierarchical level corresponds to a higher-level business entity in a higher hierarchical level, for example, corresponding to a more general category of business entity in a higher hierarchical level. On the other hand, a single business entity at a higher hierarchical level corresponds to at least one lower-level business entity in a lower hierarchical level.

[0039] The respective business volumes corresponding to a single business entity in multiple sequentially arranged historical predetermined time periods can form a single historical time series. A single predetermined time period can be set according to business requirements, such as one day, one week, one month, etc. Multiple time periods can be continuous or discontinuous in time. Taking a single predetermined time period as one day as an example, assuming that multiple time periods are continuous, such as sampling the sales volume for consecutive days, then for the case where a single merchant is used as a business entity, the sales volume sampled for consecutive days (such as 60 days) can be used as its historical time series. Assuming that multiple time periods are discontinuous, such as only sampling the sales volume on weekends, then for the case where a single merchant is used as a business entity, the sales volume sampled on multiple weekends (such as 52 weekends) can be used as its historical time series. The historical time series can be represented by a multi-dimensional vector, an array, or a set, etc., and is not limited here. Taking the hierarchical service including the aforementioned first level and second level as an example, for the first business entity in the first level and the N business entities corresponding to it in the second level, N + 1 historical time series corresponding one by one can be obtained. Each historical time series can be represented by a matrix, etc., or can be represented by respective vectors, and is not limited here.

[0040] In an alternative embodiment, a single volume of business data in a historical time series can also be represented by data in multiple dimensions, and each dimension of a single historical time series can be expanded into multi-dimensional data. As an example, if a single volume of business data corresponds to 2D data (such as data in two dimensions of passenger flow and sales volume, and at this time the predetermined business metrics can be two business metrics of passenger flow and sales volume), then the business volume of 60 time periods can correspond to a 120-dimensional vector or array, or a 2×60-dimensional matrix, and so on.

[0041] Next, in step 202, each historical time series is encoded to obtain respective encoded vectors. It can be understood that for a single historical time series, it can describe the specific business volume situation of a single business entity within a certain period. In order to represent it abstractly, the historical time series can be encoded to explore the deep associations between the business volumes in the time series. In a possible design, this step 202 can be implemented by any reasonable model, such as a recurrent neural network (RNN), a neural network based on the attention mechanism, BERT, etc. The model implementing step 202 can be called an encoding module.

[0042] Since historical time series usually have forward and backward correlations and periodicity, therefore, it is preferable to use a recurrent neural network (RNN) to process it. In one embodiment, each historical time series can be processed by a single recurrent neural network. In this case, the historical time series of a single business entity can be regarded as input features corresponding to multiple time points. Through the processing of the recurrent unit of the recurrent neural network, a processing result can be obtained as the encoding result for the corresponding historical time series. In another embodiment, a matrix composed of multiple historical time series can also be processed by an unfolded recurrent neural network (Unrolling RNN). As Figure 3 shown, for example, each historical time series can be described in matrix form, and the recurrent unit of the recurrent neural network is unfolded and tiled into multiple copies (i.e., the recurrent neural network is unfolded). In this way, a single copy of the recurrent unit corresponds to a single business entity, and thus the output result of a single copy of the recurrent unit corresponds to the encoding result of a historical time series. For example, the output of N + 1 copies of the recurrent unit is the encoding results of N + 1 business entities.

[0043] In an alternative implementation, when the data of a single business entity for a single time period is multi-dimensional, the single input data processed by a single copy of the recurrent unit can be the corresponding multi-dimensional data. The encoding result for an encoding object can be in vector form, for example, called an encoded vector. The encoded vector can have a higher dimension than the time series to better describe the time series.

[0044] Then, according to step 203, determine the joint probability distribution satisfied by each coding vector. Among them, the joint probability distribution, abbreviated as joint distribution, is the probability distribution of a random variable composed of two or more random variables. The joint probability distribution can be, for example, a multivariate Gaussian distribution, a multivariate exponential distribution, a multivariate Poisson distribution, a multivariate student-T distribution, and so on. The joint probability distribution in this specification can be a distribution applicable to data statistical scenarios such as a multivariate Gaussian distribution or a multivariate student-T distribution. In this article, a multivariate Gaussian distribution is taken as an example for description. A multivariate Gaussian distribution can be regarded as a joint Gaussian distribution of data in multiple dimensions. For example, N + 1 coding vectors can be regarded as data in N + 1 dimensions, and more coding vectors can be regarded as data in more dimensions. The corresponding joint Gaussian distribution can be constructed by mining the multi-dimensional relationships between each coding vector.

[0045] In this specification, it is assumed that the multivariate Gaussian distribution satisfied by each coding vector is represented in the following form:

[0046]

[0047] Among them, z is a standardized parameter that satisfies the Gaussian distribution (0, 1). The multivariate Gaussian distribution represented by z is a standard multivariate Gaussian distribution. x represents a variable matrix composed of each coding vector, and μ x is the mean matrix of the multivariate Gaussian distribution, Σ represents the covariance matrix corresponding to the variable matrix x, and the element in the i-th row and j-th column is the covariance between variable x i and x j (such as the covariance between the i-th coding vector and the j-th coding vector). n is the total number of business entities in each level of the hierarchical service. For example Figure 1 in the hierarchical service architecture shown, n = 7. In the case where the variables are independent of each other, Σ is a diagonal matrix, and z, μ x and Σ satisfy:

[0048]

[0049] And:

[0050]

[0051] In the standardized multivariate Gaussian distribution, σ i 2 = 1, Σ is an n×n-dimensional identity matrix, each coding vector corresponds to the standard Gaussian distribution N(0, I), and I represents the identity matrix. μ x and Σ are the distribution parameters of the joint Gaussian distribution constructed for each coding vector.

[0052] The process of determining the multivariate Gaussian distribution satisfied by each coding vector can also be regarded as the process of determining the mean matrix μ of the multivariate Gaussian distribution x and the covariance matrix Σ. Generally speaking, the process of determining other joint probability distributions for each coding vector can also be regarded as the process of determining other parameters of other joint probability distributions besides the independent variables.

[0053] Furthermore, in step 204, based on the sampling sequences obtained by sampling each business entity according to the joint probability distribution respectively, determine the prediction sequences for each business entity respectively. In practice, sampling can be performed for each business entity that needs to predict the time series to obtain each sampling sequence. Among them, the business entities that need to predict the time series can be all business entities in the hierarchical business or some business entities, which is not limited here.

[0054] It can be understood that due to some statistical characteristics of the hierarchical structure, in the historical time series, for the lowest-level time series, according to the description of the business scenario above, there is a mutual constraint relationship between the business volumes (or time series) of each business entity. That is, the business volume (or time series) of the business entity in the higher-order level is the sum of the business volumes (or time series) of each business entity corresponding to it in a certain lower-order level. For example Figure 1 in the example shown, the time series corresponding to business entity B is the sum of the time series corresponding to business entities D and E. This can also be called the consistency constraint. In Figure 2 the embodiment shown, the consistency constraint includes: the sum of the prediction sequence corresponding to the first business entity and the prediction sequences of the corresponding N business entities.

[0055] Since the business volumes of business entities in other higher-order levels are all determined by superimposing the business volumes of business entities in the lowest-order level, in a possible design, in the sampling process, first sample the business entities in the lowest-order level, and then sum the corresponding sampling results as the sampling results of the higher-order level. For example Figure 1 in it, directly sample the prediction sequences corresponding to business entities D, E, F, and G from the joint probability distribution, then determine the prediction sequence corresponding to business entity B according to the prediction sequences corresponding to business entities D and E, determine the prediction sequence corresponding to business entity C according to the prediction sequences corresponding to business entities F and G, and then determine the prediction sequence corresponding to business entity A according to the prediction sequences corresponding to business entities B and C. In this way, it can be ensured that each prediction sequence satisfies the consistency constraint.

[0056] According to another possible design, for the randomness of sampling, corresponding sampling sequences can be sampled from the joint probability distribution for each business entity as the preliminary prediction results, and the consistency constraint between business entities can be achieved by adjusting the preliminary prediction results. In particular, according to one embodiment, the sampling results are represented by a sampling matrix composed of each sampling sequence If described, the preliminary prediction results can be adjusted by a predetermined matrix to obtain prediction results that meet the consistency. The matrix used to adjust the preliminary prediction results can be called an optimization matrix. The optimization matrix can process the preliminary prediction results in a predetermined manner to make it meet the consistency constraint, such as Figure 1 the y in A,t = y B,t + y C,t ; y B,t = y D,t + y E,t ; y C,t = y F,t + y G,t .

[0057] According to the above formula, in order to construct the consistency constraint, one side of the equation describing the consistency constraint can be made a constant. For example, move the right side to the left side or move the left side to the right side to make one side equal to 0. In one embodiment, the product of the optimization matrix R and the time series matrix y composed of each time series to be predicted can be used as the consistency constraint, such as denoted as Ry = 0. Each row / column of y can correspond to the time series to be predicted of a business entity. In order to satisfy the consistency constraint when Ry = 0, the optimization matrix R can be constructed according to the hierarchical structure of the business entities (which can be provided in advance by the corresponding business parties). According to the business volume law of the hierarchical business entities, in the process of constructing the optimization matrix, the sum matrix can be constructed first. Since the business entities at the lowest hierarchical level are independent of each other, they can be described by the identity matrix. Taking Figure 1 the implementation architecture shown as an example, the business entities D, E, F, G at the lowest hierarchical level can be described by a 4×4 diagonal matrix, that is:

[0058]

[0059] Assuming that each row of the above matrix represents a business entity, the business entities at the higher hierarchical level can all be represented by the vectors corresponding to the business entities at the lowest hierarchical level. For example, Figure 1 the time series of the business entities A, B, C at the higher hierarchical level in C,t can all be represented by the business entities D, E, F, G at the lowest hierarchical level. For example, according to y F,t = y G,t + y B,t = y D,t + yE,t The vector corresponding to B is obtained as (1, 1, 0, 0), and according to y A,t = y B,t + y C,t the vector corresponding to A is obtained as (1, 1, 1, 1). Thus, the matrix formed by the vectors corresponding to A, B, and C can be called the sum matrix, denoted as the sum matrix S for example, as follows:

[0060]

[0061] To make Ry = 0 hold, an identity matrix representing business entities other than the lowest - order level and the opposite matrix -S of the sum matrix S can be concatenated with the corresponding identity matrix I S In the case where each row of y corresponds to a business entity, the concatenation result is denoted as R = [I S |-S], for example, as follows:

[0062]

[0063] It is easy to understand that R can also be constructed in other forms that satisfy the consistency constraint, such as the concatenation matrix of a negative identity matrix and the sum matrix. Assuming the prediction sequence is y, under the consistency constraint Ry = 0, one row of R describes a constraint. Regarding y as a column vector represented as (y A , y B , y C , y D , y E , y F , y G ), for example, the product of the first row and y describes that the predicted value of business entity A in the first future time period is equal to the sum of the predicted values of the lowest - order level business entities D, E, F, and G (i.e., the difference is 0), and so on, thus forming a consistency constraint on the final time - series matrix y. T , such as the product of the first row and y describes that the predicted value of business entity A in the first future time period is equal to the sum of the predicted values of the lowest - order level business entities D, E, F, and G (i.e., the difference is 0), and so on, thus forming a consistency constraint on the final time - series matrix y.

[0064] In a specific example, a time - series matrix y can be adjusted as little as possible on the basis of such that the time - series matrix y becomes the prediction matrix Y composed of the final predicted time series. For example, by solving based on the constraint Ry = 0 to make the modulus of y (such as ) the smallest, as the prediction matrix Y. Here, the modulus can be represented by the first - order norm, the second - order norm, etc. For example, it is the modulus of the difference between the time - series matrix y and the sampling matrix . It can be understood that here, the prediction matrix in the undetermined state is represented by y, and the determined prediction matrix is represented by Y. Each row / column of Y can describe the time series predicted for a business entity.

[0065] According to a specific method, the Lagrange multiplier method can be used to obtain a closed-form solution:

[0066]

[0067] In other embodiments, the initial prediction results of sampling can also be optimized through other reasonable constraint methods, and accordingly, other forms of optimization matrices and other solution methods are constructed, which will not be elaborated here. When the optimization form is determined, the optimization matrix can be a hyperparameter matrix constructed according to historical time series. In this way, both the randomness of sampling and the consistency of the final prediction results are ensured.

[0068] Among them, the dimension of the sampling result or prediction result can be any dimension. For example, y A , y B , y C , y D , y E , y F , y G can be any mutually consistent dimensions, such as all being 7-dimensional. In other words, the length of the prediction sequence and the length of the historical time series can be independent of each other. For example, the sales volume in 30 days can be used as the historical time series to predict the time series of sales volume in the next 30 days, 10 days, 3 days, 40 days, etc.

[0069] It can be understood that the joint probability distribution obtained through step 203 is a complex distribution. When sampling this distribution, in order to ensure that the prediction process is differentiable and continuous, according to a possible design, the reparameterization trick can be used for sampling. Taking the joint probability distribution as a multivariate Gaussian distribution as an example, according to the reparameterization trick, sampling a variable Z from the distribution N(μ, σ 2 ) is equivalent to sampling a variable ε from the standard distribution N(0, 1), and then letting Z = μ + ε × σ. Specifically for step 204, a variable z can be first sampled from the standard multivariate Gaussian distribution, that is, z ∼ N(0, I), and then converted to the sampling result from the joint multivariate probability distribution, such as: Among them, the dimension of the sampling variable z can be the same as or different from the dimension of the historical time series. The process of sampling from the standard multivariate Gaussian distribution will not be elaborated here. The sampling result can be described in matrix form as: each row / column corresponds to a business entity (corresponding to a vector ), and each column / row corresponds to a future time period. It can be understood that the above multivariate probability distribution is a continuous distribution. Therefore, the sampling process is for a continuous distribution, and gradient-related methods can be used to adjust the model parameters.

[0070] It should be noted that the above prediction process can be implemented through a prediction model, and each step is implemented as each module of the prediction model. For example, step 202 is implemented through an encoding module, step 203 is implemented through a mapping module, and step 204 is implemented through a sampling module. At this time, during the training process of the prediction model, the parameters of the encoding module, the parameters of the mapping module, or the joint probability distribution parameters (optionally, sampling parameters can also be included) can all be regarded as undetermined parameters. Using the historical time series of each business entity in the hierarchical business structure as sample data and the subsequent time series of the historical time series as sample labels, the model loss is determined, and each undetermined parameter is adjusted. The model loss can be measured by various reasonable methods such as variance and cross entropy, which are not limited here. Various appropriate methods can be used to adjust the model parameters. Among them, in a continuously differentiable prediction model, gradient-related methods such as the gradient descent method and the Newton method can be used to adjust each undetermined parameter in the direction of reducing the model loss. Using gradients is usually a better method for parameter adjustment.

[0071] Reviewing the above process, the technical concept provided in this specification constructs a joint probability distribution by using the historical time series of each business entity, integrates the time series information of each hierarchical business entity, fully performs information fusion, and mines the mutual influence of data between business entities, thereby improving the accuracy of time series prediction.

[0072] According to an embodiment of another aspect, a time series prediction device for hierarchical services is also provided. This device is used to predict the time series of the business volume composition of business entities at multiple hierarchical levels on a predetermined business indicator, where each hierarchical level corresponds to at least one business entity.

[0073] Figure 4 A time series prediction device 400 for hierarchical services in an embodiment is shown. As Figure 4 shown, the device 400 includes: an acquisition unit 401 configured to acquire each historical time series corresponding one by one to each business entity in multiple hierarchical levels, where each historical time series corresponds to each business volume of a single business entity corresponding to each of a plurality of sequentially arranged predetermined time periods; an encoding unit 402 configured to encode each historical time series to obtain corresponding encoding vectors; a mapping unit 403 configured to determine the joint probability distribution satisfied by each encoding vector; a prediction unit 404 configured to determine each prediction sequence for each business entity based on each sampling sequence obtained by sampling each business entity according to the joint probability distribution by a sampling module, where each prediction sequence describes each business volume of a single business entity in subsequent multiple predetermined time periods, and each prediction sequence satisfies a consistency constraint, and the consistency constraint includes: for any subsequent time period, the business volume of a single business entity at a higher hierarchical level is consistent with the sum of the business volumes of at least one business entity corresponding to it at a lower hierarchical level.

[0074] According to a possible design, the consistency constraint is described by a predetermined optimization matrix; the prediction unit 404 is further configured to:

[0075] Use the sampling matrix obtained by arranging each sampling sequence in turn as the preliminary prediction result;

[0076] Use the optimization matrix to adjust the preliminary prediction result to satisfy the consistency constraint, so as to obtain each prediction sequence.

[0077] In one embodiment, the optimization matrix is determined based on the hierarchical architectures of each business entity in the multi-level hierarchy, and is formed by splicing the identity matrix for describing the business entity of the lowest level and the negative matrix of the sum matrix for describing the business entities of other levels.

[0078] In one embodiment, the business entity of the lowest level corresponds one-to-one with the rows / columns of the corresponding identity matrix, and a single business entity of each other level corresponds to a single row / column in the sum matrix, and this single row / column is formed by superimposing the rows / columns of the corresponding identity matrix of each business entity corresponding to this single business entity at the lowest level.

[0079] In some alternative implementation manners, the consistency constraint is defined by the product of the optimization matrix and the sampling matrix being 0, and the prediction unit 404 is further configured to:

[0080] Under the consistency constraint, solve the timing matrix that minimizes the difference from the preliminary prediction result as the predicted matrix obtained by prediction;

[0081] Determine each prediction sequence according to the rows / columns in the predicted matrix.

[0082] In one embodiment, the encoding unit 402 can be a single recurrent neural network or an unfolded recurrent neural network composed of multiple recurrent neural networks corresponding one-to-one with each business entity.

[0083] In a possible implementation, the joint probability distribution is a multivariate Gaussian distribution, and the multivariate Gaussian distribution corresponds to a first mean matrix and a first covariance matrix. The apparatus 400 further includes a sampling unit (not shown). The sampling unit is configured to determine a single sampling sequence in the following manner:

[0084] Sample a first vector in the standardized joint probability distribution, where the mean matrix of the standardized joint probability distribution is a 0 matrix and the covariance matrix is an identity matrix;

[0085] Determine a single sampling sequence based on the superimposed result of the product of the square root of the covariance matrix and the first vector and the mean matrix.

[0086] It should be noted that Figure 4 the device 400 shown Figure 2 corresponds to the Figure 2 method described. The corresponding descriptions in the method embodiments of

[0087] also apply to the device 400 and will not be elaborated here. Figure 2 According to an embodiment of another aspect, a computer-readable storage medium is further provided, on which a computer program is stored. When the computer program is executed in a computer, the computer is caused to execute the method described in combination with

[0088] According to an embodiment of still another aspect, a computing device is further provided, including a memory and a processor. An executable code is stored in the memory. When the processor executes the executable code, the method described in combination with Figure 2 is implemented.

[0089] Those skilled in the art should be able to realize that in the above one or more examples, the functions described in the embodiments of this specification can be implemented by hardware, software, firmware, or any combination thereof. When implemented using software, these functions can be stored in a computer-readable medium or transmitted as one or more instructions or codes on a computer-readable medium.

[0090] The specific implementation manners described above further elaborate on the purpose, technical solutions, and beneficial effects of the technical concept of this specification. It should be understood that the above description is only the specific implementation manners of the technical concept of this specification and is not used to limit the protection scope of the technical concept of this specification. Any modifications, equivalent replacements, improvements, etc. made on the basis of the technical solutions of the embodiments of this specification should be included in the protection scope of the technical concept of this specification.

Claims

1. A time series prediction method for hierarchical services, which is used to predict the time series of the business volume composition of business entities at multiple hierarchical levels on a predetermined business indicator, wherein, a single hierarchical level corresponds to at least one business entity; the method includes: obtaining respective historical time series corresponding to each business entity in the multiple hierarchical levels, where a single historical time series includes respective business volumes corresponding to a single business entity in a plurality of predetermined time periods arranged in sequence; encoding each historical time series to obtain respective encoded vectors; determining the joint probability distribution satisfied by each encoded vector; based on respective sampling sequences obtained by sampling each business entity according to the joint probability distribution, determining respective prediction sequences for each business entity, wherein a single prediction sequence describes the respective business volumes of a single business entity in subsequent multiple predetermined time periods, and each prediction sequence satisfies a consistency constraint, and the consistency constraint includes: for any subsequent time period, the business volume of a single business entity at a higher hierarchical level is consistent with the sum of the business volumes of the respective business entities corresponding to it at a lower hierarchical level.

2. The method according to claim 1, wherein, the consistency constraint is described by a pre-determined optimization matrix; the determining respective prediction sequences for each business entity based on respective sampling sequences obtained by sampling each business entity according to the joint probability distribution includes: using a sampling matrix obtained by arranging each sampling sequence in sequence as a preliminary prediction result; adjusting the preliminary prediction result using the optimization matrix to make it satisfy the consistency constraint, so as to obtain respective prediction sequences.

3. The method according to claim 2, wherein, the optimization matrix is determined based on the hierarchical structure of each business entity in the multiple hierarchical levels, and is formed by splicing the identity matrix for describing the business entities at the lowest hierarchical level and the negative opposite matrix of the sum matrix for describing the business entities at other hierarchical levels, or by splicing the negative identity matrix and the sum matrix.

4. The method according to claim 3, wherein, the business entities at the lowest hierarchical level correspond one-to-one with the rows / columns of the corresponding identity matrix, and a single business entity at each other hierarchical level corresponds to a single row / column in the sum matrix, and this single row / column is formed by superimposing the respective rows / columns of the respective business entities corresponding to this single business entity at the lowest hierarchical level in the corresponding identity matrix.

5. The method according to claim 2, wherein, the consistency constraint is defined by the product of the optimization matrix and the sampling matrix being 0, and the adjusting the preliminary prediction result using the optimization matrix to make it satisfy the consistency constraint, so as to obtain respective prediction sequences includes: under the consistency constraint, solving a time series matrix with the smallest difference from the preliminary prediction result as a predicted prediction matrix; determining respective prediction sequences according to the rows / columns in the prediction matrix.

6. The method according to claim 2, wherein, the difference between the time series matrix and the preliminary prediction result is measured by the norm of the difference between the time series matrix and the sampling matrix.

7. The method according to claim 1, Among them, encoding each of the historical time series is implemented by one of the following neural networks: a single recurrent neural network; an unfolded recurrent neural network composed of multiple recurrent neural networks corresponding to each business entity one by one.

8. The method according to claim 1, wherein, the joint probability distribution is a multivariate Gaussian distribution, the multivariate Gaussian distribution corresponds to a first mean matrix and a first covariance matrix, and a single sampling sequence is determined by the following method: sampling a first vector in the standardized multivariate Gaussian distribution, wherein the mean matrix of the standardized multivariate Gaussian distribution is a zero matrix and the covariance matrix is an identity matrix; determining the single sampling sequence based on the superimposed result of the first mean matrix superimposed on the product of the square root of the first covariance matrix and the first vector.

9. A time series prediction device for hierarchical services, which is used to predict the time series composed of the business volumes of business entities at multiple hierarchical levels on a predetermined business indicator, wherein, a single hierarchical level corresponds to at least one business entity; the device includes: an acquisition unit configured to acquire each historical time series corresponding to each business entity in the multiple hierarchical levels, and a single historical time series corresponds to each business volume corresponding to a single business entity in a plurality of sequentially arranged predetermined time periods; an encoding unit configured to encode each historical time series to obtain corresponding encoded vectors; a mapping unit configured to determine the joint probability distribution satisfied by each encoded vector; a prediction unit configured to determine each prediction sequence for each business entity respectively based on each sampling sequence obtained by the sampling module sampling each business entity according to the joint probability distribution, wherein a single prediction sequence describes each business volume of a single business entity in multiple subsequent predetermined time periods, and each prediction sequence satisfies a consistency constraint, and the consistency constraint includes: for any subsequent time period, the business volume of a single business entity at a higher hierarchical level is consistent with the sum of the business volumes of at least one business entity corresponding to it at a lower hierarchical level.

10. A computer-readable storage medium, on which a computer program is stored, and when the computer program is executed in a computer, the computer is made to execute the method according to any one of claims 1-8.

11. A computing device, including a memory and a processor, characterized in that the memory stores executable code, and when the processor executes the executable code, the method according to any one of claims 1-8 is implemented.

Citation Information

Cited By

  • Business state prediction method and device, electronic equipment and storage medium

    CN121029545A

  • Business state prediction method and device, electronic equipment and storage medium

    CN121029545B