Data prediction method and data prediction device

By combining hierarchical models with machine learning algorithms and utilizing a combination of individual and group models, the problem of low accuracy in medium- and long-term network traffic predictions is solved, achieving more accurate network traffic predictions.

CN114077912BActive Publication Date: 2025-09-23HUAWEI TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202010817375.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-08-14
Publication Date
2025-09-23
Estimated Expiration
2040-08-14

AI Technical Summary

Technical Problem

The accuracy of long-term network traffic prediction in existing technologies is low, resulting in the inability to effectively and reasonably plan network resources.

Method used

A hierarchical model is combined with a machine learning algorithm. By combining individual models and group models, predictions are made using the historical data and time characteristics of the target object. The individual model is used to predict the initial value, and the group model is used to correct the deviation value to improve the prediction accuracy.

Benefits of technology

The accuracy of medium- and long-term network traffic prediction is improved, and it can simultaneously meet the individual characteristics of the target object and the group characteristics of the classification to which it belongs, thus optimizing the prediction results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114077912B_ABST
    Figure CN114077912B_ABST
Patent Text Reader

Abstract

The present application provides a data prediction method and device, including: obtaining target time information, the target time information is used to represent the time information corresponding to the prediction of the prediction item of the target object; inputting the target time information into a pre-trained first model to obtain an initial prediction value of the prediction item, the first model is used to predict the initial prediction value of the prediction item corresponding to the time information when the time information is input; inputting the target time identifier into a pre-trained second model to obtain a deviation value of the prediction item, the second model is used to predict the deviation value of the prediction item corresponding to the time identifier when the time identifier is input, the second model is obtained by learning the correlation relationship between multiple residual values ​​and historical time identifiers through a machine learning algorithm or a statistical algorithm; obtaining the prediction result of the prediction item according to the initial prediction value of the prediction item and the deviation value of the prediction item. The solution based on the present application can improve the accuracy of the prediction result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of the combination of machine learning and big data, and specifically, to a data prediction method and a data prediction device. Background Art

[0002] With the continuous advancement of mobile communication technology, the application of communication networks is increasing, and the demand for communication network traffic is also growing. The continuous development and market promotion of new services by operators pose significant challenges to the mobile network experience. Network traffic forecasting is fundamental to resolving network congestion, improving user experience, and rationally allocating and utilizing network resources to increase bandwidth utilization. For example, by predicting future service traffic growth or changes, various professional departments can support decision-making, evaluation, capacity expansion, new construction, and network assurance. Medium- and long-term traffic forecasting, which predicts user network traffic trends over the next three months or longer, is primarily used in scenarios such as operator agile capacity expansion, holiday support, and annual planning.

[0003] Currently, the accuracy of the prediction results for medium- and long-term network traffic is low, which makes it impossible to effectively and rationally plan network resources; therefore, how to improve the accuracy of data prediction methods has become an urgent problem that needs to be solved. Summary of the Invention

[0004] The present application provides a data prediction method and a data prediction device. The data prediction method provided in the embodiments of the present application improves the accuracy of prediction results in medium- and long-term data prediction scenarios.

[0005] In a first aspect, a data prediction method is provided, comprising: obtaining target time information, wherein the target time information is used to represent the time information for predicting a prediction item of a target object; inputting the target time information into a pre-trained first model to obtain an initial prediction value of the prediction item, wherein the first model is used to predict the initial prediction value of the prediction item corresponding to the time information when the time information is input; inputting the target time identifier into a pre-trained second model to obtain a deviation value of the prediction item, wherein the target time identifier is obtained based on the target time information, and the second model is used to predict the deviation value of the prediction item corresponding to the time identifier when the time identifier is input. The deviation value of the predicted item, the second model is obtained by learning the association relationship between multiple residual values ​​and historical time identifiers through a machine learning algorithm or a statistical algorithm, the multiple residual values ​​refer to the residual values ​​output by multiple objects of the same target classification as the target object, the multiple residual values ​​include a first residual value, the first residual value refers to the difference between the initial predicted value of the predicted item obtained by inputting historical time information into the first model and the true value of the predicted item corresponding to the historical time information, the historical time identifier is obtained according to the historical time information; the prediction result of the predicted item is obtained according to the initial predicted value of the predicted item and the deviation value of the predicted item.

[0006] It should be understood that the pre-trained first model can refer to individual models of different target objects, that is, the first model can be trained using the historical data of the prediction project of a target object and the historical time information corresponding to the historical data. The second model can be a group model trained for a class of target objects, that is, the second model refers to learning a unified trend factor based on the class of target objects, believing that the class of target objects has a stronger correlation and a relatively unified change trend, thereby establishing a group model for the class of target objects; the second model can predict the deviation value of the prediction project of the class of target objects when predicting the prediction project. The deviation value can refer to the difference between the predicted value obtained when theoretically predicting the prediction project of the target object and the future true value of the prediction project of the target object.

[0007] In a possible implementation, the prediction item of the target object may refer to the network traffic of the target base station cell.

[0008] For example, the prediction target of the target object is taken as the network traffic of the target base station cell; the first model can be trained based on multiple sample data of the target base station cell, and one sample data in the multiple sample data may include the historical time information of the target base station cell and the network traffic of the target base station cell corresponding to the historical time information.

[0009] Assume that the target classification to which the target base station cell belongs, that is, the same type of base station cells with the same attributes as the target base station cell, includes the target base station cell, base station cell A and base station cell B; then obtain the historical time information A, input the historical time information A into the first model of the target base station cell, and obtain the initial prediction value A of the network traffic corresponding to the historical time information A output by the first model of the target base station cell; similarly, input the historical information A into the first model of base station cell A respectively, and obtain the initial prediction value B of the network traffic corresponding to the historical time information A output by the first model of base station cell A; and input the historical information A into the first model of base station cell B, and obtain the historical time information A output by the first model of base station cell B. The initial predicted value C of the corresponding network traffic can be obtained; at the same time, the actual value A of the historical network traffic corresponding to the historical time information A of the target base station cell, the actual value B of the historical network traffic corresponding to the historical time information A of the base station cell A, and the actual value C of the historical network traffic corresponding to the historical time information A of the base station cell B can also be obtained; further, three residual values ​​can be obtained, namely the residual value between the initial predicted value A of the network traffic and the actual value A of the network traffic, the residual value between the initial predicted value B of the network traffic and the actual value B of the network traffic, and the residual value between the initial predicted value C of the network traffic and the actual value C of the network traffic; a training sample for training the second model may include the historical time identifier A and the above three residual values.

[0010] Among them, the historical time identifier A can be obtained based on the historical time information A; for example, the historical time identifier A can refer to a time sequence number. For example, assuming that the first day's time identifier in the data set used to train the second model can be 0, the historical time identifier A can be determined by the historical time information A and the time base.

[0011] In a possible implementation manner, the prediction item of the target object may refer to the physical resource block utilization rate of the target base station cell.

[0012] In a possible implementation manner, the prediction item of the target object may refer to the number of users in the target base station cell.

[0013] In a possible implementation, the predicted item of the target object may refer to the merchandise sales volume of the target store.

[0014] In a possible implementation, the predicted item of the target object may refer to the webpage traffic of the target website.

[0015] In an embodiment of the present application, when predicting the prediction items of the target object, the initial prediction value output by the pre-trained first model can be used; the initial prediction value output by the first model can be corrected to a certain extent by the pre-trained second model, that is, a deviation value is output; the prediction result obtained by the initial prediction value and the deviation value of the prediction item of the target object can simultaneously meet the individual characteristics of the target object and the group characteristics of the target classification to which the target object belongs, thereby improving the accuracy of the prediction result of the prediction item.

[0016] In one possible implementation, the prediction result of the prediction item can be expressed by the following formula:

[0017] f(t)=C(t)+S(t);

[0018] Where f(t) represents the prediction result, C(t) represents the group trend term, and S(t) represents the individual regularity term. The individual regularity term is obtained based on the first model of the target object, that is, it is fitted based on the historical time series of the target object using a machine learning operator. The group trend term is obtained based on the second model of the target classification corresponding to the target object, that is, it is fitted based on the time-varying difference between the predicted value and the true value for the same type of object using a machine learning algorithm or a statistical algorithm.

[0019] In one possible implementation, the prediction item of the target object may refer to the network traffic of the target base station cell, and the target time information may refer to the time characteristics of the time series to be predicted, wherein the time series to be predicted may refer to the network traffic of the target base station cell in a period of time in the future; for example, with the day as the granularity, the input characteristics of the network traffic of the target base station cell on a certain day in the future are predicted to be the time characteristics corresponding to this day.

[0020] In one possible implementation, the prediction item of the target object may refer to the physical resource block utilization of the target base station cell, and the target time information may refer to the time characteristics of the time series to be predicted, wherein the time series to be predicted may refer to the physical resource block utilization of the target base station cell in a period of time in the future; for example, with the day as the granularity, the input characteristics for predicting the physical resource block utilization of the target base station cell on a certain day in the future are the time characteristics corresponding to this day.

[0021] It should be noted that one or more items included in the target time information may correspond to the time feature tensor corresponding to the historical time information in the training data for training the first model.

[0022] For example, the historical time features may include, but are not limited to, one or more of the following:

[0023] festival: whether it is a festival, 1 if yes, 0 if no;

[0024] Holiday: Is it a holiday? 1 if yes, 0 if no (including holidays and ordinary weekends);

[0025] vacation: whether it is winter or summer vacation, 1 if yes, 0 if no (optional);

[0026] time_corr: day number (incremental indefinitely), starting from 1 and incrementing indefinitely, 1, 2, 3...;

[0027] week_idx: week number (incremental indefinitely), starting from 1 and incrementing continuously, 1, 2, 3...;

[0028] day_in_holiday: the day of the holiday;

[0029] day_in_workday: the day of the workday;

[0030] days_to_next_workday: the number of days to the next workday;

[0031] days_to_next_day_off: the number of days to the next day off;

[0032] length_holiday: length of holiday;

[0033] week_of_month: counts the weeks of the month starting from the 1st of each month, i.e. 1 to 7 for the first week, 8 to 14 for the second week, etc. The value range is 1 to 5.

[0034] In one possible implementation, the first model may be a machine learning model based on a tree ensemble model; for example, the first model may refer to an extreme gradient boosting tree model; a classification gradient boosting tree model; a lightweight gradient booster, etc.

[0035] In a possible implementation, the second model may refer to a linear model, a simple neural network, or other models.

[0036] In conjunction with the first aspect, in some implementations of the first aspect, the method further includes:

[0037] The target objects are classified to obtain the target classification; and the second model is obtained by training based on historical data of the prediction items corresponding to multiple objects included in the target classification.

[0038] In an embodiment of the present application, the target object can be classified to obtain the target category to which the target object belongs; generally, objects of one category have stronger associations, and a unified trend factor of a category of objects can be learned through historical data of multiple objects in the target category; thus, when predicting the prediction item of the target object, the initial prediction value can be adjusted according to the trend factor, thereby improving the accuracy of the prediction result of the prediction item.

[0039] In conjunction with the first aspect, in certain implementations of the first aspect, classifying the target object to obtain the target classification includes:

[0040] The target classification is obtained according to a time series of the target object, where the time series is used to represent a change trend of historical data of the forecast items of the target object over time.

[0041] For example, the target classification of a target object can be determined based on the similarity between its time series and the time series of multiple other objects. A time series is a sequence of values ​​of the same statistical indicator arranged in chronological order. For example, similarity between time series can mean that multiple objects whose predicted items have the same or similar time-varying trends as the target object's predicted items can be assigned to the same target classification as the target object.

[0042] In a possible implementation, multiple objects can be classified by measuring the similarity of their time series. The measurement methods of the similarity of time series include but are not limited to: Pearson correlation, simple Euclidean distance, dynamic time warping measurement, etc.

[0043] In the embodiments of the present application, target objects with similar time series are usually located in similar scenarios and have similar characteristic patterns of prediction items. By classifying the target objects, it is convenient to subsequently learn trend factors of prediction items of a class of target objects.

[0044] In conjunction with the first aspect, in certain implementations of the first aspect, obtaining the target classification according to the time series of the target object includes:

[0045] The target classification is obtained according to the time series of the target object and the spatial features of the target object, wherein the spatial features include the spatial coordinates of the target object, the functional area type of the location of the target object, and the spatial similarity.

[0046] For example, the target classification of the target object may be determined based on the similarity between the time series of the target object and the time series of multiple other objects, and the similarity between the spatial features of the target object and the spatial features of multiple other objects.

[0047] In one possible implementation, spatial features may include but are not limited to: spatial coordinates, functional area types of locations, and spatial similarity; for example, multiple objects with the same or similar land utilization rates as the target object may be classified into the same target category; or, multiple objects with geographical locations close to the target object may be classified into the same target category; or, multiple objects with functional area types close to the geographical locations of the target object may be classified into the same target category; or, multiple objects with spatial similarities in geographical locations of the target object may be classified into the same target category.

[0048] In the embodiments of the present application, target objects with similar time series and spatial features are usually located in similar scenarios, and the characteristic rules of the prediction items are also similar; by classifying the target objects, it is convenient to subsequently learn the trend factors of the prediction items of a class of target objects.

[0049] In combination with the first aspect, in certain implementations of the first aspect, the first model is obtained by training with multiple sample data, wherein one sample training data among the multiple sample training data includes the historical time information and the historical data of the predicted item of the target object corresponding to the historical time information.

[0050] In combination with the first aspect, in certain implementations of the first aspect, the first model is a model obtained through hyperparameter optimization processing, and the hyperparameters in the hyperparameter optimization processing are determined according to the target classification.

[0051] In an embodiment of the present application, in order to improve the accuracy of the first model of the target object, the hyperparameters of the first model may be optimized.

[0052] In one possible implementation, the traffic data distribution for each type of base station cell is similar. In order to further optimize the accuracy of the individual model while taking into account the model performance, in the embodiment of the present application, the first model corresponding to each base station cell in the same type of base station cell is optimized by hyperparameter optimization, that is, the final model of the same type of base station cell ultimately shares the hyperparameters, but the individual models of each base station cell have their own independent internal parameters.

[0053] In combination with the first aspect, in certain implementations of the first aspect, the first model and the second model refer to models of different layers included in the same hierarchical model.

[0054] It should be noted that the first model can refer to an individual model; the second model can refer to a group model; wherein the two layers of individual learning and group learning are coupled with each other; for different target objects, the individual models are independent of each other; for target objects of different categories, the group models are independent of each other.

[0055] In conjunction with the first aspect, in certain implementations of the first aspect, the predicted item of the target object includes any one of the following:

[0056] The network traffic of the target base station cell, the physical resource block utilization rate of the target base station cell, the number of users in the target base station cell, the sales volume of the target store, and the web traffic of the target website.

[0057] In a second aspect, a data prediction device is provided, comprising:

[0058] An acquisition unit is used to acquire target time information, wherein the target time information is used to represent the time information corresponding to the prediction of the prediction item of the target object; a processing unit is used to input the target time information into a pre-trained first model to obtain an initial prediction value of the prediction item, wherein the first model is used to predict the initial prediction value of the prediction item corresponding to the time information when the time information is input; a target time identifier is input into a pre-trained second model to obtain a deviation value of the prediction item, wherein the target time identifier is obtained based on the target time information, and the second model is used to predict the deviation value of the prediction item corresponding to the time identifier when the time identifier is input. The deviation value of the predicted item, the second model is obtained by learning the association relationship between multiple residual values ​​and historical time identifiers through a machine learning algorithm or a statistical algorithm, the multiple residual values ​​refer to the residual values ​​output by multiple objects of the same target classification as the target object, the multiple residual values ​​include a first residual value, the first residual value refers to the difference between the initial prediction value of the predicted item obtained by inputting historical time information into the first model and the true value of the predicted item corresponding to the historical time information, the historical time identifier is obtained according to the historical time information; the prediction result of the predicted item is obtained according to the initial prediction value of the predicted item and the deviation value of the predicted item.

[0059] It should be noted that, in the embodiments of the present application, the data prediction device may refer to a computing device or a chip in a computing device configured in the cloud.

[0060] The computing device may be a device with data prediction capabilities, for example, any device with computing capabilities known in the art, such as a server or computer. Alternatively, the computing device may refer to a chip with computing capabilities, such as a chip in a server or a chip in a computer. The computing device may include memory and a processor. The memory may be used to store program code, and the processor may be used to call the program code stored in the memory to implement the corresponding functions of the computing device. The processor and memory included in the computing device may be implemented as chips, and are not specifically limited here.

[0061] It should be understood that the pre-trained first model can refer to individual models of different target objects, that is, the first model can be obtained by training the historical data of the prediction project of a target object and the time features corresponding to the historical data. The second model can be a group model trained for a class of target objects, that is, the second model refers to learning a unified trend factor based on the class of target objects, believing that the class of target objects has a stronger correlation and a relatively unified change trend, thereby establishing a group model for the class of target objects; the second model can predict the deviation value of the prediction project of the class of target objects when predicting the prediction project. The deviation value can refer to the difference between the initial prediction value of the prediction project of the target object and the true value of the prediction project of the target object.

[0062] In an embodiment of the present application, when predicting the prediction items of the target object, the initial prediction value output by the pre-trained first model can be used; the initial prediction value output by the first model can be corrected to a certain extent by the pre-trained second model, that is, a deviation value is output; the prediction result obtained by the initial prediction value and the deviation value of the prediction item of the target object can simultaneously meet the individual characteristics of the target object and the group characteristics of the target classification to which the target object belongs, thereby improving the accuracy of the prediction result of the prediction item.

[0063] In one possible implementation, the prediction result of the prediction item can be expressed by the following formula:

[0064] f(t)=C(t)+S(t);

[0065] Where f(t) represents the final prediction result, C(t) represents the group trend term, and S(t) represents the individual regularity term. The individual regularity term is obtained based on the first model of the target object, that is, it is fitted based on the historical time series of the target object through a machine learning operator. The group trend term is obtained based on the second model of the target classification corresponding to the target object, that is, it is fitted based on the change of the difference between the predicted value and the true value over time for the same type of object through a machine learning algorithm or a statistical algorithm.

[0066] In one possible implementation, the prediction item of the target object may refer to the network traffic of the target base station cell, and the target time information may refer to the time characteristics of the time series to be predicted; the time series to be predicted may refer to the network traffic of the target base station cell in a period of time in the future; for example, when the granularity is daily, the input characteristics of the network traffic of the target base station cell on a certain day in the future are predicted to be the time characteristics corresponding to this day.

[0067] In one possible implementation, the prediction item of the target object may refer to the physical resource block utilization of the target base station cell, and the target time information may refer to the time characteristics of the time series to be predicted; the time series to be predicted may refer to the physical resource block utilization of the target base station cell in a period of time in the future; for example, with the day as the granularity, the input characteristics for predicting the physical resource block utilization of the target base station cell on a certain day in the future are the time characteristics corresponding to this day.

[0068] It should be noted that historical time information is used as training data when training the first model. The historical time information may refer to the time feature tensor corresponding to each historical time point in the time series; the time feature of the time series to be predicted, that is, the above-mentioned target time information, may include the same tensor as the historical time information.

[0069] Exemplarily, the time features of the time series to be predicted, that is, the above-mentioned target time information, may include one or more of the following: date identification (for example, the day of the year, the day of the month, the week of the month, the day of the week, etc.), day sequence index features (for example, 0, 1, 2, 3...), week sequence index features (for example, 0, 1, 2, 3...), holiday features (for example, whether the time to be predicted is a holiday, the day of the holiday, the length of the holiday, etc.), winter and summer vacation features (for example, whether the time to be predicted is during winter and summer vacation), etc.

[0070] In one possible implementation, the first model may be a machine learning model based on a tree ensemble model; for example, the first model may refer to an extreme gradient boosting tree model; a classification gradient boosting tree model; a lightweight gradient booster, etc.

[0071] In a possible implementation, the second model may refer to a linear model, a simple neural network, or other models.

[0072] In conjunction with the second aspect, in some implementations of the second aspect, the processing unit is further configured to:

[0073] Classifying the target object to obtain the target classification;

[0074] The second model is obtained by training based on historical data of the prediction items corresponding to multiple objects included in the target classification.

[0075] In an embodiment of the present application, the target object can be classified to obtain the target category to which the target object belongs; generally, objects of one category have stronger associations, and a unified trend factor of a category of objects can be learned through historical data of multiple objects in the target category; thus, when predicting the prediction item of the target object, the initial prediction value can be adjusted according to the trend factor, thereby improving the accuracy of the prediction result of the prediction item.

[0076] In conjunction with the second aspect, in some implementations of the second aspect, the processing unit is specifically configured to:

[0077] The target classification is obtained according to a time series of the target object, where the time series is used to represent a change trend of historical data of the forecast items of the target object over time.

[0078] In one possible implementation, the target classification of the target object can be determined based on the similarity between the time series of the target object and the time series of multiple other objects, that is, the objects among the multiple objects that belong to the same target classification as the target object are determined; wherein, the time series refers to a series of values ​​of the same statistical indicator arranged in the order of their occurrence time.

[0079] In the embodiments of the present application, target objects with similar time series are usually located in similar scenarios and have similar characteristic patterns of prediction items. By classifying the target objects, it is convenient to subsequently learn trend factors of prediction items of a class of target objects.

[0080] In conjunction with the second aspect, in some implementations of the second aspect, the processing unit is specifically configured to:

[0081] The target classification is obtained according to the time series of the target object and the spatial features of the target object, wherein the spatial features include the spatial coordinates of the target object, the functional area type of the location of the target object, and the spatial similarity.

[0082] In a possible implementation, the target classification of the target object may be determined based on the similarity between the time series of the target object and the time series of other objects, and the similarity between the spatial features of the target object and the spatial features of other objects.

[0083] In the embodiments of the present application, target objects with similar time series and spatial features are usually located in similar scenarios, and the characteristic rules of the prediction items are also similar; by classifying the target objects, it is convenient to subsequently learn the trend factors of the prediction items of a class of target objects.

[0084] In combination with the second aspect, in certain implementations of the second aspect, the first model is obtained by training with multiple sample data, wherein one sample training data among the multiple sample training data includes the historical time information and the historical data of the predicted item of the target object corresponding to the historical time information.

[0085] In combination with the second aspect, in certain implementations of the second aspect, the first model is a model obtained through hyperparameter optimization processing, and the hyperparameters in the hyperparameter optimization processing are determined according to the target classification.

[0086] In an embodiment of the present application, in order to improve the accuracy of the first model of the target object, the hyperparameters of the first model may be optimized.

[0087] In one possible implementation, the timing patterns of traffic data for each type of base station cell are similar. In order to further optimize the accuracy of the individual model while taking into account the model performance, the first model corresponding to each base station cell in the same type of base station cell is optimized through hyperparameter optimization in an embodiment of the present application. That is, the first models corresponding to the same type of base station cells ultimately share hyperparameters, but the individual models of each base station cell have their own independent internal parameters.

[0088] In combination with the second aspect, in certain implementations of the second aspect, the first model and the second model refer to models of different layers included in the same hierarchical model.

[0089] It should be noted that the first model can refer to an individual model; the second model can refer to a group model; wherein the two layers of individual learning and group learning are coupled with each other; for different target objects, the individual models are independent of each other; for target objects of different categories, the group models are independent of each other.

[0090] In conjunction with the second aspect, in certain implementations of the second aspect, the predicted item of the target object includes any one of the following:

[0091] The network traffic of the target base station cell, the physical resource block utilization rate of the target base station cell, the number of users in the target base station cell, the sales volume of the target store, and the web traffic of the target website.

[0092] According to a third aspect, a data prediction device is provided, comprising: a memory for storing a program; a processor for executing the program stored in the memory, wherein when the program stored in the memory is executed, the processor is configured to: obtain target time information, wherein the target time information is used to represent the time information corresponding to the prediction of the prediction item of the target object; input the target time information into a pre-trained first model to obtain an initial prediction value of the prediction item, wherein the first model is used to predict the initial prediction value of the prediction item corresponding to the time information when the time information is input; input the target time identifier into a pre-trained second model to obtain a deviation value of the prediction item, wherein the target time identifier is obtained based on the target time information, and the first model is used to input the target time identifier into a pre-trained second model to obtain a deviation value of the prediction item. The second model is used to predict the deviation value of the predicted item corresponding to the time identifier when a time identifier is input. The second model is obtained by learning the association relationship between multiple residual values ​​and historical time identifiers through a machine learning algorithm or a statistical algorithm. The multiple residual values ​​refer to the residual values ​​output by multiple objects of the same target classification as the target object. The multiple residual values ​​include a first residual value. The first residual value refers to the difference between the initial prediction value of the predicted item obtained by inputting historical time information in the first model and the true value of the predicted item corresponding to the historical time information. The historical time identifier is obtained based on the historical time information; the prediction result of the predicted item is obtained based on the initial prediction value of the predicted item and the deviation value of the predicted item.

[0093] In a possible implementation manner, the processor included in the above-mentioned apparatus is further configured to execute the data prediction method in any implementation manner of the first aspect.

[0094] It should be understood that the expansion, limitation, explanation and description of the relevant content in the above-mentioned first aspect also apply to the same content in the third aspect.

[0095] In a fourth aspect, a computer-readable medium is provided, which stores a program code for execution by a device, wherein the program code includes a method for executing the data prediction method in the above-mentioned first aspect and any one of the implementations of the first aspect.

[0096] In a fifth aspect, a computer program product comprising instructions is provided. When the computer program product is run on a computer, the computer is caused to execute the data prediction method in the above-mentioned first aspect and any one of the implementations of the first aspect.

[0097] In a sixth aspect, a chip is provided, comprising a processor and a data interface, wherein the processor reads instructions stored in a memory through the data interface and executes the data prediction method in the above-mentioned first aspect and any one of the implementation methods of the first aspect.

[0098] Optionally, as an implementation method, the chip may further include a memory, in which instructions are stored, and the processor is used to execute the instructions stored in the memory. When the instructions are executed, the processor is used to execute the data prediction method in the above-mentioned first aspect and any one of the implementation methods of the first aspect. BRIEF DESCRIPTION OF THE DRAWINGS

[0099] Figure 1 Schematic diagram of the system architecture of the data prediction method provided in the embodiment of the present application;

[0100] Figure 2 This is a schematic diagram of a system architecture provided by an embodiment of the present application;

[0101] Figure 3 is a schematic diagram of the hardware structure of a chip provided in an embodiment of the present application;

[0102] Figure 4 is a schematic diagram of a system architecture for applying the data prediction method according to an embodiment of the present application;

[0103] Figure 5 is a schematic flow chart of the data prediction method provided in an embodiment of the present application;

[0104] Figure 6 is a schematic flow chart of a training method for a network traffic prediction model provided in an embodiment of the present application;

[0105] Figure 7 Schematic diagram of a training method for a network traffic prediction model provided in an embodiment of the present application;

[0106] Figure 8 is a schematic diagram of a network traffic prediction method provided by an embodiment of the present application;

[0107] Figure 9 is a schematic block diagram of the data prediction device provided by this application;

[0108] Figure 10 Schematic diagram of the hardware structure of the data prediction device provided in the embodiment of the present application. DETAILED DESCRIPTION

[0109] The following will describe the technical solutions in the embodiments of this application in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0110] For example, Figure 1It is a schematic diagram of the system architecture provided in the embodiment of the present application.

[0111] like Figure 1 As shown, the system 100 includes a classification module 110 , a hierarchical modeling module 120 and a prediction module 130 .

[0112] The classification module 110 is used to classify multiple objects (eg, base station cells, stores, or websites).

[0113] In one example, taking the classification of base station cells as an example, multiple base station cells can be classified based on the similarity of their time series; wherein, the time series refers to a series of values ​​of the same statistical indicator arranged in the order of their occurrence time; the similarity of the time series may refer to the similarity of the trend of the value of the same statistical indicator changing over time.

[0114] For example, the objects can be classified by measuring the similarity of their time series. The measurement methods of the similarity of time series include but are not limited to: Pearson correlation, simple Euclidean distance, dynamic time warping measurement, etc.

[0115] In another example, taking the classification of base station cells as an example, the base station cells can be classified based on the similarity of the time series and the similarity of the spatial features of multiple base station cells; wherein the spatial features may include the spatial coordinates of the base station cells, the functional area type of the location of the base station cells, and the spatial similarity between a certain base station cell and other base station cells.

[0116] It should be understood that the above-mentioned classification process of multiple objects can adopt an unsupervised classification method, the goal of which is to classify objects with similar characteristics into one category, and then further learn the trend factor of the category of objects. Exemplarily, the hierarchical modeling module 120 is used to learn individual rules and group trends respectively; wherein, the learning of individual rules refers to establishing an individual model of a base station cell based on the correlation between the historical traffic data and the time series of the cell; the learning of group trends refers to learning a unified trend factor based on a category of base station cells, believing that the cells of a category have a stronger correlation and a relatively unified change trend, thereby establishing a group model for a category of cells.

[0117] It should be noted that the hierarchical modeling module 120 may include two levels: individual learning and group learning. Individual learning may refer to learning the changing trend of the predicted items of a target object over time based on the sample time characteristics of the target object. Group learning may refer to learning the deviation between the predicted value and the true value of a class of target objects in the evaluation set. The two layers of individual learning and group learning are mutually coupled. For example, the difference between the predicted value and the true value output by each individual model of a class of target objects can be used to learn the group model. Each layer is relatively independent. For example, for different target objects, the individual models are independent of each other; for target objects of different categories, the group models are independent of each other.

[0118] For example, taking the target object as a base station cell as an example, the individual model of a base station cell can be trained based on multiple sample data of the base station cell. One sample data in the multiple sample data may include the sample value of the prediction item of the base station cell at a certain historical moment and the time feature corresponding to the certain historical moment.

[0119] For example, the time characteristics of a base station cell may also be referred to as a time characteristic tensor, and may include but not be limited to the following characteristics:

[0120] Date identifier (for example, the day of the year, the day of the month, the week of the month, the day of the week, etc.), daily sequential index features (for example, 0, 1, 2, 3...), weekly sequential index features (for example, 0, 1, 2, 3...), holiday features (for example, whether it is a holiday, the day of the holiday, the length of the holiday, etc.), winter and summer vacation features (for example, whether it is winter or summer vacation).

[0121] Exemplarily, in an embodiment of the present application, the individual model can be a tree ensemble model obtained by machine learning; for example, the individual model can refer to an extreme gradient boosting tree model (extreme gradient boosting, XGBoost); a categorical gradient boosting tree model (categorical boosting, CatBoost); a light gradient boosting machine (LightGBM), etc.

[0122] It should be understood that for each of the multiple base station cells, individual modeling can be performed using machine learning operators, and the internal parameters of each model can be fitted based on the training data. For example, a group model for a type of base station cell can refer to a model that learns the difference between the predicted value and the actual value of the type of base station cell in the evaluation set and establishes the relationship between this difference and time. The group trend model can be used to represent a linear trend between the difference and time, or a more complex method can be used to fit a nonlinear trend.

[0123] For example, in the embodiments of the present application, the group model may use a linear model, a simple neural network, or other models. Furthermore, in the embodiments of the present application, in order to improve the accuracy of the individual models of each base station cell, the hyperparameters of the individual models may be optimized.

[0124] Exemplarily, the time series patterns of traffic data for each type of base station cell are similar. In order to further optimize the accuracy of individual models while taking into account model performance, in an embodiment of the present application, the first model corresponding to each base station cell in the same type of base station cell is optimized through hyperparameter optimization, that is, the first models corresponding to the same type of base station cells ultimately share hyperparameters, but the individual models of each base station cell have their own independent internal parameters.

[0125] It should be noted that, compared to the internal parameters of the individual models, hyperparameters can be set before machine learning training; for example, the number of trees in the extreme gradient boosting tree model can be a hyperparameter, while the internal parameters of the individual models are model parameters learned during training.

[0126] For example, when performing hyperparameter optimization, an evaluation data set with a time span of at least one month can be used, with the largest proportion of cells with a mean absolute percentage error (MAPE) of weekly average traffic less than 20% being used as the optimization target (20% is a typical value, and can also be set to 15% or other values ​​based on business objectives); the hyperparameter optimization method can select the Bayesian hyperparameter optimization method, or the grid hyperparameter optimization method, etc.

[0127] Exemplarily, the prediction module 130 is configured to predict the prediction items of the target object based on the time characteristics of the base station cell.

[0128] For example, the prediction items of the target object include any one of the following: network traffic of the target base station cell, physical resource block utilization of the target base station cell, number of users of the target base station cell, sales traffic of the target store, and network traffic of the target website.

[0129] It should be understood that the above is an example of the prediction items of the target object. The data prediction method provided in the embodiment of the present application is applicable to various medium- and long-term data prediction scenarios. This application does not make any specific limitations on the prediction items of the target object.

[0130] Figure 2 A system architecture 200 provided by an embodiment of the present application is shown.

[0131] exist Figure 2 In the embodiment of the present application, the data acquisition device 260 is used to collect training data. For the pre-trained first model of the embodiment of the present application, the first model can be trained by the training data collected by the data acquisition device 260.

[0132] Exemplarily, in an embodiment of the present application, for the first model, the training data may include a historical time series and historical values ​​of predicted items of the target object corresponding to the historical time series.

[0133] Exemplarily, in an embodiment of the present application, for the second model, the training data may include multiple residual values ​​and historical time identifiers corresponding to the multiple residual values, wherein the multiple residual values ​​refer to the residual values ​​output by multiple objects of the same target classification as the target object, and the multiple residual values ​​include a first residual value, which refers to the deviation between the initial prediction value of the prediction item of the target object obtained by inputting the historical time identifier in the first model and the true value of the prediction item of the target object.

[0134] After collecting the training data, the data collection device 260 stores the training data in the database 230 , and the training device 220 obtains the target model / rule 201 through training based on the training data maintained in the database 230 .

[0135] The following describes how the training device 220 obtains the target model / rule 201 based on the training data.

[0136] For example, the training device 220 processes the input first model training data, and compares the predicted initial value of the predicted item of the target object output by the first model with the true value of the predicted item of the target object until the difference between the predicted initial value output by the training device 220 and the true value is less than a certain threshold, thereby completing the training of the first model.

[0137] It should be noted that, in actual applications, the training data maintained in the database 230 may not necessarily be collected by the data collection device 260, but may also be received from other devices.

[0138] It should also be noted that the training device 220 does not necessarily train the target model / rule 201 entirely based on the training data maintained by the database 230. It is also possible to obtain training data from the cloud or other places for model training. The above description should not be used as a limitation on the embodiments of the present application.

[0139] The target model / rule 201 obtained by training the training device 220 can be applied to different systems or devices, such as Figure 2 The execution device 210 shown in the figure may be a terminal, such as a mobile phone terminal, a tablet computer, a laptop computer, an augmented reality (AR) / virtual reality (VR), a vehicle terminal, etc. It may also be a server, or a cloud server, etc. Figure 2 In the embodiment of the present application, the execution device 210 is configured with an input / output (I / O) interface 212 for data interaction with an external device. The user can input data to the I / O interface 212 through the client device 240. The input data may include: training samples input by the client device.

[0140] The preprocessing module 213 and the preprocessing module 214 are used to preprocess the input data received by the I / O interface 212; in an embodiment of the present application, the preprocessing module 213 and the preprocessing module 214 may be omitted (or only one of the preprocessing modules may be present), and the calculation module 211 may be directly used to process the input data.

[0141] When the execution device 210 preprocesses the input data, or when the computing module 211 of the execution device 210 performs calculations and other related processing, the execution device 210 can call the data, code, etc. in the data storage system 250 for corresponding processing, and can also store the data, instructions, etc. obtained from the corresponding processing in the data storage system 250.

[0142] Finally, the I / O interface 212 returns the processing result, such as the prediction result of the predicted item of the target object, to the client device 240 to provide it to the user.

[0143] It is worth noting that the training device 220 can generate corresponding target models / rules 201 based on different training data for different goals or different tasks. The corresponding target models / rules 201 can be used to achieve the above goals or complete the above tasks, thereby providing users with the desired results.

[0144] exist Figure 2In the case shown in FIG, in one case, the user can manually give input data, and the manual giving can be operated through the interface provided by the I / O interface 212 .

[0145] In another embodiment, the client device 240 can automatically send input data to the I / O interface 212. If the client device 240 requires user authorization to automatically send input data, the user can set the corresponding permissions in the client device 240. The user can view the results output by the execution device 210 on the client device 240. The specific presentation form can be a display, sound, action, or other specific methods. The client device 240 can also serve as a data acquisition terminal, collecting the input data input into the I / O interface 212 and the output results output from the I / O interface 212 as new sample data, and storing them in the database 230. Of course, the collection can also be performed without going through the client device 240, and the I / O interface 212 can directly store the input data input into the I / O interface 212 and the output results output from the I / O interface 212 as new sample data in the database 230.

[0146] It is worth noting that Figure 2 This is only a schematic diagram of a system architecture provided by an embodiment of the present application, and the positional relationship between the devices, components, modules, etc. shown in the figure does not constitute any limitation. Figure 2 In the embodiment, the data storage system 250 is an external memory relative to the execution device 210; in other cases, the data storage system 250 can also be placed in the execution device 210.

[0147] For example, the first model in this application may refer to Figure 1 Individual models shown; for example, the first model may refer to a tree ensemble model; the second model may refer to Figure 1 The group model shown; for example, the second model can be a linear model or a neural network model.

[0148] Figure 3 This is a schematic diagram of the hardware structure of a chip provided in an embodiment of the present application.

[0149] like Figure 3 As shown, the chip includes a neural-network processing unit 300 (NPU); the chip can be set as Figure 2 The execution device 210 shown in FIG. 2 is used to complete the calculation work of the calculation module 211. The chip can also be set in Figure 2 The training device 220 shown is used to complete the training work of the training device 220 and output the target model / rule 201.

[0150] The neural network processor 300 is mounted on the main central processing unit (CPU) as a coprocessor, and tasks are assigned by the main CPU; the core part of the NPU300 is the operation circuit 303, and the controller 304 controls the operation circuit 303 to extract data from the memory (weight memory or input memory) and perform operations.

[0151] In some implementations, the computing circuit 303 includes multiple processing engines (PEs) therein.

[0152] In some implementations, the arithmetic circuit 303 is a two-dimensional systolic array; the arithmetic circuit 203 may also be a one-dimensional systolic array or other electronic circuitry capable of performing mathematical operations such as multiplication and addition.

[0153] In some implementations, the arithmetic circuit 303 is a general-purpose matrix processor.

[0154] For example, assume there are input matrix A, weight matrix B, and output matrix C. Operation circuit 303 retrieves the corresponding data of matrix B from weight memory 302 and caches it on each PE in operation circuit 303. Operation circuit 303 obtains the matrix A data from input memory 301 and performs matrix operations on matrix B. The partial or final results of the matrix are stored in accumulator 308. Vector calculation unit 307 can further process the output of operation circuit 303, such as vector multiplication, vector addition, exponential operation, logarithmic operation, size comparison, etc.

[0155] For example, the vector calculation unit 307 can be used for network calculations of non-convolutional / non-FC layers in a neural network, such as pooling, batch normalization, local response normalization, etc.

[0156] In some implementations, the vector calculation unit 307 can store the processed output vector to the unified memory 306. For example, the vector calculation unit 307 can apply a nonlinear function to the output of the operation circuit 303; for example, accumulate a vector of values ​​to generate an activation value.

[0157] In some implementations, the vector calculation unit 307 generates normalized values, merged values, or both.

[0158] In some implementations, the vector of processed outputs can be used as activation input to the operational circuit 303, eg, for use in subsequent layers in a neural network.

[0159] For example, the unified memory 306 can be used to store input data and output data. The weight data is directly stored in the input memory 301 and / or the unified memory 306 through the direct memory access controller 305 (DMAC), the weight data in the external memory is stored in the weight memory 302, and the data in the unified memory 306 is stored in the external memory.

[0160] For example, the bus interface unit 310 (BIU) may be used to implement interaction between the main CPU, the DMAC, and the instruction fetch memory 309 via a bus.

[0161] For example, an instruction fetch buffer 309 connected to the controller 304 can be used to store instructions used by the controller 304. The controller 304 can be used to call instructions cached in the instruction fetch buffer 309 to control the operation process of the computing accelerator.

[0162] Generally, the unified memory 306, the input memory 301, the weight memory 302 and the instruction fetch memory 309 can all be on-chip memories; the external memory is a memory outside the NPU, which can be a double data rate synchronous dynamic random access memory (DDR SDRAM), a high bandwidth memory (HBM) or other readable and writable memory.

[0163] It should be noted that the operations in the first model and the second model in the embodiment of the present application can be performed by the operation circuit 303 or the vector calculation unit 307.

[0164] Currently, the accuracy of the network traffic prediction values ​​obtained from medium- and long-term network traffic prediction is low, which makes it impossible to effectively and rationally plan network resources.

[0165] In view of this, the present application proposes a data prediction method, which, when predicting the prediction items of the target object, can be based on the initial prediction value output by the pre-trained first model; the pre-trained second model can correct the initial prediction value output by the first model to a certain extent, that is, output a deviation value; the prediction result obtained by the initial prediction value and the deviation value of the prediction item of the target object can simultaneously meet the individual characteristics of the target object and the group characteristics of the target classification to which the target object belongs, thereby improving the accuracy of the prediction item of the prediction item object.

[0166] Figure 4 The system architecture 400 for applying the data prediction method of the embodiment of the present application may include a local device 420, a local device 430, an execution device 410, and a data storage system 450, wherein the local device 420 and the local device 430 may be connected to the execution device 410 via a communication network.

[0167] Execution device 410 can be implemented by one or more servers. Optionally, execution device 410 can be used in conjunction with other computing devices, such as data storage devices, routers, and load balancers. Execution device 410 can be deployed at a single physical site or distributed across multiple physical sites. Execution device 410 can use data in data storage system 450 or invoke program code in data storage system 450 to implement the data prediction method of the present embodiment.

[0168] Exemplarily, the data storage system 450 may be deployed in the local device 420 or the local device 430 ; for example, the data storage system 450 may be used to store user behavior logs.

[0169] It should be noted that the above-mentioned execution device 410 can also be called a cloud device, and in this case the execution device 410 can be deployed in the cloud.

[0170] Specifically, the execution device 410 can perform the following process: obtaining target time information, where the target time information is used to represent the time information corresponding to the prediction of the prediction item of the target object; inputting the target time information into a pre-trained first model to obtain an initial prediction value of the prediction item, wherein the first model is used to predict the initial prediction value of the prediction item corresponding to the time information when the time information is input; inputting the target time identifier into a pre-trained second model to obtain a deviation value of the prediction item, wherein the target time identifier is obtained based on the target time information, and the second model is used to predict the time identifier for the prediction item when the time identifier is input. The deviation value of the predicted item corresponding to the second model is obtained by learning the association relationship between multiple residual values ​​and historical time identifiers through a machine learning algorithm or a statistical algorithm. The multiple residual values ​​refer to the residual values ​​output by multiple objects of the same target classification as the target object. The multiple residual values ​​include a first residual value. The first residual value refers to the difference between the initial prediction value of the predicted item obtained by inputting historical time information into the first model and the true value of the predicted item corresponding to the historical time information. The historical time identifier is obtained based on the historical time information; the prediction result of the predicted item is obtained based on the initial prediction value of the predicted item and the deviation value of the predicted item.

[0171] Through the above process execution device 410 can obtain the pre-trained first model and the pre-trained second model through training, and obtain the prediction result of the prediction item of the target object.

[0172] In a possible implementation, the above-mentioned execution device 410 method may be an offline method executed in the cloud.

[0173] For example, after operating their respective user devices (e.g., local device 420 and local device 430), users can store operation logs in data storage system 450, and execution device 410 can call the data in data storage system 450 to complete the training process of the first model and the second model. Each local device can represent any computing device, such as a personal computer, a computer workstation, a smart phone, a tablet computer, a smart camera, a smart car or other type of cellular phone, a media consumption device, a wearable device, a set-top box, a game console, etc. Each user's local device can interact with execution device 410 through a communication network with any communication mechanism / communication standard. The communication network can be a wide area network, a local area network, a point-to-point connection, etc., or any combination thereof.

[0174] In one implementation, the local device 420 and the local device 430 can obtain relevant parameters of the pre-trained first model and the pre-trained second model from the execution device 410, and use the pre-trained first model and the pre-trained second model on the local device 420 and the local device 430 to predict the prediction items of the target object and obtain the prediction results.

[0175] In another implementation, a pre-trained first model and a pre-trained second model can be directly deployed on the execution device 410. The execution device 410 obtains target time information from the local device 420 and the local device 430, and obtains prediction results of the prediction items of the target object based on the pre-trained first model and the second model.

[0176] Exemplarily, the data storage system 450 may be deployed in the local device 420 or the local device 430 to store user behavior logs of the local device.

[0177] Exemplarily, the data storage system 450 can be independent of the local device 420 or the local device 430 and deployed separately on a storage device. The storage device can interact with the local device to obtain the user's behavior log in the local device and store it in the storage device.

[0178] The following combination Figures 5 to 8 The embodiments of the present application are described in detail.

[0179] Figure 5 It is a schematic flow chart of the data prediction method provided in the embodiment of the present application. Figure 5 The method 500 shown includes steps S510 to S540, and steps S510 to S540 are described in detail below.

[0180] S510: Obtain target time information.

[0181] The target time information may be used to indicate the time information for predicting the prediction item of the target object.

[0182] It should be understood that the above-mentioned target time information may refer to the time feature of the time series to be predicted, that is, the time feature tensor corresponding to the time point corresponding to the predicted item of the predicted item object.

[0183] In one possible implementation, the prediction item of the target object may refer to the network traffic of the target base station cell, and the target time information may refer to the time characteristics of the time series to be predicted; the time series to be predicted may refer to the network traffic of the target base station cell in a period of time in the future; for example, when the granularity is daily, the input characteristics of the network traffic of the target base station cell on a certain day in the future are predicted to be the time characteristics corresponding to this day.

[0184] In one possible implementation, the prediction item of the target object may refer to the physical resource block utilization of the target base station cell, and the target time information may refer to the time characteristics of the time series to be predicted; the time series to be predicted may refer to the physical resource block utilization of the target base station cell in a period of time in the future; for example, with the day as the granularity, the input characteristics for predicting the physical resource block utilization of the target base station cell on a certain day in the future are the time characteristics corresponding to this day.

[0185] Optionally, in one possible implementation, the prediction items of the target object may include any one of the following: the network traffic of the target base station cell, the physical resource block (PRB) utilization rate of the target base station cell, the number of users of the target base station cell, or the sales volume of goods in the target store and the web traffic of the target website, etc.

[0186] It should be understood that the above is an example of the prediction items of the target object. The data prediction method provided in the embodiment of the present application is applicable to various medium- and long-term data prediction scenarios. This application does not make any specific limitations on the prediction items of the target object.

[0187] Exemplarily, one or more features included in the target time information correspond to the time features corresponding to the historical time information in the training data used when training the first model; that is, the historical time information used to train the first model may refer to the time feature tensor corresponding to each historical time point in the time series; the target time information may be the same tensor included in the historical time information.

[0188] For example, the historical time information may include, but is not limited to, one or more of the following:

[0189] festival: whether it is a festival, 1 if yes, 0 if no;

[0190] Holiday: Is it a holiday? 1 if yes, 0 if no (including holidays and ordinary weekends);

[0191] vacation: whether it is winter or summer vacation, 1 if yes, 0 if no (optional);

[0192] time_corr: day number (incremental indefinitely), starting from 1 and incrementing indefinitely, 1, 2, 3...;

[0193] week_idx: week number (incremental indefinitely), starting from 1 and incrementing continuously, 1, 2, 3...;

[0194] day_in_holiday: the day of the holiday;

[0195] day_in_workday: the day of the workday;

[0196] days_to_next_workday: the number of days to the next workday;

[0197] days_to_next_day_off: the number of days to the next day off;

[0198] length_holiday: length of holiday;

[0199] week_of_month: counts the weeks of the month starting from the 1st of each month, i.e. 1 to 7 for the first week, 8 to 14 for the second week, etc. The value range is 1 to 5.

[0200] S520: Input the target time information into a pre-trained first model to obtain an initial prediction value of the prediction item.

[0201] The first model is used to predict the initial prediction value of the prediction item corresponding to the time information when the time information is input.

[0202] It should be noted that the pre-trained first model can refer to individual models of different target objects, that is, the first model can be obtained by training the historical data of the prediction project of a target object and the historical time information corresponding to the historical data. The first model can be used to represent the correlation between the initial prediction value of the prediction project and the corresponding time characteristics.

[0203] Optionally, in a possible implementation, the first model is obtained by training multiple sample data, wherein one sample training data among the multiple sample training data includes a sample value of a prediction item of a certain historical time information and a sample time feature corresponding to the historical time information. The sample time feature corresponding to the historical time information can also be called a sample time feature tensor.

[0204] For example, taking the target base station cell of the target object as an example, the time feature of a certain day can be used as input data (i.e., training input feature X), and the downlink traffic value of the target base station cell on that day (or indicators such as PRB utilization and number of users) can be used as output data (i.e., training label y), thereby forming a training sample.

[0205] S530: Input the target time identifier into the pre-trained second model to obtain a deviation value of the predicted item.

[0206] Among them, the target time identifier is obtained based on the target time information, and the second model is used to predict the deviation value of the predicted item corresponding to the time identifier when the time identifier is input. The second model is obtained by learning the association relationship between multiple residual values ​​and historical time identifiers through a machine learning algorithm or a statistical algorithm. The multiple residual values ​​refer to the residual values ​​output by multiple objects with the same target classification as the target object. The multiple residual values ​​include a first residual value. The first residual value refers to the difference between the initial predicted value of the predicted item obtained by inputting the historical time information in the first model and the true value of the predicted item corresponding to the historical time information. The historical time identifier is obtained based on the historical time information.

[0207] It should be understood that the second model can predict the deviation value of a type of target object when predicting the prediction item. The deviation value may refer to the difference between the initial prediction value obtained when theoretically predicting the prediction item of the target object and the future true value of the prediction item of the target object.

[0208] Exemplarily, the process of training the second model is described with an example. Assume that the target classification to which the target base station cell belongs, that is, the same type of base station cells with the same attributes as the target base station cell, includes the target base station cell, base station cell A and base station cell B; then obtain historical time information A, input historical time information A into the first model of the target base station cell, and obtain the initial prediction value A of the network traffic corresponding to the historical time information A output by the first model of the target base station cell; similarly, input historical information A into the first model of base station cell A respectively, and obtain the initial prediction value B of the network traffic corresponding to the historical time information A output by the first model of base station cell A; and input historical information A into the first model of base station cell B, and obtain the historical time information A output by the first model of base station cell B. The initial predicted value C of the corresponding network traffic can be obtained; at the same time, the actual value A of the historical network traffic corresponding to the historical time information A of the target base station cell, the actual value B of the historical network traffic corresponding to the historical time information A of the base station cell A, and the actual value C of the historical network traffic corresponding to the historical time information A of the base station cell B can also be obtained; further, three residual values ​​can be obtained, namely the residual value between the initial predicted value A of the network traffic and the actual value A of the network traffic, the residual value between the initial predicted value B of the network traffic and the actual value B of the network traffic, and the residual value between the initial predicted value C of the network traffic and the actual value C of the network traffic; a training sample for training the second model may include the historical time identifier A and the above three residual values.

[0209] Among them, the historical time identifier A can be obtained based on the historical time information A; for example, the historical time identifier A can be a time serial number. For example, the first day time identifier in the data set used to train the second model can be 0, and the historical time identifier A can be determined by the historical time information A and the time base.

[0210] It should be noted that the specific steps of the training process of the second model can be found in the subsequent Figure 6 Step S630 in is not described again here.

[0211] It should be understood that the second model can be a model trained for a class of target objects (for example, base station cells), that is, the second model refers to learning a unified trend factor based on a class of target objects, believing that a class of target objects has a stronger correlation and has a relatively unified change trend, thereby establishing a group model for a class of target objects; the second model can predict the deviation value of a class of target objects when predicting the prediction item, and the deviation value can refer to the difference between the initial prediction value of the prediction item of the target object and the true value of the prediction item of the target object.

[0212] Exemplarily, the target time identifier may be a time sequence number. For example, the time identifier of the first day in the evaluation data set used to train the second model may be 0. Then the time identifier may be determined by the time in the time series of the target to be tested relative to the time reference in the evaluation set.

[0213] For example, the first day of the evaluation set is June 1st, and the corresponding time identifier is 0; the time obtained from the target time information is June 10th, so the target time identifier can be 9, that is, the time identifier of June 10th relative to the time base June 1st.

[0214] Optionally, in a possible implementation, the method includes: classifying the target object to obtain a target classification; and training based on historical data of prediction items corresponding to multiple objects included in the target classification to obtain a second model.

[0215] Exemplarily, the second model is taken as an example of a group model corresponding to a type of base station cell; the group model of the same type of base station cell can be trained by learning the difference between the predicted value and the true value of the same type of cell in the evaluation set to obtain the relationship between the difference and time; the true value can refer to the historical network traffic usage value of the sample base station cell in the evaluation set; the predicted value can refer to the predicted traffic value obtained by inputting the time series corresponding to the true value into the first model of the base station cell.

[0216] Optionally, in a possible implementation, classifying the target object to obtain the target classification includes: obtaining the target classification based on a time series of the target object, wherein the time series is used to represent a change trend of historical data of the prediction item of the target object over time.

[0217] For example, the target classification of the target object can be determined based on the similarity between the time series of the target object and the time series of multiple other objects, that is, the objects among multiple objects that belong to the same target classification as the target object can be determined; among them, the time series refers to a series of values ​​​​of the same statistical indicator arranged in the order of their occurrence time.

[0218] Exemplarily, the target object may refer to a target base station cell, and the traffic trends of each base station cell in a plurality of base station cells may be clustered according to the similarity of the time series of each base station cell through a kmeans algorithm based on dynamic time warping, thereby dividing the plurality of base station cells into several categories of base station cells; that is, the target classification of the target base station cell may be determined through the time series of the target base station cell.

[0219] Optionally, in a possible implementation, in addition to the above-mentioned classification by time series, the spatial characteristics of the target object can also be added, that is, the target classification of the target object can be determined based on the time series and spatial characteristics of the target object, wherein the spatial characteristics can include the spatial coordinates of the target object, the functional area type of the target object's location, and spatial similarity, etc.

[0220] For example, the target classification of the target object may be determined based on the similarity between the time series of the target object and the time series of multiple other objects, and the similarity between the spatial features of the target object and the spatial features of multiple other objects.

[0221] For details, please refer to the following classification processing steps. Figure 6 Step S620 in is not described again here.

[0222] Furthermore, in an embodiment of the present application, in order to improve the accuracy of each first model, the hyperparameters of the first model may be optimized.

[0223] Optionally, in a possible implementation, the first model is a model obtained through hyperparameter optimization processing.

[0224] Exemplarily, for each type of target object, the data distribution of the predicted items is similar. In order to further optimize the accuracy of the first model while taking into account the model performance, in the embodiment of the present application, the first model of the target object is optimized by adopting the classification hyperparameter optimization method, that is, the models corresponding to multiple objects included in the same target classification can share hyperparameters, but the models corresponding to each object have their own independent internal parameters.

[0225] It is important to note that, in contrast to the internal parameters of a model, hyperparameters are set before machine learning training; the number of trees in the Extreme Gradient Boosting Trees model is a hyperparameter, while the weights in a neural network are model parameters learned during training.

[0226] Optionally, in a possible implementation, the first model and the second model refer to models at different levels included in the same hierarchical model.

[0227] In the embodiments of the present application, the first model and the second model may refer to two independent models, or the first model and the second model may refer to two strongly coupled models; Figure 1 As shown, the first model can refer to an individual model; the second model can refer to a group model; wherein the two layers of individual learning and group learning are coupled with each other; for different target objects, the individual models are independent of each other; for target objects of different categories, the group models are independent of each other.

[0228] S540: Obtain a prediction result for the prediction item according to the initial prediction value of the prediction item and the deviation value of the prediction item.

[0229] It should be understood that when predicting the prediction items of the target object, the initial prediction value output by the pre-trained first model can be used as the basis; the pre-trained second model can correct the initial prediction value output by the first model to a certain extent, that is, output a deviation value. The prediction result obtained by the initial prediction value and the deviation value of the prediction item of the target object can simultaneously meet the individual characteristics of the target object and the group characteristics of the target classification to which the target object belongs, thereby improving the accuracy of the prediction item of the prediction item object.

[0230] The data prediction method provided in the embodiment of the present application may include a model training phase and a model online prediction phase. Figures 6 to 8 The two stages are described in detail.

[0231] Training phase:

[0232] Figure 6 is a schematic flow chart of a training method for a network traffic prediction model provided by an embodiment of the present application. The training method 600 can be executed by a training device; for example, the training method can be executed by Figure 2The execution device 210 in the embodiment executes, or Figure 4 The execution device 410 in the embodiment of the present invention may also execute the execution device 410, or the execution device 410 may ... including steps S610 to S640, which are described in detail below.

[0233] S610: Obtain training data.

[0234] Exemplarily, the training data may refer to historical data of sample targets; for example, the sample targets may refer to network traffic, physical resource block (PRB) utilization, number of users, etc. of a base station cell.

[0235] It should be noted that Figure 6 The example of the network traffic prediction scenario as a sample target is used for illustration. The training method of the embodiment of the present application can also be applied to models of other medium- and long-term prediction scenarios, including but not limited to: sales forecast models, website traffic forecast models and other multi-time series medium- and long-term prediction scenarios.

[0236] In one example, for a network traffic prediction scenario, the acquired training data may include: historical traffic values ​​of multiple sample base station cells and time features corresponding to the historical traffic values.

[0237] In one example, in a scenario of store sales prediction, the acquired training data may include: historical sales values ​​of multiple sample stores and time features corresponding to the historical sales values.

[0238] For example, the historical traffic value obtained is the historical network traffic usage of a sample base station cell on a certain day. The time characteristics may include but are not limited to one or more of the following:

[0239] festival: whether it is a festival, 1 if yes, 0 if no;

[0240] Holiday: Is it a holiday? 1 if yes, 0 if no (including holidays and ordinary weekends);

[0241] vacation: whether it is winter or summer vacation, 1 if yes, 0 if no (optional);

[0242] time_corr: day number (incremental indefinitely), starting from 1 and incrementing indefinitely, 1, 2, 3...;

[0243] week_idx: week number (incremental indefinitely), starting from 1 and incrementing continuously, 1, 2, 3...;

[0244] day_in_holiday: the day of the holiday;

[0245] day_in_workday: the day of the holiday;

[0246] days_to_next_workday: the days of the holiday;

[0247] days_to_next_day_off: the days of the holiday;

[0248] length_holiday: length of holiday;

[0249] week_of_month: counts the weeks of the month starting from the 1st of each month, i.e. 1 to 7 for the first week, 8 to 14 for the second week, etc. The value range is 1 to 5.

[0250] For example, Figure 7 As shown, a sample time feature combination tensor can be generated based on the acquired training data, i.e., the input data; wherein, the sample time feature combination tensor may include one or more of the above-mentioned time dimension features; in addition, the combination tensor may also include spatial dimension features, for example, spatial dimension features may include: land use, lighting, and information points (POI), etc. Each POI may contain four aspects of information: name, category, coordinates, and classification. Comprehensive POI information can remind users of detailed information about road branches and surrounding buildings.

[0251] It should be understood that in the sales forecasting scenario, the time series corresponding to the historical sales values ​​may also include one or more of the above features.

[0252] Furthermore, based on the sample time feature combination tensor of each sample base station cell and the historical traffic value corresponding to the sample time feature combination tensor, the correlation relationship between the historical traffic value and the time feature of each sample base station cell in multiple cells can be obtained. This correlation relationship can be used to represent the changing trend of the network traffic usage of a sample base station cell over time.

[0253] Step 620: Classification processing.

[0254] The classification process refers to clustering multiple sample base station cells.

[0255] In one example, the category to which the target base station cell belongs can be determined based on the similarity between the time series of the target base station cell and the time series of multiple other sample base station cells, that is, the sample base station cells among the multiple sample base station cells that belong to the same target category as the target base station cell are determined; wherein, the time series refers to a series in which the numerical values ​​of the same statistical indicator are arranged in the chronological order of their occurrence.

[0256] It should be understood that base station cells with similar time series usually have similar scenarios and similar characteristics of network traffic usage values; by clustering multiple sample base station cells, it is convenient to subsequently learn the traffic trend factors of a class of sample base station cells; usually, sample base station cells of the same class have a stronger correlation and have a relatively unified network traffic change trend; therefore, the goal of classification processing is to classify sample base station cells with similar time series into the same class.

[0257] In another example, multiple sample base station cells can be classified according to the similarity of time series and the similarity of spatial features; that is, the category to which the target base station cell belongs can be determined based on the similarity between the time series of the target base station cell and the time series of multiple other sample base station cells, and the similarity between the spatial features of the target base station cell and the spatial features of multiple other sample base station cells; wherein the spatial features may include the spatial coordinates of the target base station cell, the functional area type of the location of the target base station cell, and the spatial similarity between the target base station cell and other base station cells.

[0258] Exemplarily, the traffic trends of each sample base station cell in multiple sample base station cells can be clustered using a kmeans algorithm based on dynamic time warping (DTW), that is, according to the historical traffic values ​​of each sample base station cell and the correlation relationship between the time characteristics corresponding to the historical traffic values, so that the multiple sample base station cells can be divided into several categories of base station cells.

[0259] To determine the optimal number of categories, the optimal k value can be determined by using the kmeans sum of squared errors (SSE) for both n and (n+1) to be less than a threshold (n > 1, with a threshold of 10%). Alternatively, the SSE SSE can be replaced by the SSE SSE. Furthermore, clustering algorithms can employ the k-shape algorithm to improve computational efficiency.

[0260] It should be noted that when spatial features are added to the classification process, the k-shape algorithm may no longer be applicable; other unsupervised classification methods can be used, such as the kmeans algorithm or density-based spatial custering of applications with noise (DBSCAN).

[0261] S630: Hierarchical model training.

[0262] It should be noted that hierarchical model training includes training individual models for each sample base station cell and training group models for the same type of cells; the individual model is used to represent the correlation between the network traffic value of a sample base station cell and the corresponding time characteristics; the group model is used to represent the correlation between the network traffic trend and time of a type of cell. The group model learns a unified trend factor based on the same type of base station cells, and believes that cells of the same type have stronger correlations and relatively unified change trends, thereby establishing a group model for the same type of cells.

[0263] It should be understood that the individual model of the target base station cell may refer to Figure 5 The first model shown in the target base station; the group model corresponding to the target classification may refer to Figure 5 The second model is shown.

[0264] For example, the individual model of each sample base station cell may be trained using a historical traffic value of the sample base station cell and a time series corresponding to the historical traffic value to obtain a correlation between the time series and the network traffic.

[0265] For example, the individual model corresponding to each sample base station cell can use the time characteristics of a certain day as input data (i.e., training input feature X) and the downlink traffic value of that day (or PRB utilization, number of users, and other indicators) as output data (i.e., training label y), thereby forming a training sample.

[0266] like Figure 7 As shown, for base station cell C1, individual model 1 can be obtained through training, and for base station cell C2, individual model 2 can be obtained through training; similarly, for base station cell Cm, individual model m can be obtained through training.

[0267] Exemplarily, the group model of the same type of base station cells can be trained by learning the difference between the predicted value and the true value of the same type of cells in the evaluation set to obtain the relationship between the difference and time; the true value can refer to the historical network traffic usage value of the sample base station cells obtained in the evaluation set; the predicted value can refer to the predicted traffic value obtained by inputting the time series corresponding to the true value into the individual model of the base station cell.

[0268] For example, a population model can be used to represent a linear trend of differences over time, or a more complex approach can be used to fit a nonlinear trend.

[0269] For example, a population model can be used to represent a functional relationship over time without an intercept term, and can be expressed as s*f(t), where s is the coefficient and t is the time sequence number starting from the evaluation set (i.e., the first day of the evaluation set, t=0) (the granularity is the same as the original data; if the original data is daily, then this is the day number). f(t) can be in linear or logarithmic form. The parameter s can be obtained by fitting or by optimizing the population model through grid search.

[0270] Furthermore, in an embodiment of the present application, in order to improve the accuracy of the individual model of each base station cell, the hyperparameters of the individual model may be optimized.

[0271] For example, for each type of target object, the data distribution of the predicted items is similar. In order to further optimize the accuracy of the first model while taking into account the model performance, in the embodiment of the present application, the first model of the target object is optimized by adopting the method of classification super-parameter optimization, that is, the models corresponding to multiple objects included in the same target classification can share hyper-parameters, but the models corresponding to each object have their own independent internal parameters.

[0272] It should be understood that, in contrast to the internal parameters of the individual models, hyperparameters can be set prior to machine learning training; for example, the number of trees in an extreme gradient boosted tree model can be a hyperparameter, while the internal parameters of the individual models are model parameters learned during training.

[0273] For example, when performing classification hyperparameter optimization, an evaluation data set with a time span of at least one month can be used, with the largest proportion of cells with a mean absolute percentage error (MAPE) of weekly average traffic less than 20% being used as the optimization target (20% is a typical value, and can also be set to 15% or other values ​​based on business objectives); the hyperparameter optimization method can select the Bayesian hyperparameter optimization method, or the grid hyperparameter optimization method, etc.

[0274] In one possible implementation, if the computational efficiency of the network traffic in the base station cell is high, hyperparameter optimization of the individual model may not be performed; for example, the individual model can use a machine learning operator similar to catboost, and the default parameters can already achieve good prediction results.

[0275] S640: Model training is completed.

[0276] In an embodiment of the present application, a pre-trained individual model and a pre-trained group model can be obtained through the above-mentioned training method; when predicting the network traffic of the target base station cell, the predicted value of the network traffic of the target base station cell can be obtained based on the pre-trained individual model corresponding to the above-mentioned target base station cell and the pre-trained group model corresponding to the classification to which the target base station cell belongs. The network traffic prediction value of the target base station cell is obtained based on the individual prediction value output by the pre-trained individual model and the group trend value output by the pre-trained group model; the individual prediction value output by the individual model can be corrected to a certain extent through the group trend item, that is, through hierarchical learning, the network traffic prediction result of the target base station cell satisfies both individual characteristics and group characteristics, thereby improving the accuracy of network traffic prediction.

[0277] Prediction stage:

[0278] Figure 8 is a schematic diagram of the network traffic prediction method provided by the embodiment of the present application. The training method 700 can be executed by a prediction device; for example, the method can be performed by Figure 2 The execution device 210 in the embodiment executes, or Figure 4 Alternatively, it may be executed by the execution device 410 in the process, or it may be executed by the local device 420. The process includes steps S710 to S740, which are described in detail below.

[0279] S710: Obtain the time characteristics of the time series to be predicted, that is, obtain target time information.

[0280] Exemplarily, the prediction items of the target object may refer to the network traffic of the target base station cell, the physical resource block utilization of the target base station cell, the number of users of the target base station cell; or it may be the sales volume of goods in the target store and the web page traffic of the target website, etc.

[0281] It should be noted that Figure 8 The network traffic prediction scenario is used as an example. The prediction method of the embodiment of the present application can also be applied to other medium- and long-term prediction scenarios, including but not limited to: sales forecast, website traffic forecast and other multi-time series medium- and long-term prediction scenarios.

[0282] In one example, for the sales volume of goods in target stores, the target stores can be classified according to their time series and spatial characteristics, and then hierarchical prediction can be performed to obtain prediction results.

[0283] In one example, when predicting the webpage traffic of a target website, although there are generally no spatial features, historical time series can be directly used for classification and then hierarchical prediction.

[0284] For example, one or more features included in the time features of the time series to be predicted correspond to the time features corresponding to historical moments in the training data used when training the individual model; therefore, the predicted time features may include, but are not limited to, one or more of the following:

[0285] festival: whether it is a festival, 1 if yes, 0 if no;

[0286] Holiday: Is it a holiday? 1 if yes, 0 if no (including holidays and ordinary weekends);

[0287] vacation: whether it is winter or summer vacation, 1 if yes, 0 if no (optional);

[0288] time_corr: day number (incremental indefinitely), starting from 1 and incrementing indefinitely, 1, 2, 3...;

[0289] week_idx: week number (incremental indefinitely), starting from 1 and incrementing continuously, 1, 2, 3...;

[0290] day_in_holiday: the day of the holiday;

[0291] day_in_workday: the day of the workday;

[0292] days_to_next_workday: the number of days to the next workday;

[0293] days_to_next_day_off: the number of days to the next day off;

[0294] length_holiday: length of holiday;

[0295] week_of_month: counts the weeks of the month starting from the 1st of each month, i.e. 1 to 7 for the first week, 8 to 14 for the second week, etc. The value range is 1 to 5.

[0296] In other words, the time feature combination tensor of the target base station cell can be generated based on the need to predict whether a certain day in the future of the network traffic is a festival, a holiday, a winter or summer vacation, the day number, the week number, the day number of the holiday, the day number of the holiday, the day number of the holiday, the day number of the holiday, the length of the holiday, and the number of weeks in the month.

[0297] S720: Target individual model processing.

[0298] For example, the pre-trained individual model corresponding to the target base station cell can be determined through the target base station cell; the time characteristics of the acquired time series to be predicted can be input into the target individual model to obtain the initial prediction value of the prediction item, that is, the individual item prediction value; the individual item prediction value can be used to represent the predicted traffic value of the target base station cell in the future time based on the historical time characteristics of the target base station cell and the historical traffic values ​​corresponding to the historical time characteristics, which are output by fitting the traffic trend through machine learning operators.

[0299] S730: Target group model processing.

[0300] For example, a target group model corresponding to the target base station cell may be determined based on the target base station cell; for example, multiple base station cells may be classified during the training phase to obtain pre-trained group models corresponding to base station cells of the same type.

[0301] Exemplarily, the category to which the target base station cell belongs may be determined according to the identification information of the target base station cell; and then the target group model may be determined according to the category to which the target base station cell belongs.

[0302] Furthermore, the target time identifier can be determined based on the time characteristics of the time series to be predicted, and the target time identifier can be input into the pre-trained target group model to obtain the deviation value of the prediction item, namely the group trend prediction value; the group trend prediction value can be used to represent the difference between the prediction value output by the target individual model and the actual traffic usage value of the target base station cell. The group trend prediction value is fitted by the difference between the prediction value and the actual value of a certain type of base station cell machine learning operator over time.

[0303] It should be understood that the target group model is used to represent the association between the time identifier and the deviation value, wherein the deviation value refers to the difference between the network traffic prediction value of the base station cell by the individual model and the network true value of the base station cell; therefore, the time identifier can be a time serial number. For example, the time identifier of the first day in the evaluation data set used to train the group model can be 0, and the time identifier can be determined by the time in the time series of the target to be measured relative to the time base in the evaluation set.

[0304] For example, the first day of the evaluation set is June 1st, and the corresponding time identifier is 0; according to the time characteristics of the time series to be predicted, the time obtained is June 10th, so the target time identifier can be 9, that is, the time identifier of June 10th relative to the time base June 1st.

[0305] It should be noted that S730 and S740 may be executed simultaneously, or S740 may be executed first and then S730. This application does not impose any limitation on the order of S730 and S740.

[0306] S740: Output the prediction result.

[0307] Exemplarily, the network traffic corresponding to the time feature of the time series to be predicted of each target base station cell is obtained by adding the predicted value of the regularity item of the target base station cell and the predicted value of the trend item of the category to which the target base station cell belongs.

[0308] For example, the prediction result of the target base station cell can be expressed by the following formula:

[0309] f(t)=C(t)+S(t);

[0310] Where f(t) represents the final traffic prediction value, C(t) represents the group trend term, and S(t) represents the individual regularity term. The individual regularity term is obtained by fitting the time series of the target base station cell using a machine learning operator, while the group trend term is fitted by the time-varying difference between the predicted value and the actual value of the machine learning operator for the same type of cell.

[0311] In an embodiment of the present application, when predicting the network traffic of a target base station cell, the network traffic prediction value of the target base station cell is obtained based on the individual prediction value output by a pre-trained individual model and the group trend value output by a pre-trained group model; the individual prediction value output by the individual model can be corrected to a certain extent through the group trend item, that is, through hierarchical learning, the network traffic prediction result of the target base station cell satisfies both individual characteristics and group characteristics, thereby improving the accuracy of network traffic prediction.

[0312] It should be understood that the above examples are intended to help those skilled in the art understand the embodiments of the present application, and are not intended to limit the embodiments of the present application to the specific numerical values ​​or specific scenarios illustrated. Those skilled in the art can obviously make various equivalent modifications or variations based on the above examples, and such modifications or variations also fall within the scope of the embodiments of the present application.

[0313] Combined with the above Figures 1 to 8 , describes in detail the data prediction method provided by the embodiment of the present application; the following will be combined with Figures 9 and 10 , describing the device embodiments of the present application in detail. It should be understood that the data prediction device in the embodiments of the present application can execute the various data prediction methods in the aforementioned embodiments of the present application, that is, the specific working processes of the various products below can refer to the corresponding processes in the aforementioned method embodiments.

[0314] Figure 9 It is a schematic block diagram of the data prediction device provided in this application.

[0315] It should be understood that the data prediction device 800 can execute Figures 5 to 8The data prediction device 800 includes an acquisition unit 810 and a processing unit 820.

[0316] Among them, the acquisition unit 810 is used to obtain the target time information, and the target time information is used to represent the time information corresponding to the prediction of the prediction item of the target object; the processing unit 820 is used to input the target time information into a pre-trained first model to obtain the initial prediction value of the prediction item, wherein the first model is used to predict the initial prediction value of the prediction item corresponding to the time information when the time information is input; the target time identifier is input into a pre-trained second model to obtain the deviation value of the prediction item, wherein the target time identifier is obtained based on the target time information, and the second model is used to predict the time identifier corresponding to the time identifier when the time identifier is input. The deviation value of the predicted item corresponding to the second model is obtained by learning the association relationship between multiple residual values ​​and historical time identifiers through a machine learning algorithm or a statistical algorithm. The multiple residual values ​​refer to the residual values ​​output by multiple objects of the same target classification as the target object. The multiple residual values ​​include a first residual value. The first residual value refers to the difference between the initial prediction value of the predicted item obtained by inputting historical time information into the first model and the true value of the predicted item corresponding to the historical time information. The historical time identifier is obtained based on the historical time information; the prediction result of the predicted item is obtained based on the initial prediction value of the predicted item and the deviation value of the predicted item.

[0317] Optionally, as an embodiment, the processing unit 820 is further configured to:

[0318] Classifying the target object to obtain the target classification;

[0319] The second model is obtained by training based on historical data of the prediction items corresponding to multiple objects included in the target classification.

[0320] Optionally, as an embodiment, the processing unit 820 is specifically configured to:

[0321] The target classification is obtained according to a time series of the target object, where the time series is used to represent a change trend of historical data of the forecast items of the target object over time.

[0322] Optionally, as an embodiment, the processing unit 820 is specifically configured to:

[0323] The target classification is obtained according to the time series of the target object and the spatial features of the target object, wherein the spatial features include the spatial coordinates of the target object, the functional area type of the location of the target object, and the spatial similarity.

[0324] Optionally, as an embodiment, the first model is obtained by training multiple sample data, wherein one sample training data among the multiple sample training data includes the historical time information and the historical data of the predicted item of the target object corresponding to the historical time information.

[0325] Optionally, as an embodiment, the first model is a model obtained through hyperparameter optimization processing, and the hyperparameters in the hyperparameter optimization processing are determined according to the target classification.

[0326] Optionally, as an embodiment, the first model and the second model refer to models of different layers included in the same hierarchical model.

[0327] Optionally, as an embodiment, the prediction item of the target object includes any one of the following:

[0328] The network traffic of the target base station cell, the physical resource block utilization rate of the target base station cell, the number of users in the target base station cell, the sales volume of the target store, and the web traffic of the target website.

[0329] It should be noted that the processing device 800 is implemented in the form of a functional unit. The term "unit" here can be implemented in the form of software and / or hardware, and is not specifically limited to this.

[0330] For example, a "unit" may be a software program, a hardware circuit, or a combination of the two that implements the aforementioned functionality. The hardware circuit may include an application specific integrated circuit (ASIC), an electronic circuit, a processor (e.g., a shared processor, a dedicated processor, or a group processor) and memory for executing one or more software or firmware programs, combined logic circuits, and / or other suitable components that support the described functionality.

[0331] Therefore, the units of each example described in the embodiments of this application can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0332] Figure 10 Schematic diagram of the hardware structure of the data prediction device provided in the embodiment of the present application.

[0333] Figure 10The data prediction device 900 shown (the data prediction device 900 may be a computer device) includes a memory 910, a processor 920, a communication interface 930, and a bus 940. The memory 910, the processor 920, and the communication interface 930 are connected to each other via the bus 940.

[0334] The memory 910 may be a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 910 may store a program. When the program stored in the memory 910 is executed by the processor 920, the processor 920 is used to execute the various steps of the data prediction method of the embodiment of the present application; for example, executing Figures 5 to 8 The steps shown.

[0335] It should be understood that the data prediction device shown in the embodiment of the present application can be a computing device or a chip configured in a computing device in the cloud.

[0336] The computing device may be a device with data prediction capabilities, for example, any device with computing capabilities known in the art, such as a server or computer. Alternatively, the computing device may refer to a chip with computing capabilities, such as a chip in a server or a chip in a computer. The computing device may include memory and a processor. The memory may be used to store program code, and the processor may be used to call the program code stored in the memory to implement the corresponding functions of the computing device. The processor and memory included in the computing device may be implemented as chips, and are not specifically limited here.

[0337] For example, the memory can be used to store relevant program instructions of the data prediction method provided in the embodiments of the present application, and the processor can be used to call the relevant program instructions of the data prediction method stored in the memory.

[0338] The processor 920 can be a general-purpose central processing unit (CPU), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits to execute relevant programs to implement the data prediction method of the method embodiment of the present application.

[0339] The processor 920 may also be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the data prediction method of the present application may be completed by hardware integrated logic circuits in the processor 920 or software instructions.

[0340] The above-mentioned processor 920 can also be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. The various methods, steps and logic block diagrams disclosed in the embodiments of this application can be implemented or executed. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in the embodiments of this application can be directly reflected as being executed by a hardware decoding processor, or can be executed by a combination of hardware and software modules in the decoding processor. The software module can be located in a mature storage medium in the field such as random access memory, flash memory, read-only memory, programmable read-only memory or electrically erasable programmable memory, registers, etc. The storage medium is located in the memory 910, and the processor 920 reads the information in the memory 910, and completes the implementation of this application in combination with its hardware. Figure 9 The functions required to be performed by the units included in the data prediction device shown, or the functions required to perform the method embodiments of the present application Figures 5 to 8 The data prediction method shown.

[0341] The communication interface 930 uses a transceiver device such as, but not limited to, a transceiver to implement communication between the data prediction device 900 and other devices or a communication network.

[0342] The bus 940 may include a path for transmitting information between various components of the data prediction apparatus 900 (eg, the memory 910 , the processor 920 , and the communication interface 930 ).

[0343] It should be noted that although the above-mentioned data prediction device 900 only shows a memory, a processor, and a communication interface, in the specific implementation process, those skilled in the art should understand that the data prediction device 900 may also include other devices necessary for normal operation. At the same time, according to specific needs, those skilled in the art should understand that the above-mentioned data prediction device 900 may also include hardware devices that implement other additional functions. In addition, those skilled in the art should understand that the above-mentioned data prediction device 900 may also only include the devices necessary to implement the embodiments of the present application, and does not necessarily include Figure 10 All devices shown in .

[0344] For example, an embodiment of the present application further provides a chip comprising a transceiver unit and a processing unit. The transceiver unit may be an input / output circuit or a communication interface; the processing unit may be a processor, microprocessor, or integrated circuit integrated on the chip; and the chip may execute the data prediction method described in the above method embodiment.

[0345] Illustratively, an embodiment of the present application further provides a computer-readable storage medium having instructions stored thereon, which, when executed, execute the data prediction method in the above method embodiment.

[0346] Illustratively, an embodiment of the present application further provides a computer program product comprising instructions, which, when executed, execute the data prediction method in the above method embodiment.

[0347] It should be understood that the processor in the embodiments of the present application may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc.

[0348] It should also be understood that the memory in the embodiments of the present application may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of random access memory (RAM) are available, such as static RAM (SRAM), dynamic random access memory (DRAM), synchronous DRAM (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link DRAM (SLDRAM), and direct rambus RAM (DR RAM).

[0349] The above embodiments can be implemented in whole or in part by software, hardware, firmware or any other combination. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer program are loaded or executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center via a wired (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that contains one or more available media sets. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium. The semiconductor medium can be a solid-state drive.

[0350] It should be understood that the term "and / or" as used herein simply describes a relationship between associated objects, indicating that three possible relationships exist. For example, "A and / or B" can represent: A alone, A and B together, or B alone. A and B can be singular or plural. Furthermore, the character " / " as used herein generally indicates an "or" relationship between the associated objects, but it may also indicate an "and / or" relationship. For specific understanding, please refer to the context.

[0351] In this application, "at least one" means one or more, and "plurality" means two or more. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, or c can mean: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or plural.

[0352] It should be understood that in the various embodiments of the present application, the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.

[0353] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0354] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0355] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0356] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0357] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0358] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.

[0359] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.

Claims

1. A data prediction method, characterized in that: include: Acquire target time information, where the target time information is used to indicate the time information corresponding to the prediction of the prediction item of the target object; Inputting the target time information into a pre-trained first model to obtain an initial prediction value of the prediction item, wherein the first model is used to predict the initial prediction value of the prediction item corresponding to the time information when the time information is input, and the first model is obtained by training with multiple sample data, wherein one sample training data among the multiple sample training data includes the historical time information and the historical data of the prediction item corresponding to the historical time information; Input the target time identifier into a pre-trained second model to obtain a deviation value of the predicted item, wherein the target time identifier is obtained based on the target time information, and the second model is used to predict the deviation value of the predicted item corresponding to the time identifier when the time identifier is input, and the second model is obtained by learning the association relationship between multiple residual values ​​and historical time identifiers through a machine learning algorithm or a statistical algorithm, and the multiple residual values ​​refer to the residual values ​​output by multiple objects of the same target classification as the target object, and the multiple residual values ​​include a first residual value, and the first residual value refers to the difference between an initial predicted value of the predicted item obtained by inputting historical time information into the first model and a true value of the predicted item corresponding to the historical time information, and the historical time identifier is obtained based on the historical time information; Correcting the initial prediction value of the prediction item according to the deviation value of the prediction item to obtain a prediction result of the prediction item; The target classification is obtained based on the time series of the target object and the spatial characteristics of the target object, and the prediction items of the target object include any one of the following: network traffic, number of users and utilization rate of physical resource blocks.

2. The method according to claim 1, wherein Also includes: Classifying the target object to obtain the target classification; The second model is obtained by training based on historical data of the prediction items corresponding to multiple objects included in the target classification.

3. The method according to claim 2, wherein The classifying the target object to obtain the target classification includes: The target classification is obtained according to the time series of the target object, where the time series is used to represent a change trend of historical data of the forecast items of the target object over time.

4. The method according to claim 3, wherein Obtaining the target classification according to the time series of the target object includes: The target classification is obtained according to the time series of the target object and the spatial features of the target object, wherein the spatial features include the spatial coordinates of the target object, the functional area type of the location of the target object, and the spatial similarity.

5. The method according to claim 1, wherein The first model is a model obtained through hyperparameter optimization processing, and the hyperparameters in the hyperparameter optimization processing are determined according to the target classification.

6. The method according to claim 1, wherein The first model and the second model refer to models of different layers included in the same hierarchical model.

7. The method according to any one of claims 1 to 6, characterized in that The network traffic includes the network traffic of the target base station cell, the physical resource block utilization includes the physical resource block utilization of the target base station cell, and the number of users includes the number of users of the target base station cell.

8. A data prediction device, characterized in that: include: An acquiring unit, configured to acquire target time information, wherein the target time information is used to indicate time information corresponding to a prediction item of a target object; A processing unit is used to input the target time information into a pre-trained first model to obtain an initial prediction value of the prediction item, wherein the first model is used to predict the initial prediction value of the prediction item corresponding to the time information when the time information is input, and the first model is obtained by training with multiple sample data, wherein one sample training data among the multiple sample training data includes the historical time information and the historical data of the prediction item corresponding to the historical time information; input the target time identifier into a pre-trained second model to obtain a deviation value of the prediction item, wherein the target time identifier is obtained based on the target time information, and the second model is used to predict the deviation value of the prediction item corresponding to the time identifier when the time identifier is input. , the second model is obtained by learning the association relationship between multiple residual values ​​and historical time identifiers through a machine learning algorithm or a statistical algorithm, the multiple residual values ​​refer to the residual values ​​output by multiple objects of the same target classification as the target object, the multiple residual values ​​include a first residual value, the first residual value refers to the difference between the initial prediction value of the prediction item obtained by inputting historical time information into the first model and the true value of the prediction item corresponding to the historical time information, the historical time identifier is obtained based on the historical time information; the initial prediction value of the prediction item is corrected according to the deviation value of the prediction item to obtain the prediction result of the prediction item; wherein, the prediction item of the target object includes any one of the following: network traffic, number of users and utilization rate of physical resource blocks.

9. The device according to claim 8, wherein The processing unit is further configured to: Classifying the target object to obtain the target classification; The second model is obtained by training based on historical data of the prediction items corresponding to multiple objects included in the target classification.

10. The device according to claim 9, wherein The processing unit is specifically configured to: The target classification is obtained according to a time series of the target object, where the time series is used to represent a change trend of historical data of the forecast items of the target object over time.

11. The device according to claim 10, wherein The processing unit is specifically configured to: The target classification is obtained according to the time series of the target object and the spatial features of the target object, wherein the spatial features include the spatial coordinates of the target object, the functional area type of the location of the target object, and the spatial similarity.

12. The device according to claim 8, wherein The first model is a model obtained through hyperparameter optimization processing, and the hyperparameters in the hyperparameter optimization processing are determined according to the target classification.

13. The device according to claim 8, wherein The first model and the second model refer to models of different layers included in the same hierarchical model.

14. The device according to any one of claims 8 to 13, characterized in that The network traffic includes the network traffic of the target base station cell, the physical resource block utilization includes the physical resource block utilization of the target base station cell, and the number of users includes the number of users of the target base station cell.

15. A data prediction device, characterized in that: The system comprises at least one processor and a memory, wherein the at least one processor is coupled to the memory and is configured to read and execute instructions in the memory to perform the method according to any one of claims 1 to 7.

16. A computer-readable medium, characterized in that The computer-readable medium stores a program code, and when the computer program code is run on a computer, the computer is caused to perform the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Prophet-CEEMDAN-ARIMA-based commodity sales volume prediction method and Prophet-CEEMDAN-ARIMA-based commodity sales volume prediction device

    CN111461786A