A base station traffic prediction method and related device
By selecting neighboring base stations based on geographical distance and causal relationships, and by utilizing attention mechanisms and multi-scale decomposition methods, the problems of single and blind feature selection in base station traffic prediction are solved, thereby improving prediction accuracy and efficiency.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- CHINA MOBILE COMM LTD RES INST
- Filing Date
- 2022-03-04
- Publication Date
- 2026-05-29
AI Technical Summary
Existing base station traffic prediction methods suffer from limited or blind selection of feature data, resulting in poor prediction accuracy.
By filtering neighboring base stations based on geographical distance and causal relationships between them, extracting feature data using an attention mechanism, and employing multiple predictors for multi-scale decomposition and weighted summation, base station traffic can be predicted.
It improves the accuracy and efficiency of base station traffic prediction, reduces interference from redundant features, and enhances the effectiveness of feature selection.
Smart Images

Figure CN116744325B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of communication technology, and in particular to a base station traffic prediction method and related equipment. Background Technology
[0002] In recent years, the widespread use of mobile devices has led to a surge in traffic to mobile base stations (BS). In this context, accurately predicting base station traffic in advance, and managing and planning base stations accordingly to avoid network congestion during peak periods, is crucial for stable network services.
[0003] However, existing base station traffic prediction methods suffer from a certain degree of singularity or blindness in extracting feature data, resulting in poor accuracy in base station traffic prediction. Summary of the Invention
[0004] This invention provides a base station traffic prediction method and related equipment to solve the problem of poor accuracy in existing base station traffic prediction.
[0005] To solve the above-mentioned technical problems, the present invention is implemented as follows:
[0006] In a first aspect, embodiments of the present invention provide a base station traffic prediction method, comprising:
[0007] Based on the geographical distance information between base stations, determine the neighboring base stations of the base station to be predicted;
[0008] Identify neighboring base stations among the nearby base stations, and find a causal relationship between the historical traffic records of the neighboring base stations and the historical traffic records of the base station to be predicted;
[0009] Feature data is obtained by extracting features from the historical traffic records of the base station to be predicted and the neighboring base stations;
[0010] Predict the predicted traffic of the base station to be predicted based on the feature data.
[0011] Optionally, determining the neighboring base stations among the nearby base stations includes:
[0012] Determine whether there is a causal relationship between the historical traffic records of the base station to be predicted and the historical traffic records of the i-th neighboring base station, where i = 1, ..., k, and k is the total number of neighboring base stations of the base station to be predicted;
[0013] If a causal relationship exists, the i-th neighboring base station is determined as a neighboring base station.
[0014] Optionally, the step of extracting features from the historical traffic records of the base station to be predicted and the neighboring base stations to obtain feature data includes:
[0015] Based on the attention mechanism, the attention score between the base station to be predicted and the neighboring base stations is calculated using the historical traffic records of the base station to be predicted and the neighboring base stations.
[0016] The spatial features are obtained by weighted summation of the historical traffic records of the neighboring base stations based on the attention scores.
[0017] Feature data is obtained based on the historical traffic records of the base station to be predicted and the spatial characteristics.
[0018] Optionally, the step of weighted summing of the historical traffic records of the neighboring base stations based on the attention score to obtain spatial features includes:
[0019] The first attention distribution and the second attention distribution of the attention score are calculated using the softmax function and the softmin function, respectively.
[0020] The positive spatial features of the neighboring base stations are obtained by weighted summation of the historical traffic records of the neighboring base stations using the first attention distribution, and the negative spatial features of the neighboring base stations are obtained by weighted summation of the historical traffic records of the neighboring base stations using the second attention distribution.
[0021] The positive and negative spatial features of the neighboring base stations are determined as spatial features.
[0022] Optionally, predicting the predicted traffic of the base station to be predicted based on feature data includes:
[0023] The feature data is input into N predictors respectively, and N first predicted flows are output, where N>2. The N predictors are divided into M+1 groups of predictors, M≥1. The first predicted flows obtained by different groups of predictors are at least partially different in scale.
[0024] The first predicted flow obtained by each predictor group is decomposed into a first predicted series set. Each first predicted series set includes P series terms, where p is the index of the series term, and 0≤p≤P-1.
[0025] The weighted sum of the series terms after decomposition is obtained by using a weight matrix; the sum of the weight parameters of the series terms p in each first prediction series set is 1.
[0026] Optionally, the indexes of the first predicted flow decompositions obtained from the different predictors in each group of predictors may differ.
[0027] Optionally, the predictor is a temporal convolutional network.
[0028] In a second aspect, embodiments of the present invention provide a base station traffic prediction device, comprising:
[0029] The first filtering module is used to determine the neighboring base stations of the base station to be predicted based on the geographical distance information between base stations.
[0030] The second filtering module is used to determine the neighboring base stations among the nearby base stations, and there is a causal relationship between the historical traffic records of the neighboring base stations and the historical traffic records of the base station to be predicted.
[0031] The feature extraction module is used to extract features from the historical traffic records of the base station to be predicted and the neighboring base stations to obtain feature data;
[0032] The prediction module is used to predict the predicted traffic of the base station to be predicted based on feature data.
[0033] Optional, the second filtering module includes:
[0034] The judgment module is used to determine whether there is a causal relationship between the historical traffic records of the base station to be predicted and the historical traffic records of the i-th neighboring base station, i = 1, ..., k, where k is the total number of neighboring base stations of the base station to be predicted;
[0035] The filtering submodule is used to determine the i-th neighboring base station as a neighboring base station if a causal relationship exists.
[0036] Optional feature extraction module, including:
[0037] The calculation module is used to calculate the attention score between the base station to be predicted and the neighboring base stations based on the attention mechanism and using the historical traffic records of the base station to be predicted and the neighboring base stations.
[0038] The spatial feature module is used to perform a weighted summation of the historical traffic records of the neighboring base stations based on the attention score to obtain spatial features;
[0039] The feature extraction submodule is used to obtain feature data based on the historical traffic records of the base station to be predicted and the spatial features.
[0040] Optional spatial feature modules include:
[0041] An attention distribution module is used to calculate the first attention distribution and the second attention distribution of the attention score using the softmax function and the softmin function, respectively.
[0042] The spatial feature submodule is used to obtain the positive spatial features of the neighboring base station by weighted summation of the historical traffic records of the neighboring base station using the first attention distribution, and to obtain the negative spatial features of the neighboring base station by weighted summation of the historical traffic records of the neighboring base station using the second attention distribution.
[0043] The positive and negative spatial features of the neighboring base stations are then identified as spatial features.
[0044] Optionally, the prediction module includes:
[0045] The predictor module is used to input the feature data into N predictors respectively and output N first predicted flows, where N>2, the N predictors are divided into M+1 groups of predictors, M≥1, and the first predicted flows obtained by different groups of predictors are at least partially different in scale.
[0046] The series decomposition module is used to decompose the first predicted flow obtained by each group of predictors into a first predicted series set. Each first predicted series set includes P series terms, where p is the index of the series term, and 0≤p≤P-1.
[0047] The aggregation module is used to perform a weighted summation of the decomposed series terms using a weight matrix to obtain the predicted traffic of the base station to be predicted; the sum of the weight parameters of the series terms p in each first prediction series set is 1.
[0048] Optionally, the indexes of the first predicted flow decompositions obtained from the different predictors in each group of predictors may differ.
[0049] Optionally, the predictor is a temporal convolutional network.
[0050] Thirdly, embodiments of the present invention provide an electronic device, including a processor.
[0051] The processor is used to determine the neighboring base stations of the base station to be predicted based on the geographical distance information between base stations.
[0052] Identify neighboring base stations among the nearby base stations, and find a causal relationship between the historical traffic records of the neighboring base stations and the historical traffic records of the base station to be predicted;
[0053] Feature data is obtained by extracting features from the historical traffic records of the base station to be predicted and the neighboring base stations;
[0054] Predict the predicted traffic of the base station to be predicted based on the feature data.
[0055] Optionally, determining the neighboring base stations among the nearby base stations includes:
[0056] Determine whether there is a causal relationship between the historical traffic records of the base station to be predicted and the historical traffic records of the i-th neighboring base station, where i = 1, ..., k, and k is the total number of neighboring base stations of the base station to be predicted;
[0057] If a causal relationship exists, the i-th neighboring base station is determined as a neighboring base station.
[0058] Optionally, the step of extracting features from the historical traffic records of the base station to be predicted and the neighboring base stations to obtain feature data includes:
[0059] Based on the attention mechanism, the attention score between the base station to be predicted and the neighboring base stations is calculated using the historical traffic records of the base station to be predicted and the neighboring base stations.
[0060] The spatial features are obtained by weighted summation of the historical traffic records of the neighboring base stations based on the attention scores.
[0061] Feature data is obtained based on the historical traffic records of the base station to be predicted and the spatial characteristics.
[0062] Optionally, the step of weighted summing of the historical traffic records of the neighboring base stations based on the attention score to obtain spatial features includes:
[0063] The first attention distribution and the second attention distribution of the attention score are calculated using the softmax function and the softmin function, respectively.
[0064] The positive spatial features of the neighboring base station are obtained by weighted summation of the historical traffic records of the neighboring base station using the first attention distribution, and the negative spatial features of the neighboring base station are obtained by weighted summation of the historical traffic records of the neighboring base station using the second attention distribution.
[0065] The positive and negative spatial features of the neighboring base stations are determined as spatial features.
[0066] Optionally, predicting the predicted traffic of the base station to be predicted based on feature data includes:
[0067] The feature data is input into N predictors respectively, and N first predicted flows are output, where N>2. The N predictors are divided into M+1 groups of predictors, M≥1. The first predicted flows obtained by different groups of predictors are at least partially different in scale.
[0068] The first predicted flow obtained by each predictor group is decomposed into a first predicted series set. Each first predicted series set includes P series terms, where p is the index of the series term, and 0≤p≤P-1.
[0069] The weighted sum of the series terms after decomposition is obtained by using a weight matrix; the sum of the weight parameters of the series terms p in each first prediction series set is 1.
[0070] Optionally, the indexes of the first predicted flow decompositions obtained from the different predictors in each group of predictors may differ.
[0071] Optionally, the predictor is a temporal convolutional network.
[0072] Fourthly, embodiments of the present invention provide an electronic device, including: a processor, a memory, and a program stored in the memory and executable on the processor, wherein when the program is executed by the processor, it implements the steps of the base station traffic prediction method as described in the first aspect above.
[0073] Fifthly, embodiments of the present invention provide a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps of the base station traffic prediction method as described in the first aspect above.
[0074] In this embodiment of the invention, neighboring base stations of the base station to be predicted are determined based on geographical distance information between base stations; neighboring base stations among the neighboring base stations are determined, and a causal relationship exists between the historical traffic records of the neighboring base stations and the historical traffic records of the base station to be predicted; feature extraction is performed on the historical traffic records of the base station to be predicted and the neighboring base stations to obtain feature data; and the predicted traffic of the base station to be predicted is predicted based on the feature data. By selecting neighboring base stations through geographical distance information and causal relationships, spatial factors are fully considered during the feature data selection process to avoid the uniformity of feature data selection. Simultaneously, the causal relationships between base station traffic are also fully considered, thereby avoiding blind selection of feature data, reducing redundant feature interference, and improving the effectiveness of feature selection, thus improving the accuracy and efficiency of base station traffic prediction. Attached Figure Description
[0075] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments of the present invention will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0076] Figure 1 This is a flowchart of a base station traffic prediction method provided in an embodiment of the present invention;
[0077] Figure 2 This is a flowchart of a neighbor base station screening method provided in an embodiment of the present invention;
[0078] Figure 3 This is a schematic diagram of the prediction task decomposition of a predictor provided in an embodiment of the present invention;
[0079] Figure 4 This is a flowchart of another base station traffic prediction method provided in an embodiment of the present invention;
[0080] Figure 5This is a schematic diagram of a base station traffic prediction device provided in an embodiment of the present invention;
[0081] Figure 6 This is a schematic diagram of an electronic device provided in an embodiment of the present invention. Detailed Implementation
[0082] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0083] Traditional traffic forecasting methods, such as support vector machine regression and differential mobile autoregression, are widely used due to their efficiency and ease of use. With the rise of deep learning, a series of neural network models, such as recurrent neural network models, have also been applied to traffic forecasting. However, the above methods model the traffic forecasting problem as a simple sequence forecasting problem, considering only time characteristics. But traffic forecasting is different from general time series forecasting problems. Due to the influence of factors such as the mobility of mobile device users, the traffic characteristics between base stations are not independent and have strong spatial distribution characteristics. If these models that only consider time characteristics are used, it is difficult to obtain the expected results. Therefore, it is necessary to consider the influence of spatial characteristics.
[0084] In improved traffic prediction methods, spatial feature modeling is introduced. One approach is to implicitly model spatial features, using information from neighboring base stations as input to the model. This means inputting features from neighboring base stations along with those of the base station being predicted into the time series prediction model. This method doesn't add an extra spatial feature extraction unit to the model. However, this approach can easily lead to confusion between primary and secondary features, and excessive features may result in feature redundancy, reducing prediction accuracy. While this type of method can also extract the impact of surrounding base stations on the current base station's traffic, it doesn't explicitly model spatial features into the model. The basic model of this type is essentially the same as a model that only considers time features, except that it includes some historical traffic features from surrounding base stations as input.
[0085] Another approach involves explicitly modeling spatial features, using a dedicated spatial feature extraction unit to process them. Common practices include first selecting "neighbor base stations" based on clustering algorithms, then aggregating features based on these neighbor base stations, or using graph neural networks to construct an adjacency matrix based on the features of surrounding base stations, aggregating features using graph convolution, and finally integrating these spatial features into temporal features, ultimately achieving prediction through time series processing algorithms. However, these algorithms exhibit some degree of randomness in extracting spatial features, resulting in poor interpretability and the introduction of interference information that affects prediction accuracy. For example, clustering-based spatial feature extraction methods, being unsupervised learning methods, often produce clustering results that vary significantly across different metrics, leading to poor interpretability and difficulty in evaluating the results.
[0086] It is evident that existing base station traffic prediction methods typically only consider traffic characteristics or only consider traffic characteristics and geographical information. When extracting spatial features, they exhibit a certain degree of singularity or blindness, resulting in poor accuracy in base station traffic prediction.
[0087] In this embodiment of the invention, a base station traffic prediction method is proposed to solve the problem of poor accuracy in existing base station traffic prediction.
[0088] See Figure 1 , Figure 1 This is a flowchart of a base station traffic prediction method provided in an embodiment of the present invention, such as... Figure 1 As shown, the method includes the following steps:
[0089] Step 101: Based on the geographical distance information between base stations, determine the neighboring base stations of the base station to be predicted.
[0090] In a communication system, each base station has a relatively fixed geographical location, and the geographical distance between base stations is known, for example, it can be calculated using the latitude and longitude information of the base stations. Taking a particular base station as the base station to be predicted, the neighboring base stations of the base station to be predicted can be determined based on the spatial distance information between the base station to be predicted and other base stations.
[0091] Optionally, the nearby base stations can be base stations within a preset distance range from the base station to be predicted, or they can be the k closest base stations to the base station to be tested. Optionally, the k closest base stations to the base station to be tested can be selected using the K-nearest neighbor algorithm or the weighted K-nearest neighbor algorithm.
[0092] In this embodiment of the invention, step 101 above uses the geographical distance information between base stations to perform the first round of screening of neighboring base stations. The neighboring base stations can be understood as some of the target base stations corresponding to the input features of the traffic prediction model. It can be understood that the target base stations include not only the neighboring base stations, but also the base station to be predicted itself. That is, the feature data corresponding to the historical traffic records of the base station to be predicted and the historical traffic records of the neighboring base stations are used as the input features of the traffic prediction model.
[0093] Step 102: Determine the neighboring base stations among the nearby base stations, and find a causal relationship between the historical traffic records of the neighboring base stations and the historical traffic records of the base station to be predicted.
[0094] In this embodiment of the invention, step 102 above performs a second round of screening of neighboring base stations based on the causal relationship between the historical traffic records of the base station to be predicted and the neighboring base stations. That is, it determines whether a neighboring base station is a neighboring base station based on whether there is a causal relationship between the historical traffic records of the neighboring base station and the historical traffic records of the base station to be predicted. This can be understood as determining whether the historical traffic records of the neighboring base station have influenced the historical traffic records of the base station to be predicted; if they have, the neighboring base station is identified as a neighboring base station.
[0095] Optionally, step 102 includes:
[0096] Determine whether there is a causal relationship between the historical traffic records of the base station to be predicted and the historical traffic records of the i-th neighboring base station, where i = 1, ..., k, and k is the total number of neighboring base stations of the base station to be predicted;
[0097] If a causal relationship exists, the i-th neighboring base station is determined as a neighboring base station.
[0098] Optionally, causal testing methods can be used to determine whether a causal relationship exists between the historical traffic records of the base station to be predicted and the historical traffic records of neighboring base stations, thereby identifying the neighboring base stations among the neighboring base stations. Specific causal detection methods may include Granger causality tests, dynamic causality models, transfer entropy models, and other causal detection methods.
[0099] Taking the Granger causality test as an example, in step 102, historical traffic records from two base stations are used to filter neighboring base stations through the Granger causality test. For example, for two time series Y and Z, where Y represents the traffic data of the base station to be predicted and Z represents the traffic data of a neighboring base station, to determine whether Z affects Y, given the lag parameter h, the test equation can be obtained:
[0100]
[0101] Where, η i ,γ i It is the regression coefficient, y t ,z t These are the actual flow values in time series Y and Z, respectively, and corresponding to y. t-i and z t-i Let Y and Z be the actual flow rates of time series Y and Z, respectively, offset from time t by i points (or, in other words, the actual flow rates of the i points preceding the predicted time t in time series Y and Z). u is the estimated flow rate value in the time series Y. t It can be considered as white noise, where t is the time point.
[0102] The null hypothesis that Z has no effect on Y:
[0103] H0: γ1=γ2=...,=γ h =0
[0104] Then, the p-value can be calculated using the F-statistic to perform hypothesis testing and determine whether Z affects Y. In other words, it determines whether a neighboring base station with traffic Z will affect the current base station to be predicted, thus selecting neighboring base stations. The calculation method for the F-statistic is shown in the following formula:
[0105]
[0106] SSE r SSE represents the sum of squared residuals of the model after applying constraints (assuming the null hypothesis is true). u This represents the sum of squared residuals of the model under no constraints, and T represents the sample size, i.e., the length of the historical flow.
[0107] The sum of squared residuals is calculated as follows:
[0108]
[0109] For SSE u , The calculation method is as follows:
[0110]
[0111] For SSE r , The calculation method is as follows:
[0112]
[0113] F can be obtained by looking up the table. α(h, T-2h), where α is the significance level and the p-value is the minimum significance level for rejecting the null hypothesis. Here, the p-value is used to determine whether to accept the null hypothesis. If F ≥ F α If (h, T-2h), then the null hypothesis is rejected. That is, the null hypothesis is that Z has no effect on Y. Calculations determine that F ≥ F0. α (h,T-2h), reject the null hypothesis, that is, believe that Z will affect Y, that is, determine the nearby base station corresponding to Z as the neighbor base station.
[0114] In this embodiment of the invention, steps 101 and 102 above, through two rounds of screening of neighboring base stations, are based on spatial distance and causal relationship respectively, which can improve the effectiveness of subsequent spatial feature extraction, reduce redundant feature interference, and solve the problem of low accuracy of existing methods in modeling spatial features. Figure 2 Taking steps 101 and 102, which respectively employ the K-nearest neighbor algorithm and the Granger causality test, as examples, a flowchart of the neighbor base station selection method is shown.
[0115] Step 103: Extract features from the historical traffic records of the base station to be predicted and the neighboring base stations to obtain feature data.
[0116] The neighboring base stations are screened through the above steps 101 and 102, and the neighboring base stations of the base station to be predicted are determined from the communication system. Then, feature data is obtained by feature extraction based on the historical traffic records of the base station to be predicted and the neighboring base stations.
[0117] Step 104: Predict the predicted traffic of the base station to be predicted based on the feature data.
[0118] After determining the characteristic data, the traffic of the base station to be predicted can be predicted using a traffic prediction model.
[0119] In this embodiment of the invention, neighboring base stations of the base station to be predicted are determined based on geographical distance information between base stations; neighboring base stations among the neighboring base stations are identified, and a causal relationship exists between the historical traffic records of the neighboring base stations and the historical traffic records of the base station to be predicted; feature extraction is performed on the historical traffic records of the base station to be predicted and the neighboring base stations to obtain feature data; and the predicted traffic of the base station to be predicted is predicted based on the feature data. By selecting neighboring base stations through geographical distance information and causal relationships, spatial factors are fully considered during the feature data selection process to avoid the uniformity of feature data selection. Simultaneously, the causal relationships between base station traffic are also fully considered, thereby avoiding blind selection of feature data, reducing redundant feature interference, and improving the effectiveness of feature selection, thus improving the accuracy and efficiency of base station traffic prediction.
[0120] This invention also provides another base station traffic prediction method, the method comprising the following steps:
[0121] Step 201: Based on the geographical distance information between base stations, determine the neighboring base stations of the base station to be predicted.
[0122] Step 201 above performs a first round of filtering of neighboring base stations based on the geographical distance information between base stations. The optional implementation methods described above can be found in [reference needed]. Figure 1 To avoid repetition, the relevant descriptions in the embodiments shown will not be repeated in this embodiment.
[0123] Step 202: Based on the geographical distance information between base stations, determine the neighboring base stations of the base station to be predicted.
[0124] Step 202 above performs a second round of screening of neighboring base stations based on the causal relationship between the historical traffic records of the base station to be predicted and the neighboring base stations. Optional implementation methods described above can be found in [reference needed]. Figure 1 To avoid repetition, the relevant descriptions in the embodiments shown will not be repeated in this embodiment.
[0125] Step 203: Extract features from the historical traffic records of the base station to be predicted and the neighboring base stations to obtain feature data.
[0126] The neighboring base stations are screened through the above steps 201 and 202 to determine the neighboring base stations of the base station to be predicted from the communication system. Then, feature data is obtained by feature extraction based on the historical traffic records of the base station to be predicted and the neighboring base stations.
[0127] Optionally, step 203 specifically employs an attention mechanism for feature extraction. This can be understood as follows: after determining in step 202 that neighboring base stations have an impact on the traffic of the base station to be predicted, the attention mechanism is further used to quantitatively calculate the impact of neighboring base stations on the base station to be predicted, thereby reflecting the different degrees of impact of neighboring base stations on the base station to be predicted in the feature extraction results.
[0128] Optionally, the step of extracting features from the historical traffic records of the base station to be predicted and the neighboring base stations to obtain feature data includes:
[0129] Based on the attention mechanism, the attention score between the base station to be predicted and the neighboring base stations is calculated using the historical traffic records of the base station to be predicted and the neighboring base stations.
[0130] The spatial features are obtained by weighted summation of the historical traffic records of the neighboring base stations based on the attention scores.
[0131] Feature data is obtained based on the historical traffic records of the base station to be predicted and the spatial characteristics.
[0132] Optionally, an attention score is calculated between the historical traffic records of the current base station to be predicted and the historical traffic records of each selected neighbor base station, as shown below, x i y represents the historical traffic of neighboring base stations and the historical traffic of the base station to be predicted, respectively.
[0133] Score(x i ,y)=x i T ·y
[0134] Where i represents the index of the neighboring base station, and T represents the matrix transpose.
[0135] Attention score for neighboring base station i (Score(x)) i After normalization, the attention score can be used as the weight of the historical traffic records of neighboring base stations to calculate spatial features. The historical traffic records of neighboring base stations and the historical traffic records of the base station to be predicted can be used as traffic data of different dimensions. Spatial features are determined based on the historical traffic records of all neighboring base stations, and these spatial features serve as spatial feature data. Temporal feature data is determined based on the historical traffic records of the base station to be predicted, resulting in temporal and spatial feature data for subsequent traffic prediction.
[0136] Optionally, the step of weighted summing of the historical traffic records of the neighboring base stations based on the attention score to obtain spatial features includes:
[0137] The first attention distribution and the second attention distribution of the attention score are calculated using the softmax function and the softmin function, respectively.
[0138] The positive spatial features of the neighboring base station are obtained by weighted summation of the historical traffic records of the neighboring base station using the first attention distribution, and the negative spatial features of the neighboring base station are obtained by weighted summation of the historical traffic records of the neighboring base station using the second attention distribution.
[0139] The positive and negative spatial features of the neighboring base stations are determined as spatial features.
[0140] In the above embodiments, the softmax and softmin functions are further utilized, that is, the softmax and softmin functions Score(x) are used. iThe process involves determining two weighted distributions—a first attention distribution and a second attention distribution—based on the historical traffic records of neighboring base stations. These two attention distributions are then used to weight and sum the historical traffic records of neighboring base stations, yielding two feature results corresponding to positive and negative spatial features. Specifically, positive spatial features are determined based on the historical traffic records of all neighboring base stations, forming one dimension of feature data; negative spatial features are also determined based on the historical traffic records of all neighboring base stations, forming another dimension of feature data; and finally, feature data for the base station to be predicted is determined based on the historical traffic records of the base station to be predicted, resulting in three-dimensional feature data for subsequent traffic prediction. This further improves the accuracy of traffic prediction.
[0141] The calculation formulas for the Softmax and softmin functions are as follows:
[0142]
[0143]
[0144] Where i represents the i-th element among j elements, s i Represents the value of the i-th element, ∑ j ∑ represents the summation of terms j. j exp(s j )=exp(s1)+…+exp(s i )+…+exp(s j ), ∑ j exp(-s j )=exp(-s1)+…+exp(-s i )+…+exp(-s j ).
[0145] Specifically, using the softmax and softmin functions, Score(x) i ,y) Determine the positive and negative weight distributions of the historical traffic records of neighboring base stations, s i and s j Let si represent the attention scores of the current base station to the i-th and j-th neighbor base stations, respectively, i.e., si = Score(x) i ,y)=x i T ·y, which is also about to be Score(x) i ,y) as s i The values are substituted into the softmax and softmin functions to determine the positive and negative weight distributions of the historical traffic records of neighboring base stations.
[0146] Step 204: Input the feature data into N predictors respectively, and output N first predicted traffic, where N>2. The N predictors are divided into M+1 groups of predictors, M≥1. The first predicted traffic obtained by different groups of predictors has at least some different scales. The first predicted traffic obtained by each group of predictors is decomposed into a first prediction series set through series decomposition. Each first prediction series set includes P series terms, where p is the index of the series term, 0≤p≤P-1. The series terms after decomposition are weighted and summed using a weight matrix to obtain the predicted traffic of the base station to be predicted. The sum of the weight parameters of the series terms p in each first prediction series set is 1.
[0147] Base station traffic data exhibits surges and sharp decreases at different times, resulting in a widespread long-tail distribution. This characteristic of traffic data significantly interferes with conventional traffic prediction models: on the one hand, the large amount of low-volume traffic data makes models more inclined to predict smaller outcomes, failing to effectively predict large-volume traffic; on the other hand, the highly uneven temporal distribution of the data makes it difficult for models to predict sudden changes in traffic data. Few methods specifically design time-series processing units to address the widespread long-tail distribution problem in traffic data. Generally, simple statistical transformations, such as logarithmic transformations, are used. However, these methods merely amplify the overall trend of data change and cannot effectively extract the characteristics of traffic data at different numerical scales, offering very limited help in solving the long-tail distribution problem.
[0148] In the above embodiments of the present invention, several different predictors are specifically used to predict traffic data at different levels of scale. The predicted traffic is determined by weighting and combining the prediction results at different scales. Since the typical characteristic of long-tail distribution is that the data fluctuates greatly and a large number of small data are concentrated at the tail end of the data distribution, the above embodiments of the present invention perform scale decomposition on the original traffic data, which can capture more fine-grained traffic changes, weaken the differences between data, and thus reduce the interference of the long-tail distribution phenomenon of traffic data on traffic data prediction.
[0149] For ease of explanation, we assume the flow data is one-dimensional (it can actually be multi-dimensional, as explained in step 204; the following analysis method also applies to high-dimensional input data). Optionally, historical flow records can be normalized. For example, flow data x i It can be expressed as the sum of an infinite series:
[0150]
[0151] Where, 0≤α ij ≤k-1, α ij In this series, k is a natural number, and j represents the index of a term. For example, if k = 10, xi =0.0123, then it can be represented as Clearly, as j increases, more granular variations in flow can be achieved. To better define the scale of the flow data, we merge adjacent d terms of the series into one term; that is, terms from j×d to j×d+d-1 will be merged into one term. The merged result is shown in the following equation:
[0152]
[0153] Therefore, we can further refine the traffic data by using series terms, use different models to predict specific items, and thus obtain more precise prediction results.
[0154] Optionally, the indexes of the first predicted flow decompositions obtained from the different predictors in each group of predictors may differ.
[0155] Suppose that x i It can be decomposed into four series terms β0, β1, β2, and β3, corresponding to series indices 0-3 respectively. The series terms 0, 1, 2, and 3 can be predicted using the first set of four predictors {f0, f1, f2, f3}, respectively, to obtain the predicted series terms. Then, by using f4 from the second set of predictors {f4, f5} to predict the series terms 0 and 1, and performing series decomposition on the prediction result of f4, we can obtain... By predicting the series terms 2 and 3 using f5, and performing series decomposition on the prediction results of f5, we can obtain... right and By performing a weighted summation, we obtain the prediction result for series term 0. Similarly, we obtain the prediction results for series terms 1, 2, and 3 in sequence, thus determining x. i The prediction results.
[0156] Here, {f0,f1,f2,f3} can be understood as a group predictor, and the four first predicted flows obtained correspond to the series terms β0, β1, β2, and β3, respectively, with different series term indices among them. {f4,f5} can also be understood as a group predictor, and the two first predicted flows obtained are decomposed into {β0',β1'} and {β2',β3'}, respectively, with different series term indices among {β0',β1'} and {β2',β3'}.
[0157] Optionally, the scale decomposition task can be refined based on redundant coding, that is, the scale can be divided using redundant coding, or the scales divided by redundant coding can be further grouped.
[0158] Define N predictors as f0,...,f N-1The number N of weak predictors satisfies the following formula:
[0159]
[0160] To distinguish between the index of the series terms and the index of the predictor model, we define the index of the model as q and the index of the series terms as p, where 0 ≤ q ≤ N-1, p ≥ 0. Therefore, the prediction task of each model can be expressed in the form of the following equation:
[0161]
[0162] Where u and v are the start and end indices of the series terms, respectively. For a prediction task requiring N classifiers, these models can be divided into M+1 groups, namely G0, G1, ..., G... M The correspondence between the predictor model index and the group is shown below (the elements in the set are the indices of the predictor models):
[0163]
[0164] For each model group, it can represent complete traffic information. Therefore, for the first model group, the prediction task of model f0 is the complete traffic information, and the index of the corresponding series term is from 0 to P. For the (M+1)th model group, we hope it can capture the finest-grained traffic changes, so... Will predict separately Predict all remaining terms of the series, i.e., the corresponding series term index is 2. M -1 to P. Similarly, for the Mth model, Need to predict To predict all the remaining terms of the series, the corresponding series term index is 2. M -1 to P. Following this rule, the prediction accuracy of the (M-1)th group is half that of the Mth group, the prediction accuracy of the (M-2)th group is half that of the M-1th group, and so on. This allows us to obtain the specific prediction task for each predictor model; where P is a positive integer, or in other words, P∈N. + ,0<P<∞. Figure 3 The diagram shows the prediction task breakdown of the predictors, where there are 3 predictors.
[0165] Based on the above analysis, u and v are determined by the model index and the group number in which the model belongs. The group index is defined as m, where... The specific expressions for u and v are as follows:
[0166] u=(q-2 m +1)×2 M-m
[0167]
[0168] in, This indicates rounding down to the nearest integer.
[0169] After completing each prediction task, we will aggregate these results to obtain the final result. For each prediction result... In other words, it can be broken down into the following forms:
[0170]
[0171] The weighted sum of the decomposed results is shown in the following formula:
[0172]
[0173] Where the weighting parameter ξ pq It is learnable, after weighted summation. That is the final prediction result.
[0174] Optional, weighting parameter ξ pq It is obtained through training.
[0175] For example, assuming M = 2, then N = 1 + 2 + 4 = 7. The 7 models are divided into three groups: {f0}, {f1, f2}, and {f3, f4, f5}. Assume x i The prediction can be decomposed into five series terms: β0, β1, β2, β3, and β4. Then, f3 and f4 are used to predict β0 and β1 respectively, f5 is used to predict β2+β3+β4, f1 is used to predict β0+β1, f2 is used to predict β2+β3+β4, and f0 is used to predict β0+β1+β2+β3+β4. Performing series decomposition on the prediction result of f0 yields the first set of prediction series (β0, β1, β2, β3, β4). Performing series decomposition on the prediction results of f1 and f2 yields the second set of prediction series (β0, β1, β2, β3, β4). Performing series decomposition on the prediction result of f5, combined with the prediction results of f3 and f4, yields the third set of prediction series (β0, β1, β2, β3, β4). Finally, a weighted sum is taken of the three prediction results of β0 in the three sets of prediction series to determine the predicted... The sum of the weight parameters of the three prediction results for β0 is 1. Similarly, the following calculations are performed sequentially. and The prediction result is then used to obtain x. i The prediction results.
[0176] By decomposing the task based on redundant coding, the original data is scaled to extract more granular traffic changes, enabling multi-scale traffic data analysis. This solves the problem that existing methods cannot effectively overcome the long-tail distribution of data and improves prediction performance.
[0177] Optionally, the predictor is a temporal convolutional network (TCN).
[0178] Existing traffic prediction models typically fail to adequately model long-term traffic dependencies. While traffic characteristics fluctuate significantly in the short term, these fluctuations tend to exhibit certain patterns in the long run. Therefore, the long-term dependencies of base station traffic have a crucial impact on prediction results, and long-term dependency information can enhance model prediction performance. However, most sequence prediction algorithms have poor "memory" and struggle to model such long-term dependencies. For example, commonly used models based on recurrent neural networks and sparse fuzzy cognitive graphs all have limitations in this regard.
[0179] In the above embodiments of the present invention, by setting multiple temporal convolutional networks as predictors, that is, as time series processing units, and relying on the advantages of the temporal convolutional network structure itself in capturing long-term dependency information, long-term dependencies in time series can be extracted more effectively. Based on the long-term memory advantage of its structure, long-term dependency modeling can be achieved, solving the problem that the original method cannot effectively model long-term dependencies.
[0180] The final traffic prediction result is obtained by weighted summation of the prediction results from N predictors. These predictors can be understood as several weak predictors. To obtain fine-grained traffic information, we divide traffic prediction into multiple prediction tasks based on redundant coding. Each task uses a temporal convolutional network model as a weak predictor. These weak predictors together form the time series processing module. Using a weakly supervised model as the time series process unit to complete traffic predictions of different scales can capture more fine-grained traffic changes, weaken the differences between data, and thus reduce the interference of long-tail distribution of traffic data on traffic prediction.
[0181] Optional, Figure 4 A flowchart of another base station traffic prediction method is shown.
[0182] In this embodiment of the invention, in Figure 1The embodiment described above, based on the screening of neighboring base stations using geographical distance information and causal relationships, further extracts feature data using an attention mechanism and employs predictions from multiple predictors at different levels. This further improves the effectiveness of the feature data, reduces the interference of long-tail distribution on the prediction model, and thus improves the accuracy of base station traffic prediction.
[0183] To more clearly illustrate the embodiments of the present invention, specific implementation methods of the present invention will be described below, especially the training method of the predictor. It should be understood that the following numerical and algorithm examples are for illustrative purposes only, and the embodiments of the present invention are not limited to the following numerical values or numerical ranges and algorithms.
[0184] Suppose a user uses historical traffic data from a base station over 30 consecutive time points to predict traffic patterns for the next 10 consecutive time points:
[0185] (1) Dataset Construction. Based on the base station ID, the latitude and longitude of the base station to be predicted, historical traffic data, and the spatial distribution and historical traffic information of its surrounding base stations can be obtained. Traffic data from the base station to be predicted for the most recent 90 days are selected, with the first 58 days forming the training set and the remaining 32 days forming the validation and test sets at a 1:1 ratio. Using a sliding window approach, the training, validation, and test sets are divided into multiple data pairs from the initial data, using a format of 30 historical data points and 10 true values.
[0186] (2) Neighbor base station selection. Based on the spatial distribution of surrounding base stations, select the 50 base stations closest to the base station to be predicted. Then, perform Granger causality tests on these 50 base stations and the base station to be predicted. The specific implementation details are as follows: starting from the first nearest base station, input all the traffic data in the training set corresponding to the training set of the base station to be predicted into the Granger causality test model and calculate the p-value using the F statistic. If the p-value is less than 0.05, it is considered that there is a causal relationship between the traffic of the two base stations, and the change in the traffic data of the base station will affect the traffic data of the base station to be predicted. Otherwise, it is considered that there is no causal relationship between the traffic of the two base stations. Repeat this process until the 50th base station. After completing all tests, the base stations that are considered to have a causal relationship in the test results are taken as the "neighbor base stations" of the base station to be predicted.
[0187] (3) Spatiotemporal feature extraction. After obtaining the "neighboring base stations" of the base station to be predicted, the attention mechanism is used to aggregate the historical traffic data of the base station to be predicted data pair and the traffic data of all "neighboring base stations" at the same time. In this way, the input data of the time series processing module contains both temporal and spatial features.
[0188] (4) Task Decomposition and Model Training. After inputting into the time series processing module, the historical traffic data in the data pairs remains unchanged, and the ground truth data is decomposed according to the redundancy encoding method to obtain the ground truth required by each predictor. Here, the number of weak predictors N is set to 7, and the number of aggregated adjacency terms d is set to 2. The TCN network parameters of different weak predictors are independent of each other, and the network loss is subjected to gradient descent. The loss function is MSELoss, the initial learning rate is set to 0.0001, and dropout and early-stopping are added to prevent model overfitting. The model that minimizes the loss on the validation set is saved.
[0189] (5) Result Aggregation. The weight matrix is randomly initialized and mapped using the softmax function to ensure that the weight values are always between 0 and 1, and that the sum of the weight parameters of the same level item index is 1. The validation set and the training set are combined into a new training set. The new training set is input into the 7 saved weak predictor models respectively. After obtaining the prediction results, they are decomposed according to the corresponding level items. The weighted sum of the decomposed level items is obtained using the mapped weight matrix to obtain the prediction results. The loss between the prediction results and the original traffic true value is calculated. The loss function continues to use MSELoss to optimize and obtain the weight parameters that minimize the loss value. The training steps are the same as the training steps of the weak predictor.
[0190] In this way, the entire model is trained, and predictions can be made by inputting historical traffic values from the test set.
[0191] See Figure 5 , Figure 5 This is a schematic diagram of the structure of a base station traffic prediction device provided in an embodiment of the present invention, as shown below. Figure 5 As shown, the base station traffic prediction device 500 includes:
[0192] The first filtering module 501 is used to determine the neighboring base stations of the base station to be predicted based on the geographical distance information between base stations.
[0193] The second filtering module 502 is used to determine the neighboring base stations among the nearby base stations, and there is a causal relationship between the historical traffic records of the neighboring base stations and the historical traffic records of the base station to be predicted.
[0194] Feature extraction module 503 is used to extract features from the historical traffic records of the base station to be predicted and the neighboring base stations to obtain feature data;
[0195] The prediction module 504 is used to predict the predicted traffic of the base station to be predicted based on feature data.
[0196] Optional, the second filtering module includes:
[0197] The judgment module is used to determine whether there is a causal relationship between the historical traffic records of the base station to be predicted and the historical traffic records of the i-th neighboring base station, i = 1, ..., k, where k is the total number of neighboring base stations of the base station to be predicted;
[0198] The filtering submodule is used to determine the i-th neighboring base station as a neighboring base station if a causal relationship exists.
[0199] Optional, the feature extraction module includes:
[0200] The calculation module is used to calculate the attention score between the base station to be predicted and the neighboring base stations based on the attention mechanism and using the historical traffic records of the base station to be predicted and the neighboring base stations.
[0201] The spatial feature module is used to perform a weighted summation of the historical traffic records of the neighboring base stations based on the attention score to obtain spatial features;
[0202] The feature extraction submodule is used to obtain feature data based on the historical traffic records of the base station to be predicted and the spatial features.
[0203] Optional spatial feature modules include:
[0204] An attention distribution module is used to calculate the first attention distribution and the second attention distribution of the attention score using the softmax function and the softmin function, respectively.
[0205] The spatial feature submodule is used to obtain the positive spatial features of the neighboring base station by weighted summation of the historical traffic records of the neighboring base station using the first attention distribution, and to obtain the negative spatial features of the neighboring base station by weighted summation of the historical traffic records of the neighboring base station using the second attention distribution.
[0206] The positive and negative spatial features of the neighboring base stations are then identified as spatial features.
[0207] Optionally, the prediction module includes:
[0208] The predictor module is used to input the feature data into N predictors respectively and output N first predicted flows, where N>2, the N predictors are divided into M+1 groups of predictors, M≥1, and the first predicted flows obtained by different groups of predictors are at least partially different in scale.
[0209] The series decomposition module is used to decompose the first predicted flow obtained by each group of predictors into a first predicted series set. Each first predicted series set includes P series terms, where p is the index of the series term, and 0≤p≤P-1.
[0210] The aggregation module is used to perform a weighted summation of the decomposed series terms using a weight matrix to obtain the predicted traffic of the base station to be predicted; the sum of the weight parameters of the series terms p in each first prediction series set is 1.
[0211] Optionally, the indexes of the first predicted flow decompositions obtained from the different predictors in each group of predictors may differ.
[0212] Optionally, the predictor is a temporal convolutional network.
[0213] It should be noted that this embodiment is an implementation of the device corresponding to the embodiment of the method. Its specific implementation can be found in the relevant descriptions in the embodiment of the method. To avoid repetition, this embodiment will not repeat the description.
[0214] It should be noted that the device provided in the embodiments of the present invention is a device capable of performing the above methods. Therefore, all implementation methods in the above-described base station traffic prediction method embodiments are applicable to the base station traffic prediction device and can achieve the same or similar beneficial effects.
[0215] Specifically, embodiments of the present invention also provide an electronic device, including a processor 605, optionally, see [link to previous document]. Figure 6 As shown, the electronic device also includes a bus 601, a transceiver 602, an antenna 603, a bus interface 604, and a memory 606.
[0216] The aforementioned processor 605 is used to determine the neighboring base stations of the base station to be predicted based on the geographical distance information between base stations;
[0217] Identify neighboring base stations among the nearby base stations, and find a causal relationship between the historical traffic records of the neighboring base stations and the historical traffic records of the base station to be predicted;
[0218] Feature data is obtained by extracting features from the historical traffic records of the base station to be predicted and the neighboring base stations;
[0219] Predict the predicted traffic of the base station to be predicted based on the feature data.
[0220] Optionally, determining the neighboring base stations among the nearby base stations includes:
[0221] Determine whether there is a causal relationship between the historical traffic records of the base station to be predicted and the historical traffic records of the i-th neighboring base station, where i = 1, ..., k, and k is the total number of neighboring base stations of the base station to be predicted;
[0222] If a causal relationship exists, the i-th neighboring base station is determined as a neighboring base station.
[0223] Optionally, the step of extracting features from the historical traffic records of the base station to be predicted and the neighboring base stations to obtain feature data includes:
[0224] Based on the attention mechanism, the attention score between the base station to be predicted and the neighboring base stations is calculated using the historical traffic records of the base station to be predicted and the neighboring base stations.
[0225] The spatial features are obtained by weighted summation of the historical traffic records of the neighboring base stations based on the attention scores.
[0226] Feature data is obtained based on the historical traffic records of the base station to be predicted and the spatial characteristics.
[0227] Optionally, the step of weighted summing of the historical traffic records of the neighboring base stations based on the attention score to obtain spatial features includes:
[0228] The first attention distribution and the second attention distribution of the attention score are calculated using the softmax function and the softmin function, respectively.
[0229] The positive spatial features of the neighboring base station are obtained by weighted summation of the historical traffic records of the neighboring base station using the first attention distribution, and the negative spatial features of the neighboring base station are obtained by weighted summation of the historical traffic records of the neighboring base station using the second attention distribution.
[0230] The positive and negative spatial features of the neighboring base stations are determined as spatial features.
[0231] Optionally, predicting the predicted traffic of the base station to be predicted based on feature data includes:
[0232] The feature data is input into N predictors respectively, and N first predicted flows are output, where N>2. The N predictors are divided into M+1 groups of predictors, M≥1. The first predicted flows obtained by different groups of predictors are at least partially different in scale.
[0233] The first predicted flow obtained by each predictor group is decomposed into a first predicted series set. Each first predicted series set includes P series terms, where p is the index of the series term, and 0≤p≤P-1.
[0234] The weighted sum of the series terms after decomposition is obtained by using a weight matrix; the sum of the weight parameters of the series terms p in each first prediction series set is 1.
[0235] Optionally, the indexes of the first predicted flow decompositions obtained from the different predictors in each group of predictors may differ.
[0236] Optionally, the predictor is a temporal convolutional network.
[0237] exist Figure 6 In this document, a bus architecture (represented by bus 601) is used. Bus 601 can include any number of interconnected buses and bridges, linking various circuits including one or more processors represented by processor 605 and memory represented by memory 606. Bus 601 can also link various other circuits such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art and therefore will not be described further herein. Bus interface 604 provides an interface between bus 601 and transceiver 602. Transceiver 602 can be a single element or multiple elements, such as multiple receivers and transmitters, providing a unit for communicating with various other devices over a transmission medium. Data processed by processor 605 is transmitted over a wireless medium via antenna 603, which further receives data and transmits data to processor 605.
[0238] Processor 605 manages bus 601 and general processing, and also provides various functions, including timing, peripheral interface, voltage regulation, power management, and other control functions. Memory 606 can be used to store data used by processor 605 during operation.
[0239] Optionally, the processor 605 can be a CPU, ASIC, FPGA, or CPLD.
[0240] This invention also provides an electronic device, including: a processor, a memory, and a program stored in the memory and executable on the processor. When the program is executed by the processor, it implements the various processes of the above method embodiments and achieves the same technical effect. To avoid repetition, it will not be described again here.
[0241] This invention also provides a computer-readable storage medium storing a computer program. When executed by a processor, the computer program implements the various processes of the above-described method embodiments and achieves the same technical effects. To avoid repetition, further details are omitted here. The computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0242] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0243] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of the present invention.
[0244] The embodiments of the present invention have been described above with reference to the accompanying drawings. However, the present invention is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of the present invention without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of the present invention.
Claims
1. A base station traffic prediction method, characterized in that, include: Based on the geographical distance information between base stations, determine the neighboring base stations of the base station to be predicted; Identify neighboring base stations among the nearby base stations, and find a causal relationship between the historical traffic records of the neighboring base stations and the historical traffic records of the base station to be predicted; Feature data is obtained by extracting features from the historical traffic records of the base station to be predicted and the neighboring base stations; Predict the predicted traffic of the base station to be predicted based on the feature data.
2. The method according to claim 1, characterized in that, The step of determining the neighboring base stations among the nearby base stations includes: Determine whether there is a causal relationship between the historical traffic records of the base station to be predicted and the historical traffic records of the i-th neighboring base station, where i = 1, ..., k, and k is the total number of neighboring base stations of the base station to be predicted; If a causal relationship exists, the i-th neighboring base station is determined as a neighboring base station.
3. The method according to claim 1, characterized in that, The step of extracting features from the historical traffic records of the base station to be predicted and the neighboring base stations to obtain feature data includes: Based on the attention mechanism, the attention score between the base station to be predicted and the neighboring base stations is calculated using the historical traffic records of the base station to be predicted and the neighboring base stations. The spatial features are obtained by weighted summation of the historical traffic records of the neighboring base stations based on the attention scores. Feature data is obtained based on the historical traffic records of the base station to be predicted and the spatial characteristics.
4. The method according to claim 3, characterized in that, The step of weighted summing of the historical traffic records of the neighboring base stations based on the attention score to obtain spatial features includes: The first attention distribution and the second attention distribution of the attention score are calculated using the softmax function and the softmin function, respectively. The positive spatial features of the neighboring base station are obtained by weighted summation of the historical traffic records of the neighboring base station using the first attention distribution, and the negative spatial features of the neighboring base station are obtained by weighted summation of the historical traffic records of the neighboring base station using the second attention distribution. The positive and negative spatial features of the neighboring base stations are determined as spatial features.
5. The method according to claim 1, characterized in that, The step of predicting the predicted traffic of the base station to be predicted based on feature data includes: The feature data is input into N predictors respectively, and N first predicted flows are output, where N>2. The N predictors are divided into M+1 groups of predictors, M≥1. The first predicted flows obtained by different groups of predictors are at least partially different in scale. The first predicted flow obtained by each predictor group is decomposed into a first predicted series set. Each first predicted series set includes P series terms, where p is the index of the series term, and 0≤p≤P-1. The weighted sum of the series terms after decomposition is obtained by using a weight matrix; the sum of the weight parameters of the series terms p in each first prediction series set is 1.
6. The method according to claim 5, characterized in that, The first predicted flow decomposition outputs from different predictors in each predictor group yields different indexes for the series terms.
7. The method according to claim 5, characterized in that, The predictor is a temporal convolutional network.
8. A base station traffic prediction device, characterized in that, include: The first filtering module is used to determine the neighboring base stations of the base station to be predicted based on the geographical distance information between base stations. The second filtering module is used to determine the neighboring base stations among the nearby base stations, and there is a causal relationship between the historical traffic records of the neighboring base stations and the historical traffic records of the base station to be predicted. The feature extraction module is used to extract features from the historical traffic records of the base station to be predicted and the neighboring base stations to obtain feature data; The prediction module is used to predict the predicted traffic of the base station to be predicted based on feature data.
9. An electronic device, characterized in that, Including processors, The processor is used to determine the neighboring base stations of the base station to be predicted based on the geographical distance information between base stations. Identify neighboring base stations among the nearby base stations, and find a causal relationship between the historical traffic records of the neighboring base stations and the historical traffic records of the base station to be predicted; Feature data is obtained by extracting features from the historical traffic records of the base station to be predicted and the neighboring base stations; Predict the predicted traffic of the base station to be predicted based on the feature data.
10. An electronic device, characterized in that, include: A processor, a memory, and a program stored in the memory and executable on the processor, wherein the program, when executed by the processor, implements the steps of the base station traffic prediction method as described in any one of claims 1 to 7.
11. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the steps of the base station traffic prediction method as described in any one of claims 1 to 7.