Traffic prediction method and device
By combining the prediction methods of PCA, CNN and IGRU models, the historical traffic data is dimensionalized and feature extraction is performed, which solves the problem of low accuracy in network traffic prediction and achieves more efficient traffic prediction and resource scheduling.
Patent Information
- Application Number
- CN202510586168.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-07
- Publication Date
- 2025-07-11
AI Technical Summary
In the prior art, the accuracy of network traffic prediction is low, making it difficult to effectively utilize wireless network resources, affecting the reasonable scheduling and energy efficiency of base station resources.
A pre-trained prediction model is adopted, including the first sub-model for dimensionality reduction, the second sub-model for feature extraction, and the third sub-model for prediction. Combining PCA, CNN and IGRU models, a prediction model is constructed through covariance matrix and feature vector transformation to perform feature extraction and prediction of historical traffic data.
It improves the accuracy of traffic prediction, realizes efficient processing of historical traffic data, improves feature recognition accuracy and prediction accuracy, and supports reasonable scheduling of base station resources and energy efficiency optimization.
Smart Images

Figure CN120301784A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of communication technologies, and in particular, to a traffic prediction method and apparatus. Background Art
[0002] Nowadays, network traffic demands have become more diverse and complex. In the face of huge mobile data traffic, it is necessary to efficiently allocate base station resources and improve the user experience. A common method is to predict traffic and then reasonably schedule base station resources, make full use of wireless network resources to improve energy efficiency, and contribute to achieving the goal of energy conservation and emission reduction. However, in current research on network traffic prediction, most use basic models for prediction, and the information captured by the models and the prediction accuracy are poor. Summary of the Invention
[0003] Embodiments of this application provide a traffic prediction method and apparatus to at least solve the technical problem of low prediction accuracy in related technologies.
[0004] According to one aspect of the embodiments of this application, a traffic prediction method is provided, which is characterized by including: obtaining a prediction period; using a pre-trained prediction model to predict the base station to be predicted, and obtaining a prediction result, where the prediction result is used to represent the traffic of the base station to be predicted during the prediction period, and the prediction model includes: a first sub-model, a second sub-model, and a third sub-model. The first sub-model is used to perform dimensionality reduction on historical traffic data, the second sub-model is used to extract features from the dimension-reduced historical traffic data, and the third sub-model is used to perform prediction based on the extracted feature vector to obtain the prediction result; outputting the prediction result.
[0005] Optionally, the prediction model is obtained through the following method, including: connecting the first sub-model, the second sub-model, and the third sub-model in sequence to obtain the prediction model; obtaining the historical traffic data, and performing normalization processing on the historical traffic data to obtain processed historical traffic data; using the processed historical traffic data to complete the training of the prediction model.
[0006] Optionally, training the prediction model using the processed historical traffic data includes: constructing a covariance matrix based on the processed historical traffic data, where the covariance matrix is used to represent the linear correlation degree between the features of the historical traffic data; obtaining a plurality of eigenvalues of the covariance matrix and the eigenvectors corresponding to the plurality of eigenvalues; selecting a preset number of eigenvectors with the largest eigenvalues from the plurality of eigenvectors to be determined as target eigenvectors; constructing a first matrix based on the target eigenvectors, and forming a second matrix with the processed historical traffic data, multiplying the first matrix and the second matrix to obtain the transformed historical traffic data; determining the transformed historical traffic data as the input of the second sub-model to complete the training of the prediction model.
[0007] Optionally, using the transformed historical traffic data as the input of the second sub-model to complete the training of the prediction model includes: extracting features from the transformed historical traffic data using the second sub-model to obtain an embedding vector; determining the embedding vector as the input of the third sub-model to complete the training of the prediction model.
[0008] Optionally, determining the embedding vector as the input of the third sub-model to complete the training of the prediction model includes: constructing the third sub-model, where the third sub-model includes: a target forget gate and a target update gate; determining the embedding vector as an input vector, and sequentially processing the input vector through the target forget gate and the target update gate to obtain an output result.
[0009] Optionally, the method further includes: receiving the input vector at the current moment and the hidden state vector at the previous moment through the target forget gate; after processing the input vector at the current moment and the hidden state vector at the previous moment through the target forget gate, obtaining a first output vector; performing a scaling process on the first output vector to obtain a scaled first output vector.
[0010] Optionally, the method further includes: receiving the input vector at the current moment, the hidden state vector at the previous moment, and the state vector at the previous moment through the target update gate; processing the input vector at the current moment, the hidden state vector at the previous moment, and the state vector at the previous moment through the target update gate to obtain a second output vector; performing a scaling process on the second output vector to obtain a scaled second output vector.
[0011] Optionally, the method further includes: obtaining the scaled first output vector and the scaled second output vector; determining the state vector at the current moment according to the scaled first output vector and the scaled second output vector.
[0012] Optionally, the method further includes: determining a candidate hidden state vector according to the state vector at the current moment; and determining the hidden state vector at the current moment according to the candidate hidden state vector and the second output vector after scaling processing.
[0013] According to another aspect of the embodiments of the present application, there is also provided a traffic prediction device, including: an acquisition module, configured to acquire a prediction period; a prediction module, configured to perform prediction on the to-be-predicted base station by using a pre-trained prediction model to obtain a prediction result, where the prediction result is used to represent the traffic volume of the to-be-predicted base station during the prediction period, and the prediction model includes: a first sub-model, a second sub-model, and a third sub-model, where the first sub-model is configured to perform dimensionality reduction on historical traffic data, the second sub-model is configured to extract features from the historical traffic data after dimensionality reduction, and the third sub-model is configured to perform prediction based on the extracted feature vector to obtain the prediction result; and an output module, configured to output the prediction result.
[0014] According to yet another aspect of the embodiments of the present application, there is also provided a communication device, including: a memory and a processor, where the memory is used to store program instructions; and the processor is connected to the memory and is configured to execute the above-mentioned traffic prediction method.
[0015] According to still another aspect of the embodiments of the present application, there is also provided a computer program product, including computer instructions, which, when executed by a processor, implement the above-mentioned traffic prediction method.
[0016] In the embodiments of the present application, a prediction period is acquired; a pre-trained prediction model is used to perform prediction on the to-be-predicted base station to obtain a prediction result, where the prediction result is used to represent the traffic volume of the to-be-predicted base station during the prediction period, and the prediction model includes: a first sub-model, a second sub-model, and a third sub-model, where the first sub-model is used to perform dimensionality reduction on historical traffic data, the second sub-model is used to extract features from the historical traffic data after dimensionality reduction, and the third sub-model is used to perform prediction based on the extracted feature vector to obtain the prediction result; and the prediction result is output. By using the prediction model to perform prediction on the to-be-predicted base station and obtaining the prediction result, the purpose of performing dimensionality reduction on the historical traffic data and then extracting features to obtain the traffic data during the prediction period is achieved, thereby realizing the technical effect of sequentially processing the traffic data through three sub-models to improve the feature recognition accuracy and the prediction accuracy, and further solving the technical problem of low prediction accuracy in the related art. Description of the Drawings
[0017] The accompanying drawings described herein are used to provide a further understanding of the present application, and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application, and do not constitute an improper limitation of the present application. In the drawings:
[0018] Figure 1 is a hardware structure block diagram of a computer terminal for implementing a traffic prediction method according to an embodiment of the present application;
[0019] Figure 2 is a flowchart of a traffic prediction method according to an embodiment of the present application;
[0020] Figure 3 is a flowchart of state vector processing of an IGRU model according to an embodiment of the present application;
[0021] Figure 4 is a structure diagram of a prediction model according to an embodiment of the present application;
[0022] Figure 5 is another flowchart of a traffic prediction method according to an embodiment of the present application;
[0023] Figure 6 is a comparison diagram of the true value and predicted value of traffic of a prediction model according to an embodiment of the present application;
[0024] Figure 7 is a structure diagram of a traffic prediction device according to an embodiment of the present application. Detailed implementation manners
[0025] In order to enable those skilled in the art to better understand the solution of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0026] It should be noted that the terms "first", "second", etc. in the description, claims and the above-mentioned drawings of this application are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of this application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0027] To better understand the embodiments of this application, the technical terms involved in the embodiments of this application are explained as follows:
[0028] CNN: (Convolutional Neural Networks, convolutional neural network): A deep learning network model specifically used to process and recognize data with spatial structure. The inside of the CNN mainly consists of several key parts: convolutional layer, activation function, pooling layer, fully connected layer and Softmax layer. The CNN has a good feature extraction function. Three-dimensional convolution can extract features in both time and space at the same time, and can be used for behavior recognition and video processing.
[0029] GRU: (Gated Recurrent Unit, gated recurrent unit): A deep learning model commonly used to process sequence data, especially in the fields of natural language processing and time series analysis. The GRU model is a variant of the recurrent neural network (RNN), aiming to solve the problems of long-term dependence and gradient disappearance in traditional RNN models. It controls the flow of information by introducing a gating mechanism, thereby effectively capturing long-term dependence relationships in sequence data.
[0030] IGRU: (Improved Gated Recurrent Unit, improved gated recurrent unit structure): A new type of gated unit structure proposed in this application for 5G base station traffic prediction. The IGRU combines the characteristics of the LSTM and GRU, optimizes the gated unit algorithm, and adds a data scaling module to the gated unit, effectively improving the model prediction effect.
[0031] PCA: (Principal Component Analysis) is a method widely used in the fields of data analysis and machine learning. It transforms multiple variables into a smaller number of important variables through linear transformation. These new variables are called principal components. The principal components are linearly independent and are sorted according to the variance they explain. The core idea of PCA is to map the data in the high-dimensional space to the low-dimensional space through dimensionality reduction, thereby simplifying the data and retaining most of the information.
[0032] The information collected in the embodiments of this application is information and data authorized by the user or fully authorized by all parties. For the processing of relevant data such as collection, storage, use, processing, transmission, provision, disclosure, and application, all comply with the relevant laws, regulations, and standards of the relevant regions, take necessary confidentiality measures, do not violate public order and good customs, and provide corresponding operation entrances for users to choose to authorize or reject the results of automated decision-making; if the user chooses to reject, the expert decision-making process will be entered.
[0033] To solve the problems existing in the related art, the embodiments of this application provide a traffic prediction method, which can run on Figure 1 the computer terminal shown below, and the following is an explanation of this computer terminal.
[0034] The traffic prediction method embodiments provided in the embodiments of this application can be executed on a mobile terminal, a computer terminal, or a similar computing device. Figure 1 A hardware structure block diagram of a computer terminal for implementing the traffic prediction method is shown. As Figure 1 shown, the computer terminal 10 may include one or more processors (the processors may include, but are not limited to, processing devices such as a microprocessor MCU or a field-programmable gate array FPGA), shown as 102a, 102b,..., 102n in the figure, a memory 104 for storing data, and a transmission module 106 for communication functions connected by wired and / or wireless networks. In addition, it may further include: a display, a keyboard, a cursor control device, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the I / O interface), a network interface, and a BUS bus. Those of ordinary skill in the art can understand that Figure 1 the structure shown is only schematic and does not limit the structure of the above-mentioned electronic device. For example, the computer terminal 10 may further include more or fewer components than Figure 1 shown, or have a different configuration from Figure 1 shown.
[0035] It should be noted that the above one or more processors and / or other data processing circuits can generally be referred to as "data processing circuits" herein. The data processing circuit can be embodied in software, hardware, firmware, or any combination thereof, in whole or in part. In addition, the data processing circuit can be a single independent processing module, or be incorporated in whole or in part into any one of other elements in the computer terminal 10. As involved in the embodiments of the present application, the data processing circuit is a kind of processor control (such as the selection of a variable resistance terminal path connected to an interface).
[0036] The memory 104 can be used to store software programs and modules of application software, such as the program instructions / data storage devices corresponding to the traffic prediction method in the embodiments of the present application. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory 104, that is, to implement the above-mentioned traffic prediction method. The memory 104 can include a high-speed random access memory, and can also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memories. In some instances, the memory 104 can further include a memory remotely provided relative to the processor, and these remote memories can be connected to the computer terminal 10 through a network. Examples of the above network include but are not limited to the Internet, intranet, local area network, mobile communication network, and combinations thereof.
[0037] The transmission module 106 is used to receive or send data via a network. Specific examples of the above network can include a wireless network provided by the communication provider of the computer terminal 10. In one instance, the transmission module 106 includes a network adapter (Network Interface Controller, NIC), which can be connected to other network devices through a base station and thus can communicate with the Internet. In one instance, the transmission module 106 can be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0038] The display can be, for example, a touch-screen liquid crystal display (LCD), which enables the user to interact with the user interface of the computer terminal 10.
[0039] It should be noted here that in some alternative embodiments, the above Figure 1 shown computer terminal can include hardware elements (including circuits), software elements (including computer code stored on a computer-readable medium), or a combination of both hardware elements and software elements. It should be pointed out that Figure 1 is only an example of a specific specific instance, and is intended to show the types of components that can exist in the above computer terminal.
[0040] Under the above operating environment, an embodiment of a traffic prediction method is provided in an embodiment of the present application. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0041] Figure 2 is a flowchart of a traffic prediction method according to an embodiment of the present application, as Figure 2 shown, the method includes the following steps:
[0042] Step S202, obtain a prediction period;
[0043] In step S202, the prediction period can be set according to a specific application scenario, for example: within the next week, within a month.
[0044] Step S204, use a pre-trained prediction model to predict the to-be-predicted base station, and obtain a prediction result, where the prediction result is used to represent the traffic of the to-be-predicted base station during the prediction period. The prediction model includes: a first sub-model, a second sub-model, and a third sub-model. The first sub-model is used to reduce the dimension of historical traffic data, the second sub-model is used to extract features from the dimension-reduced historical traffic data, and the third sub-model is used to perform a prediction based on the extracted feature vector to obtain the prediction result;
[0045] Step S206, output the prediction result.
[0046] Through the above steps S202 to S206, obtain a prediction period; use a pre-trained prediction model to predict the to-be-predicted base station, and obtain a prediction result, where the prediction result is used to represent the traffic of the to-be-predicted base station during the prediction period. The prediction model includes: a first sub-model, a second sub-model, and a third sub-model. The first sub-model is used to reduce the dimension of historical traffic data, the second sub-model is used to extract features from the dimension-reduced historical traffic data, and the third sub-model is used to perform a prediction based on the extracted feature vector to obtain the prediction result; output the prediction result. By predicting the to-be-predicted base station through the prediction model and obtaining the prediction result, the purpose of reducing the dimension of historical traffic data and then extracting features to obtain the traffic data during the prediction period is achieved, thereby realizing the technical effect of sequentially processing the traffic data through three sub-models to improve the feature recognition accuracy and the prediction accuracy, and further solving the technical problem of low prediction accuracy in the related art. The following is a detailed description.
[0047] In the technical solution provided in step S204 of the above traffic prediction method, the prediction model can be determined in the following manner: connect the first sub-model, the second sub-model, and the third sub-model in sequence to obtain the prediction model; obtain the historical traffic data, and perform standardization processing on the historical traffic data to obtain the processed historical traffic data; use the processed historical traffic data to complete the training of the prediction model. By predicting the historical traffic data through the prediction model, more accurate prediction data can be obtained, thereby improving the traffic prediction accuracy.
[0048] Specifically, the specific steps of using the processed historical traffic data to complete the training of the prediction model are as follows: construct a covariance matrix based on the processed historical traffic data, and the covariance matrix is used to represent the linear correlation degree between the features of the historical traffic data; obtain multiple eigenvalues of the covariance matrix and the eigenvectors corresponding to the multiple eigenvalues; select a preset number of eigenvectors with the largest eigenvalues from the multiple eigenvectors and determine them as target eigenvectors; construct a first matrix based on the target eigenvectors, and form a second matrix with the processed historical traffic data, multiply the first matrix and the second matrix to obtain the transformed historical traffic data; determine the transformed historical traffic data as the input of the second sub-model to complete the training of the prediction model. By performing transformation processing on the historical traffic data, the historical traffic data can be used as the input of the second sub-model to achieve further prediction and improve the traffic prediction accuracy.
[0049] In some embodiments of the present application, a combined model (prediction model) of PCA-CNN-IGRU is adopted. It can be understood that the first sub-model is PCA, the second sub-model is CNN, and the third sub-model is IGRU.
[0050] Figure 3 Shows the method steps for predicting the traffic of a base station, such as Figure 3 As shown, the use of PCA is to reduce the n-dimensional attributes to k dimensions (k < n). These k features are re-constructed from the original n-dimensional features, and the obtained k-dimensional features are determined as the main components. The specific steps of PCA include: data standardization. Data of different features usually have different units and magnitudes. To ensure that each feature has the same importance in PCA analysis, standardization can be performed using the mean and standard deviation. Establish a covariance matrix. The covariance matrix is used to measure the linear correlation between the features in the data set. For an n-dimensional data set, the covariance matrix is in the form of n*n, and the calculation formula for the elements in the covariance matrix is shown in the following formula:
[0051]
[0052] Wherein, c[i][j] represents the covariance between the i-th feature and the j-th feature in the covariance matrix C, x[i] and x[j] are the values of each sample in the dataset on the i-th feature and the values of each sample in the dataset on the j-th feature respectively, u[i] represents the average value of each sample in the dataset on the i-th feature, u[j] represents the average value of each sample in the dataset on the j-th feature, and n represents the number of samples in the dataset.
[0053] The calculation formula for the eigenvectors of the covariance matrix C is as follows:
[0054] |C - λI| = 0
[0055] Wherein, λ represents the eigenvalue, and I represents the identity matrix.
[0056] Sort the obtained eigenvalues from largest to smallest, select the eigenvectors corresponding to the top k largest eigenvalues, and determine them as the target eigenvectors to construct a matrix P (the first matrix). Through matrix transformation, the historical traffic data is transformed into a new k-dimensional space. Multiply the original data matrix X (the second matrix) by the matrix P obtained in the previous step to obtain the transformed matrix T (the transformed historical traffic data), that is, T = XP.
[0057] In some embodiments of the present application, using the transformed historical traffic data as the input of the second sub-model to complete the training of the prediction model includes: extracting features from the transformed historical traffic data using the second sub-model to obtain an embedding vector; determining the embedding vector as the input of the third sub-model to complete the training of the prediction model. By extracting features from the transformed historical traffic data to obtain an embedding vector, the traffic can be predicted more accurately.
[0058] It should be noted that CNN is a type of deep learning network model specifically designed to process and recognize data with spatial structures, and it is one of the representative algorithms of deep learning. Although CNN was initially designed to process image data, due to its strong feature extraction ability and advantages such as local connection, weight sharing, and downsampling, CNN is also very suitable for processing time series data. In this model, CNN is used to mine key features in time series. Its convolution operation is as follows, where w represents a one-dimensional convolution kernel, x represents the input feature vector, is the convolution operation, b is the bias vector, and f is the activation function.
[0059]
[0060] In some embodiments of the present application, determining the embedding vector as the input of the third sub-model to complete the training of the prediction model includes: constructing the third sub-model, where the third sub-model includes: a target forgetting gate and a target update gate; determining the embedding vector as the input vector, and sequentially processing the input vector through the target forgetting gate and the target update gate to obtain an output result. Processing the input vector through the third sub-model to complete the training of the prediction model, thereby improving the accuracy of traffic prediction.
[0061] Specifically, the target forgetting gate receives the input vector at the current moment and the hidden state vector at the previous moment; after processing the input vector at the current moment and the hidden state vector at the previous moment through the target forgetting gate, a first output vector is obtained; the first output vector is subjected to a scaling process to obtain a scaled first output vector. Through the scaling process, the stability and prediction performance of the entire model can be improved.
[0062] It should be noted that LSTM is composed of three gate structures: an input gate, a forgetting gate, and an output gate, and GRU is composed of two gate structures: an update gate and a reset gate. The IGRU proposed in the embodiments of the present application is composed of an improved forgetting gate (target forgetting gate) and an improved update gate (target update gate) in two parts, as Figure 4 shown, which combines the advantages of the LSTM and GRU models and transforms the gating unit.
[0063] It should be noted that the target forgetting gate includes: a forgetting gate and a scaling module, and the target update gate includes: an update gate and a scaling module.
[0064] In the embodiments of the present application, the output range of the basic forgetting gate is [0, 1]. A data scaling module is introduced into the target forgetting gate in IGRU, and the output range of the forgetting gate is changed to [0, 0.76], thereby affecting the forgetting of the current data, so that as much previous information as possible can be retained. The calculation process is as follows:
[0065] f t =σ(W f ·[h t-1 , x t +b i )
[0066] DS=tanh(f t )
[0067] In the formula, f t represents the first output vector, x t is the input vector at time t, h t-1 stores the data at time t - 1 (the hidden state vector at the previous moment), ht-1 Perform a linear transformation by multiplying with a weight matrix. Then calculate through the σ function, where σ is the Sigmoid activation function, thus obtaining the numerical result f t is between (0, 1). DS represents the data scaling module, which is used to scale the output vector. W f represents the weight matrix of the target forgetting gate in the IGRU. [h t-1 , x t means concatenating h t-1 and x t . b i represents the bias term of the target forgetting gate in the IGRU, and tanh represents the activation function.
[0068] It should be noted that the scaling process is to narrow the range of the numerical result of the first output vector from (0, 1) to a preset interval, for example: [0, 0.76]. The data scaling module adjusts the output of the gating unit from [0, 1] to a smaller interval through linear transformation, thereby reducing the proportion of information forgotten and helping the model retain more features of historical data. This operation in the IGRU is to solve specific problems, such as better retaining the long-term features of time series in traffic prediction and improving the prediction accuracy. The use of the scaling module reflects the flexibility and innovation in the design of deep learning models and can help the model adjust its behavior according to specific task requirements.
[0069] In some embodiments of the present application, receive the input vector at the current moment, the hidden state vector at the previous moment, and the state vector at the previous moment through the target update gate; process the input vector at the current moment, the hidden state vector at the previous moment, and the state vector at the previous moment through the target update gate to obtain a second output vector; perform scaling processing on the second output vector to obtain the scaled second output vector.
[0070] Among them, the state vector is determined in the following manner: obtain the scaled first output vector and the scaled second output vector; determine the state vector at the current moment according to the scaled first output vector and the scaled second output vector.
[0071] Among them, the hidden state vector at the current moment is determined in the following manner: determine the candidate hidden state vector according to the state vector at the current moment; determine the hidden state vector at the current moment according to the candidate hidden state vector and the scaled second output vector.
[0072] Specifically, the target update gate adds the state vector C at the previous moment t-1The calculation further extracts important data. The Sigmoid activation function in the target update gate has a saturation region, which reduces the learning efficiency of the model. Therefore, IGRU introduces a data scaling module in the target update gate to effectively prevent over-saturation. The calculation process is as follows:
[0073] r t = σ(W r ·[h t-1 , x t , C t-1 +b r )
[0074] DS = tanh(r t )
[0075] In the formula, W r represents the weight matrix of the target update gate, r t represents the second output vector, tanh represents the activation function, b r represents the bias term of the target update gate, DS represents the data scaling module, which is used to scale the output vector, [h t-1 , x t , C t-1 represents the concatenation of h t-1 , x t , C t-1 , C t-1 represents the state vector of the previous moment.
[0076] The state vector C t at the current moment is obtained by multiplying the state vector C t-1 of the previous moment with DS(f t )(the first output vector after scaling), and then adding DS(r t )(the second output vector after scaling) to obtain the data information retained at the current moment (the state vector C t ) at the current moment. The calculation process is as follows:
[0077] C t = DS(f t )·C t-1 + DS(r t )
[0078] When calculating the hidden state vector h t at the current moment, IGRU introduces a candidate hidden state vector Stores the obtained original hidden state vector information into And then add In the formula is used to extract the weight information in the hidden state vector information, and the result is combined with r t in the target update gatePerform operations to further determine the information with a larger weight to be carried. The calculation process is as follows:
[0079]
[0080] It should be noted that scaling is a commonly used technique in the data preprocessing stage, mainly used to change the scale of features so that all features are within a similar numerical range. There are mainly two common methods: standardization and normalization. Among them, standardization is also called Z-score normalization, and its purpose is to transform the distribution of each feature into a normal distribution with a mean of 0 and a standard deviation of 1; normalization is also called Min-Max Scaling, which scales the value of each feature to a specific range, usually [0, 1], or it can also be [-1, 1] or any other specified range.
[0081] The traffic prediction method proposed in this application reduces the complexity of the model input data compared with using a basic neural network model for prediction. The combination of CNN and IGRU can maximize the advantages of the two models, better learn the long-term dependencies in time series data, and further alleviate the problems of gradient vanishing and gradient explosion in time series prediction. The traffic prediction accuracy of 5G base stations is effectively improved.
[0082] Figure 5 Shows a specific implementation process of 5G base station traffic prediction based on a prediction model, as Figure 5 shown, including:
[0083] Obtain data;
[0084] Perform principal component analysis on the obtained data;
[0085] Preprocess the analyzed data;
[0086] Partition the preprocessed data into a training set, a test set, and a validation set;
[0087] Complete model training using the partitioned dataset, and use the trained model to predict traffic data.
[0088] Comparison of the true value and predicted value of the traffic of the PCA-CNN-IGRU model is as Figure 6 shown, where the blue curve (PCA-CNN-IGRU-TEST) represents the true value, and the yellow curve (PCA-CNN-IGRU-PREDICT) represents the predicted value. It can be seen that the two are basically the same.
[0089] Table 1 shows the prediction results of different models for the next-day traffic data. It can be seen that the PCA-CNN-IGRU model has a higher prediction accuracy.
[0090] Table 1
[0091]
[0092] Figure 7 A traffic prediction device is shown, and the device includes:
[0093] An acquisition module 70, configured to acquire a prediction period;
[0094] A prediction module 72, configured to perform prediction on the base station to be predicted by using a pre-trained prediction model to obtain a prediction result, where the prediction result is used to represent the traffic of the base station to be predicted during the prediction period, and the prediction model includes: a first sub-model, a second sub-model, and a third sub-model. The first sub-model is configured to perform dimensionality reduction on historical traffic data, the second sub-model is configured to perform feature extraction on the dimensionality-reduced historical traffic data, and the third sub-model is configured to perform prediction based on the extracted feature vectors to obtain the prediction result;
[0095] An output module 74, configured to output the prediction result.
[0096] Acquire a prediction period; perform prediction on the base station to be predicted by using a pre-trained prediction model to obtain a prediction result, where the prediction result is used to represent the traffic of the base station to be predicted during the prediction period, and the prediction model includes: a first sub-model, a second sub-model, and a third sub-model. The first sub-model is used to perform dimensionality reduction on historical traffic data, the second sub-model is used to perform feature extraction on the dimensionality-reduced historical traffic data, and the third sub-model is used to perform prediction based on the extracted feature vectors to obtain the prediction result; output the prediction result. By predicting the base station to be predicted through the prediction model and acquiring the prediction result, the purpose of performing dimensionality reduction on historical traffic data and then performing feature extraction to obtain traffic data during the prediction period is achieved, thereby realizing the technical effect of sequentially processing traffic data through three sub-models to improve feature recognition accuracy and prediction accuracy, and further solving the technical problem of low prediction accuracy in the related art.
[0097] The prediction module 72 includes: an acquisition sub-module, configured to acquire a prediction model, including: connecting the first sub-model, the second sub-model, and the third sub-model in sequence to obtain the prediction model; acquiring the historical traffic data, and performing standardization processing on the historical traffic data to obtain processed historical traffic data; using the processed historical traffic data to complete the training of the prediction model.
[0098] The obtaining sub-module includes: a training unit configured to complete the training of the prediction model by using the processed historical traffic data, including: constructing a covariance matrix based on the processed historical traffic data, where the covariance matrix is used to represent the linear correlation degree between the features of the historical traffic data; obtaining a plurality of eigenvalues of the covariance matrix and the eigenvectors corresponding to the plurality of eigenvalues; selecting a preset number of eigenvectors with the largest eigenvalues from the plurality of eigenvectors to be determined as target eigenvectors; constructing a first matrix based on the target eigenvectors, and forming a second matrix with the processed historical traffic data, multiplying the first matrix and the second matrix to obtain the transformed historical traffic data; determining the transformed historical traffic data as the input of the second sub-model to complete the training of the prediction model.
[0099] The training unit is further configured to use the transformed historical traffic data as the input of the second sub-model to complete the training of the prediction model, including: extracting features from the transformed historical traffic data by using the second sub-model to obtain an embedding vector; determining the embedding vector as the input of the third sub-model to complete the training of the prediction model.
[0100] The training unit is further configured to determine the embedding vector as the input of the third sub-model to complete the training of the prediction model, including: constructing the third sub-model, where the third sub-model includes: a target forget gate and a target update gate; determining the embedding vector as an input vector, and sequentially processing the input vector through the target forget gate and the target update gate to obtain an output result.
[0101] The training unit is further configured to receive the input vector at the current moment and the hidden state vector at the previous moment through the target forget gate; after processing the input vector at the current moment and the hidden state vector at the previous moment through the target forget gate, obtaining a first output vector; performing a scaling process on the first output vector to obtain a scaled first output vector.
[0102] The training unit is further configured to receive the input vector at the current moment, the hidden state vector at the previous moment, and the state vector at the previous moment through the target update gate; processing the input vector at the current moment, the hidden state vector at the previous moment, and the state vector at the previous moment through the target update gate to obtain a second output vector; performing a scaling process on the second output vector to obtain a scaled second output vector.
[0103] The training unit is further configured to obtain the first output vector after scaling processing and the second output vector after scaling processing; and determine the state vector at the current moment according to the first output vector after scaling processing and the second output vector after scaling processing.
[0104] The training unit is further configured to determine a candidate hidden state vector according to the state vector at the current moment; and determine the hidden state vector at the current moment according to the candidate hidden state vector and the second output vector after scaling processing.
[0105] It should be noted that Figure 7 The shown traffic prediction device is used to execute Figure 2 The shown traffic prediction method, so the relevant explanations in the above traffic prediction method also apply to this traffic prediction device, which will not be elaborated here.
[0106] An embodiment of the present application further provides a communication device, including: a memory and a processor, wherein the memory is used to store program instructions; the processor, connected to the memory, is used to execute the above traffic prediction method.
[0107] An embodiment of the present application further provides a computer program product, including computer instructions, which implement the steps of the traffic prediction method in the present application when executed by a processor.
[0108] The serial numbers of the above embodiments of the present application are only for description, and do not represent the advantages and disadvantages of the embodiments.
[0109] In the above embodiments of the present application, the descriptions of each embodiment have their own emphases. For the parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.
[0110] In several embodiments provided by the present application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only illustrative. For example, the division of the units can be a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point, the coupling or direct coupling or communication connection shown or discussed with each other can be through some interfaces, and the indirect coupling or communication connection of units or modules can be in an electrical or other form.
[0111] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they can be located in one place, or can be distributed to multiple units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0112] In addition, each functional unit in various embodiments of the present application may be integrated into one processing unit, or each unit may exist physically alone, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of a software functional unit.
[0113] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a communication device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present application. The foregoing storage medium includes: various media such as USB flash drives, read-only memories (ROMs), random access memories (RAMs), mobile hard disks, magnetic disks, or optical discs that can store program codes.
[0114] The above are only the preferred embodiments of the present application. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present application.
Claims
1. A traffic prediction method, characterized in that, Including: Obtain the prediction period; Use a pre-trained prediction model to predict the base station to be predicted, and obtain a prediction result, where the prediction result is used to represent the traffic of the base station to be predicted during the prediction period. The prediction model includes: a first sub-model, a second sub-model, and a third sub-model. The first sub-model is used to reduce the dimension of historical traffic data, the second sub-model is used to extract features from the dimension-reduced historical traffic data, and the third sub-model is used to perform prediction based on the extracted feature vectors to obtain the prediction result; Output the prediction result.
2. The method according to claim 1, characterized in that, The prediction model is obtained through the following method, including: Connect the first sub-model, the second sub-model, and the third sub-model in sequence to obtain the prediction model; Obtain the historical traffic data, and perform standardization processing on the historical traffic data to obtain the processed historical traffic data; Use the processed historical traffic data to complete the training of the prediction model.
3. The method according to claim 2, wherein Using the processed historical traffic data to complete the training of the prediction model includes: Construct a covariance matrix based on the processed historical traffic data, where the covariance matrix is used to represent the linear correlation degree between the features of the historical traffic data; Obtain multiple eigenvalues of the covariance matrix and the eigenvectors corresponding to the multiple eigenvalues; Select a preset number of eigenvectors with the largest eigenvalues from the multiple eigenvectors and determine them as target eigenvectors; Construct a first matrix based on the target eigenvectors, and form a second matrix with the processed historical traffic data, and multiply the first matrix and the second matrix to obtain the transformed historical traffic data; Determine the transformed historical traffic data as the input of the second sub-model to complete the training of the prediction model.
4. The method according to claim 3, characterized in that, Using the transformed historical traffic data as the input of the second sub-model to complete the training of the prediction model includes: Use the second sub-model to extract features from the transformed historical traffic data to obtain an embedding vector; Determine the embedding vector as the input of the third sub-model to complete the training of the prediction model.
5. The method according to claim 4, characterized in that, Determining the embedding vector as the input of the third sub-model to complete the training of the prediction model includes: Construct the third sub-model, where the third sub-model includes: a target forgetting gate and a target update gate; Determine the embedding vector as an input vector, and sequentially process the input vector through the target forgetting gate and the target update gate to obtain an output result.
6. The method according to claim 5, wherein The method further includes: Receive the input vector at the current moment and the hidden state vector at the previous moment through the target forgetting gate; After processing the input vector at the current moment and the hidden state vector at the previous moment through the target forgetting gate, obtain a first output vector; Perform scaling processing on the first output vector to obtain the scaled first output vector.
7. The method according to claim 5, wherein The method further includes: Receive the input vector at the current moment, the hidden state vector at the previous moment, and the state vector at the previous moment through the target update gate; Process the input vector at the current moment, the hidden state vector at the previous moment, and the state vector at the previous moment through the target update gate to obtain a second output vector; Perform a scaling process on the second output vector to obtain a scaled second output vector.
8. The method according to claim 5, characterized in that The method further includes: Obtain the scaled first output vector and the scaled second output vector; Determine the state vector at the current moment according to the scaled first output vector and the scaled second output vector.
9. The method according to claim 8, characterized in that The method further includes: Determine a candidate hidden state vector according to the state vector at the current moment; Determine the hidden state vector at the current moment according to the candidate hidden state vector and the scaled second output vector.
10. A flow prediction device, characterized in that, Includes: An acquisition module for acquiring a prediction period; A prediction module for predicting a base station to be predicted by using a pre-trained prediction model to obtain a prediction result, where the prediction result is used to represent the traffic of the base station to be predicted during the prediction period, and the prediction model includes: a first sub-model, a second sub-model, and a third sub-model. The first sub-model is used to reduce the dimension of historical traffic data, the second sub-model is used to extract features from the dimension-reduced historical traffic data, and the third sub-model is used to perform a prediction based on the extracted feature vector to obtain the prediction result; An output module for outputting the prediction result.
11. A communication device, characterized in that, Includes: A memory and a processor, where the memory is used to store program instructions; The processor, connected to the memory, is used to execute the traffic prediction method according to any one of claims 1 to 9.
12. A computer program product, comprising computer instructions, characterized in that, When the computer instructions are executed by the processor, the traffic prediction method according to any one of claims 1 to 9 is implemented.