Method and apparatus for training machine learning model, and method and apparatus for predicting traffic
By combining machine learning models that extract spatial and temporal features, the problem of insufficient correlation utilization in network traffic prediction is solved, achieving more efficient and accurate traffic prediction and supporting intelligent resource management of optical networks.
Patent Information
- Application Number
- CN202311011970.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-08-11
- Publication Date
- 2026-01-27
- Estimated Expiration
- 2043-08-11
AI Technical Summary
Existing technologies struggle to effectively utilize both the temporal and spatial correlations of network traffic for accurate prediction, resulting in high computational costs or poor prediction performance.
A machine learning model is employed, which includes a spatial feature extraction module, a temporal feature extraction module, and a fully connected layer. By preprocessing network traffic data, extracting spatial features, extracting temporal features, and adjusting parameters, traffic prediction values are generated. If the model does not converge, the parameters are adjusted until convergence is achieved.
It improves the accuracy and efficiency of network traffic prediction, reduces computational costs, enables better utilization of traffic correlation information, and realizes intelligent resource allocation of optical networks.
Smart Images

Figure CN117057440B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates generally to machine learning, and more specifically to the training and use of machine learning models for predicting network traffic. Background Technology
[0002] In recent years, global network traffic (e.g., the Internet (IP), optical networks, etc.) has been growing at an unprecedented rate. This growth in network traffic, combined with increasingly frequent daily population movements, has created tidal traffic patterns, posing greater challenges to the efficient allocation of existing network spectrum resources. In IP networks, optimization mechanisms based on network traffic-driven intelligent resource allocation have already begun to be used. It is hoped that similar mechanisms can be used in optical networks to optimize resource allocation. The prerequisite for such optimization mechanisms is accurate prediction of network traffic. For example, if the changing trends of network traffic can be accurately predicted, reasonable spectrum and routing can be pre-planned in optical networks, and traffic tidal patterns can be addressed by proactively adjusting the on / off states of devices in the optical network, thereby improving resource utilization efficiency and reducing energy consumption, achieving intelligent management and control of the optical network.
[0003] Therefore, a technology is needed that enables accurate prediction of network traffic. Summary of the Invention
[0004] A brief overview of this disclosure is given below to provide a basic understanding of some aspects of it. However, it should be understood that this overview is not an exhaustive summary of this disclosure. It is not intended to identify key or essential parts of this disclosure, nor is it intended to limit the scope of this disclosure. Its purpose is merely to present certain concepts of this disclosure in a simplified form as a prelude to the more detailed description that follows.
[0005] According to one aspect of this disclosure, a method is provided for training a machine learning model for predicting network traffic, wherein the machine learning model includes a spatial feature extraction module, a temporal feature extraction module, and a fully connected layer. The method includes: preprocessing historical traffic data of a target prediction region of the network, including generating a first three-dimensional matrix based on the historical traffic data, wherein the first dimension of the first three-dimensional matrix is a time dimension representing a time period, and the second and third dimensions are spatial dimensions jointly representing spatial locations within the target prediction region, and the value of each element in the first three-dimensional matrix is associated with the historical traffic value of a node at a corresponding spatial location during the corresponding time period; and spatial feature extraction, wherein the spatial feature extraction is based on the temporal feature extraction of the first three-dimensional matrix. The process involves: 1) obtaining multiple two-dimensional matrices by slicing the second three-dimensional matrix in a spatial dimension; 2) extracting temporal features, which is based on multiple sequences obtained by slicing the second three-dimensional matrix in a spatial dimension, wherein the second three-dimensional matrix is obtained by concatenating multiple output two-dimensional matrices from the spatial feature extraction step according to the original temporal dimension; 3) generating predicted values, which includes inputting the multiple feature vectors into a fully connected layer to obtain the traffic prediction value for the next time period at each node; 4) adjusting parameters, which includes comparing the traffic prediction value with historical traffic values and adjusting the model parameters based on the comparison results; and 5) re-executing the spatial feature extraction, temporal feature extraction, predicted value generation, and parameter adjustment steps based on the adjusted model parameters until the machine learning model converges.
[0006] According to another aspect of this disclosure, a method for predicting network traffic is provided, comprising: preprocessing historical traffic data of a target prediction region of the network, including generating an input three-dimensional matrix based on the historical traffic data, wherein a first dimension of the input three-dimensional matrix is a time dimension representing a time period, and a second and third dimension together represent a spatial dimension of a spatial location in the target prediction region, and the value of each element in the input three-dimensional matrix is associated with a historical traffic value of a node at a corresponding spatial location during a corresponding time period; and inputting the input three-dimensional matrix into a machine learning model trained according to the method described above, thereby obtaining a traffic prediction value for the next time period at each node.
[0007] According to another aspect of this disclosure, an apparatus for training a machine learning model for predicting network traffic is provided, comprising: a memory having instructions stored thereon; and a processor configured to execute the instructions stored in the memory to perform a method for training a machine learning model for predicting network traffic as described above.
[0008] According to another aspect of this disclosure, an apparatus for predicting network traffic is provided, comprising: a memory having instructions stored thereon; and a processor configured to execute the instructions stored in the memory to perform the method for predicting network traffic as described in the preceding aspects.
[0009] According to another aspect of this disclosure, a computer-readable storage medium is provided, comprising computer-executable instructions that, when executed by one or more processors, cause the one or more processors to perform a method for training to predict network traffic according to the foregoing aspects of this disclosure, or a method for predicting network traffic according to the foregoing aspects of this disclosure. Attached Figure Description
[0010] The accompanying drawings, which form part of this specification, illustrate embodiments of this disclosure and, together with the specification, serve to explain the principles of this disclosure.
[0011] This disclosure will be more clearly understood with reference to the accompanying drawings and the following detailed description, wherein:
[0012] Figure 1A An example of network traffic distribution at different nodes is shown;
[0013] Figure 1B An example of spatial correlation in network traffic is shown;
[0014] Figure 2 The basic structure of this disclosure is illustrated schematically;
[0015] Figure 3 A conceptual operational flow of a method for training a machine learning model for predicting network traffic according to this disclosure is illustrated schematically.
[0016] Figure 4 The spatial feature extraction according to this disclosure is illustrated schematically;
[0017] Figure 5 A conceptual configuration of the spatial feature extraction module according to this disclosure is illustrated schematically;
[0018] Figure 6 The illustration schematically shows the extraction of time features and generation of traffic prediction values according to this disclosure;
[0019] Figure 7 A conceptual configuration of the time feature extraction module according to this disclosure is illustrated schematically;
[0020] Figure 8 An operational example of training a machine learning model for predicting network traffic using a sliding window according to an embodiment of the present disclosure is illustrated schematically.
[0021] Figure 9 A conceptual operational flow of a method for predicting network traffic according to this disclosure is illustrated schematically;
[0022] Figure 10 An example of network traffic prediction results using the machine learning model according to this disclosure is illustrated.
[0023] Figure 11 An exemplary configuration of a computing device that can implement embodiments of the present disclosure is shown. Detailed Implementation
[0024] The following detailed description is based on the accompanying drawings and provides various exemplary embodiments of the present disclosure to aid in a comprehensive understanding. Various details are included in the following description to aid understanding; however, these details are considered exemplary only and not intended to limit the present disclosure, which is defined by the appended claims and their equivalents. The words and phrases used in the following description are intended only to provide a clear and consistent understanding of the present disclosure. Additionally, descriptions of well-known structures, functions, and configurations may have been omitted for clarity and brevity. Those skilled in the art will recognize that various changes and modifications can be made to the examples described herein without departing from the spirit and scope of the present disclosure.
[0025] As described in the background section above, there is a need for accurate prediction of network traffic, especially optical network traffic. However, accurate prediction of large-scale network traffic, such as optical networks, is a challenging problem.
[0026] Existing techniques for predicting network traffic recognize the temporal correlation of network traffic. Therefore, for example, the future traffic of a node can be predicted based on its historical traffic. A traditional example of network traffic prediction is to use an Autoregressive Integrated Moving Average (ARIMA) model to model the stationary characteristics (i.e., the linear component) of traffic data, thereby predicting traffic. In recent years, deep learning-based traffic prediction techniques have emerged, which can learn the non-linear components that are difficult to model using traditional methods, thus improving upon traditional modeling-based prediction methods. For example, Recurrent Neural Networks (RNNs) and Long Short-Term Memory (LSTM) networks can be applied to this deep learning-based traffic prediction technique.
[0027] However, network traffic exhibits spatial correlation in addition to temporal correlation. For example, user mobility within a communication network can lead to correlations in traffic between adjacent areas. Furthermore, users in neighboring regions may have similar network usage habits, or neighboring regions may provide the same network services to similar user groups; therefore, traffic changes in adjacent regions exhibit spatial correlation. Figure 1A and 1B This spatial correlation is illustrated schematically. Figure 1A The data shows network traffic values at different nodes in the network node grid. Figure 1B This shows the normalized value of this network traffic figure. For example... Figure 1B As shown, the horizontal and vertical axes of this matrix can represent the geographical location of each node in the network node grid (e.g., the node number in the node grid), and the grayscale of each element of the matrix represents the normalized value of the network traffic at that node relative to the maximum and minimum traffic of the entire network. Figure 1B As shown, the darker the grayscale, the greater the traffic to that node. According to... Figure 1A and 1B It can be seen that the network traffic of nodes has spatial correlation.
[0028] However, on the one hand, if we only use the time-based predictions for individual nodes to predict network traffic for the entire network, the computational cost will increase rapidly as the number of nodes grows, making it difficult to deploy in practice. On the other hand, relying solely on temporal correlations while ignoring spatial correlations for traffic prediction will not yield satisfactory prediction performance.
[0029] Currently, only a very small number of algorithms simultaneously consider temporal and spatial correlations, such as 3D Convolutional Neural Networks (3DCNN) and ResNet-LSTM. However, these algorithms still have certain limitations in mining and extracting correlations from traffic data in different regions. They may have difficulty converging, and the traffic correlation information cannot be fully utilized.
[0030] In view of the above problems, this disclosure proposes a new machine learning model for predicting network traffic.
[0031] Figure 2 The basic architecture of the scheme disclosed herein is illustrated schematically.
[0032] like Figure 2As shown, the machine learning model for predicting network traffic according to this disclosure may include, for example, a spatial feature extraction module, a temporal feature extraction module, and a fully connected layer. For example, the spatial feature extraction module is used to extract the spatial features of historical traffic data of nodes located at various spatial locations in the target prediction region of the network, the temporal feature extraction module is used to extract the temporal features of such historical traffic data, and the fully connected layer is used to implement the mapping from the extracted features to the prediction result.
[0033] like Figure 2 As shown, the input to the machine learning model according to this disclosure can be a three-dimensional matrix. One dimension of this three-dimensional matrix is a time dimension, for example, representing a time period (in other words, a time interval). For example, each coordinate of this time dimension can indicate a specific time period, such as 0:00-1:00, 1:00-2:00, etc. The other two dimensions of this three-dimensional matrix are spatial dimensions, for example, jointly representing the spatial location of the target prediction region (in other words, the node at that spatial location). For example, the node group can be represented as an I×J grid based on the spatial location of the nodes in the network, and these two spatial dimensions can respectively represent the corresponding node numbers of each node in the network node grid on the horizontal and vertical axes of the grid, for example, as shown. Figure 2 The examples shown are (1,1), (1,2), etc. For instance, these two spatial dimensions can be longitude and latitude, and each coordinate in these dimensions can indicate a specific longitude and latitude value. It should be understood that spatial dimensions can take any other form that can indicate the spatial location of nodes in the network, and are not limited to node numbers and latitude / longitude. The value of each element in this three-dimensional matrix can represent the historical flow value of a node at a corresponding spatial location during a corresponding time period. For example, a node... It can represent the network traffic of traffic node (i, j) in time period t.
[0034] The machine learning model disclosed herein can process the input three-dimensional matrix and output information representing the predicted traffic flow for the next time period at each node in the target prediction region. For example, as Figure 2 As shown, the model can output a two-dimensional matrix, where the value of each element represents the predicted flow rate for the next time period at the node at the corresponding spatial location.
[0035] The following is for reference. Figures 3-7 The training process of the machine learning model for predicting network traffic according to this disclosure is described in detail, and the machine learning model for predicting network traffic according to this disclosure is described in detail in the description of the training process.
[0036] Figure 3A conceptual operational flow of a method 30 for training a machine learning model for predicting network traffic according to this disclosure is illustrated schematically.
[0037] The method begins at step 300.
[0038] In step 302, the historical traffic data for the target prediction region of the network is preprocessed. For example, the target prediction region of the network can be a region covering the entire network. For example, the historical traffic data for the target prediction region can be pre-collected data or data downloaded from a public website.
[0039] According to this disclosure, the preprocessing step may include generating a first three-dimensional matrix based on historical traffic data. The first dimension of the first three-dimensional matrix is a time dimension representing a time period, and the second and third dimensions are spatial dimensions jointly representing spatial locations within the target prediction region. The value of each element in the first three-dimensional matrix is associated with the historical traffic value of a node at a corresponding spatial location during the corresponding time period. For example, each coordinate in the first dimension of the first three-dimensional matrix may indicate a specific time period (during which traffic was statistically analyzed), such as 0:00-1:00, 1:00-2:00, etc. For example, the second and third dimensions may be longitude and latitude, respectively, and each coordinate in these two dimensions may represent the node number of each node in the network node grid. For another example, the second and third dimensions may also indicate specific longitude and latitude values. Of course, the second and third dimensions may also take any other form that can indicate the spatial location of nodes in the network, and are not limited to node numbers and latitude / longitude.
[0040] Preferably, the preprocessing step for historical traffic data may also include some optional additional steps.
[0041] According to a preferred embodiment, outliers in historical traffic data can be replaced with predetermined values. Here, outliers can be, for example, negative values or values whose difference from the average network traffic value exceeds a predetermined threshold (e.g., values significantly exceeding the average network traffic value). For example, outliers can be replaced with 0.
[0042] According to another preferred embodiment, missing values in historical traffic data can be filled with predetermined values. For example, it is desirable to collect traffic values for each node every 15 minutes, but the actual collected traffic data is missing traffic data for one or more time periods in 15-minute increments for one or more nodes. For example, these one or more missing values can be filled with 0. Advantageously, filling missing values can ensure the stability of the structure of the final input three-dimensional matrix (i.e., the first three-dimensional matrix) and the integrity of the data in all dimensions.
[0043] According to another preferred embodiment, historical traffic data can be aggregated. For example, such aggregation may include aggregating historical traffic within multiple adjacent first time periods into traffic within a second time period of equal length to the multiple first time periods. For instance, the collected raw data may be traffic values per node every 15 minutes, and traffic values within four adjacent 15-minute time periods may be aggregated into traffic values within a 1-hour time period. Advantageously, aggregating historical traffic data can make the time interval for which the traffic to be predicted corresponds to the frequency of network scheduling, or it can also reduce the computational burden.
[0044] According to another preferred embodiment, historical traffic data can be normalized. Since the distribution of traffic data values at different nodes may vary significantly—some nodes may consistently maintain a high level of traffic data, while some peripheral nodes may have traffic data around 0—this invention considers normalizing each node individually, thereby advantageously improving the model's convergence speed and prediction accuracy. For example, for each node, traffic can be normalized based on the difference between the maximum and minimum traffic values for each time period. For instance, for the traffic value of the node at coordinate (i, j)... X can be determined i,j maximum value maxX i,j , and the minimum value minX i,j And normalize it using the following formula:
[0045]
[0046] Among them, newX i,j The value is a new normalized value, linearly placed between 0 and 1. Normalization can be performed on all nodes to obtain a normalized first three-dimensional matrix. According to this preferred embodiment, the scale on which normalization is based (i.e., the maximum and minimum flow values of the node) can be retained, so that after obtaining the predicted normalized flow values, denormalization is performed using the same scale to obtain the predicted actual flow values.
[0047] It should be noted that optional additional steps for preprocessing historical traffic data are not limited to those described above. For example, any additional steps that need to be performed as needed may be included, such as converting text-formatted data into numerical-formatted data.
[0048] Back Figure 3In step 304, spatial features can be extracted from the first three-dimensional matrix representing the historical flow of the target prediction area obtained in step 302. According to this disclosure, spatial feature extraction is based on multiple two-dimensional matrices obtained by slicing the first three-dimensional matrix in the time dimension. For example, spatial feature extraction may include slicing the first three-dimensional matrix in the time dimension to obtain multiple two-dimensional matrices corresponding to each time period, and inputting the multiple two-dimensional matrices into a spatial feature extraction module to extract spatial features, wherein the spatial feature extraction module outputs multiple output two-dimensional matrices with the same size as the input multiple two-dimensional matrices.
[0049] Figure 4 The spatial feature extraction according to this disclosure is illustrated schematically. Figure 4 Different shades of gray are used to illustrate the time dimension (i.e., the first dimension) of the first three-dimensional matrix representing historical traffic data. During spatial feature extraction, the first three-dimensional matrix can first be sliced along the time dimension; that is, according to each time period, the first three-dimensional matrix is cut into multiple two-dimensional matrices corresponding to each time period. Advantageously, this slicing avoids destroying the temporal characteristics of network traffic. For example, as... Figure 4 As shown, each two-dimensional matrix obtained after slicing along the time dimension represents the historical traffic of each node in the network within a corresponding time period. These multiple two-dimensional matrices can then be input into a neural network for extracting spatial features. According to this disclosure, spatial feature extraction can be performed on multiple two-dimensional matrices in parallel or partially in parallel, or sequentially. For example, a densely connected convolutional neural network (DenseNet) can be used to extract spatial features. According to this disclosure, the output of the neural network for extracting spatial features can be multiple output two-dimensional matrices, the same number as the input two-dimensional matrices; for example, each output two-dimensional matrix can represent the spatial features of the historical traffic of each node in the network within a corresponding time period.
[0050] According to this disclosure, all output two-dimensional matrices can be integrated. For example, multiple output two-dimensional matrices can be concatenated along their original time dimension to form a second three-dimensional matrix, thereby obtaining, as shown below. Figure 4 The historical traffic data is shown as a three-dimensional spatial feature matrix. It should be understood that each coordinate point in the time dimension of the concatenated second three-dimensional matrix is the same as each coordinate point in the time dimension of the original first three-dimensional matrix. For example, this step can be used as the last sub-step in the spatial feature extraction operation (such as...). Figure 4 As shown, it can also be the first sub-step in the temporal feature extraction operation, or it can be a separate step between the spatial feature extraction operation and the temporal feature extraction operation.
[0051] Figure 5 A conceptual configuration of the spatial feature extraction module according to this disclosure is illustrated schematically. For example... Figure 5 As shown, the spatial feature extraction module according to this disclosure may include, for example, a slicing module and a module implementing a neural network for extracting spatial features. The slicing module may, for example, slice the first three-dimensional matrix along the time dimension as described above. This disclosure preferably uses a DenseNet network as the neural network for extracting spatial features. Figure 5 As shown, a DenseNet module can include multiple basic convolutional modules. Each basic convolutional module can, for example, include at least a convolutional layer, a Linear Rectification Function (ReLU) activation function layer, and a batch regularization layer in sequence. These layers can be configured, for example, using configuration methods known in the art. Traditional feedforward neural networks (such as convolutional neural networks and residual networks like ResNet) use the output of layer (l-1) as the input of layer l to obtain the output X of layer l. l :
[0052] X l =H l (X l-1 )
[0053] Among them, X l and X l-1 Let H represent the output data of the l-th layer and the (l-1)-th layer respectively (for example, these could be two-dimensional matrices), and H... l This represents the processing performed by layer l. DenseNet then connects layer l to all layers preceding it (here, "all layers" refers to each layer in each basic convolutional module):
[0054] X l =H l ([X0,X l ,…,X l-1 ])
[0055] For example, the convolutional layers in each basic convolutional module can use 3×3 convolutions, so that information about the nine surrounding nodes, including the node itself, can be extracted for a given node. By stacking multiple convolutional layers (e.g., via the stacking of basic convolutional modules), a node can form connections with any other node, thereby completing spatial feature extraction.
[0056] Advantageously, by leveraging the more aggressive cascading strategy of DenseNet, the gradient vanishing problem in the spatial feature extraction process is alleviated on the one hand, and feature reuse is achieved through the connection of features on the channel, which strengthens feature propagation while reducing the number of parameters required for computation, thus further improving the accuracy of traffic prediction.
[0057] Specifically, the spatial feature extraction module according to this disclosure is configured such that the size of each output two-dimensional matrix is the same as the size of each input two-dimensional matrix, thereby maintaining the stability of the spatial structure between nodes. For example, this can be ensured by two aspects: First, when configuring the neural network for spatial feature extraction (e.g., DenseNet as described above), unlike the conventional DenseNet configuration, no sampling operation is used; that is, the sampling rate is set to 1 instead of any value less than 1, thus preventing the final extracted spatial feature matrix from being difficult to recover to the size of the actual network node matrix. Second, for each basic convolutional module, the format (in other words, the size) of the input two-dimensional matrix is kept uniform across the layers of the neural network used for spatial feature extraction. In other words, the size of the two-dimensional matrices output across each layer is the same. For example, in the process of extracting spatial features based on DenseNet, the size of the multiple two-dimensional matrices output between the layers of DenseNet is kept the same as the size of the multiple two-dimensional matrices input to the spatial feature extraction module. For example, padding operations can be used to ensure that the size of the two-dimensional matrices output between layers remains constant. For example, after processing by a certain layer, the output two-dimensional matrix may become smaller. In this case, an appropriate number of zero values can be filled around the periphery of the input two-dimensional matrix of that layer (e.g., outward along the first and last rows and the first and last columns of the input two-dimensional matrix) to ensure that the size of the output two-dimensional matrix is the same as the size of the original two-dimensional matrix input to the spatial feature extraction module.
[0058] Already combined Figure 5 The configuration of the spatial feature extraction module according to this disclosure is described. It should be noted that this configuration is merely exemplary, and the spatial feature extraction module according to this disclosure can be implemented as a single module to perform spatial feature extraction operations, without subdividing it into sub-modules such as slicing modules or DenseNet modules. Alternatively, the spatial feature extraction module according to this disclosure can be implemented to include more modules; for example, when matrix integration operations are performed by the spatial feature extraction module, the spatial feature extraction module may further include a matrix integration module that concatenates multiple output two-dimensional matrices along the time dimension.
[0059] Refer again Figure 3In step 306, temporal features can be extracted from the spatial feature matrix representing the historical flow of the target prediction area obtained in step 304. According to this disclosure, temporal feature extraction is based on multiple sequences obtained by slicing the second three-dimensional matrix in a spatial dimension. As mentioned above, the second three-dimensional matrix is obtained by concatenating multiple output two-dimensional matrices from the spatial feature extraction step according to their original temporal dimension. For example, temporal feature extraction may include slicing the second three-dimensional matrix (in other words, the three-dimensional spatial feature matrix of historical flow obtained based on the three-dimensional matrix representing historical flow) in a spatial dimension to obtain multiple sequences, where each sequence represents multiple spatial features of historical flow values corresponding to multiple time periods for each node, and inputting the multiple sequences into the temporal feature extraction module to extract temporal features. The temporal feature extraction model outputs multiple feature vectors corresponding to multiple nodes, and optionally, if matrix concatenation has not been performed before entering step 306, the multiple output two-dimensional matrices are first concatenated into a second three-dimensional matrix according to their original temporal dimension.
[0060] Figure 6 The left side schematically illustrates the extraction of time features according to this disclosure. For example... Figure 6 As shown, after obtaining the three-dimensional spatial feature matrix (i.e., the second three-dimensional matrix) representing historical traffic data, the matrix can first be sliced in the spatial dimension, that is, a one-dimensional vector for each node can be extracted individually. Each value in this vector can represent the spatial feature of the historical traffic value of the corresponding time period for that node. Advantageously, this slicing can avoid destroying the spatial features of network traffic. Then, the multiple sequences can be input into the time feature extraction module to extract time features. According to this disclosure, time feature extraction can be performed on multiple one-dimensional sequences in parallel or partially in parallel, or it can be performed serially. According to this disclosure, the extracted time features can be presented as multiple feature vectors, the number of which can be the same as the number of input one-dimensional vectors, and the size of each time feature vector is the same as the size of each input one-dimensional vector.
[0061] Figure 7 A conceptual configuration of the time feature extraction module according to this disclosure is illustrated schematically. For example... Figure 7As shown, the temporal feature extraction module according to this disclosure may include, for example, a slicing module and a module implementing a neural network for extracting temporal features (as shown in the dashed box). The slicing module may, for example, slice the three-dimensional spatial feature matrix (i.e., the second three-dimensional matrix) in the spatial dimension as described above. This disclosure preferably uses multiple stacked LSTM networks to extract the temporal features of network traffic. Specifically, each one-dimensional sequence obtained by spatial dimensional slicing is passed through multiple stacked LSTM networks to extract temporal features. Advantageously, stacking LSTMs allows for a deeper model depth, resulting in extracted features with deeper temporal characteristics, thereby enabling better extraction of temporal features and thus more accurate predictions. For example, each LSTM network can be configured using configuration methods known in the art.
[0062] In particular, when extracting temporal features using multiple stacked LSTM networks, a subset of neurons from the entire neural network can be used for temporal feature extraction. For example, some neurons can be randomly discarded for temporal feature extraction, thereby minimizing the risk of overfitting.
[0063] It should be noted that this invention is not limited to using LSTM to extract temporal features. For example, other recurrent neural networks, such as gated neural units, can also be used as neural networks for extracting temporal features.
[0064] Refer again Figure 3 In step 308, the traffic prediction value for each node in the target prediction region of the network for the next time period is generated. Specifically, the multiple feature vectors output in step 306 can be input into the fully connected layer, thereby mapping the extracted feature values to the traffic prediction results (in other words, determining the network traffic prediction results from the extracted network traffic feature values). For example, the fully connected layer can be configured according to any method known in the art, as long as the configured fully connected layer can achieve the mapping from features to prediction results.
[0065] Figure 6 The right side schematically illustrates the generation of traffic predictions via the fully connected layer. (Example) Figure 6 As shown, through mapping of the full-connection hierarchy, information representing the predicted traffic values of each node in the region to be predicted in the network during the next time period can be obtained. For example, this information can be in matrix form.
[0066] In step 310, the predicted traffic value of each node can be compared with its historical traffic value and the model parameters can be adjusted based on the comparison result. In step 312, the model convergence can be determined based on the comparison.
[0067] Specifically, an error function can be calculated based on the predicted and historical traffic values for each node, and the model parameters can be adjusted based on the calculation results of the error function. For example, the root mean square error (MSE) and R-squared error (R²) can be calculated as follows: 2 Error functions such as good fit, root mean square error (RMSE), and normalized mean squared error (NMSE) are used to determine whether the obtained error values are less than a predetermined threshold.
[0068]
[0069]
[0070]
[0071]
[0072] in, y represents the predicted flow rate for node i. i This represents the historical traffic value of node i, and n is the total number of nodes in the network.
[0073] If the error value calculated in step 310 is less than a predetermined threshold, then the model is determined to have converged in step 312. The training process of the machine learning model for predicting network traffic according to this disclosure ends, and the trained machine learning model can be used for subsequent network traffic prediction.
[0074] If the error value calculated in step 310 is greater than a predetermined threshold, then in step 312 it is determined that the model has not yet converged. In step 314, the model parameters can be adjusted, and a new round of training can be performed. That is, the spatial feature extraction step, the temporal feature extraction step, the predicted value generation step, and the comparison step are re-executed until the machine learning model converges. For example, the model parameters to be adjusted may include the learning rate and the number of iterations of the machine learning model. Preferably, the Adam algorithm can be used as the optimizer algorithm to help optimize the parameters at a faster speed.
[0075] Already referenced Figures 3-7This document describes the architecture of a machine learning model for predicting network traffic according to this disclosure, as well as the basic training process of such a machine learning model. According to a preferred embodiment, the machine learning model can be trained using a sliding window. According to this preferred embodiment, preprocessing the historical traffic data may further include: dividing the first three-dimensional matrix into multiple supervised learning datasets, each supervised learning dataset including a sliding window three-dimensional matrix and sample labels, wherein the temporal dimension of the sliding window three-dimensional matrix has the same length as the sliding window and the spatial dimension has the same length as the first three-dimensional matrix, and the sample labels represent the historical traffic values of each node to be predicted in the next time period. Furthermore, in this preferred embodiment, spatial feature extraction, temporal feature extraction, prediction value generation, and parameter adjustment are performed by advancing the sliding window time-by-time, and for each sliding window three-dimensional matrix. Specifically, in this preferred embodiment, the sliding window three-dimensional matrix is used instead of the first three-dimensional matrix described above as input into the spatial feature extraction module. According to this preferred embodiment, training can continue based on the parameter-adjusted model parameters until the machine learning model converges.
[0076] Figure 8 This sliding window is illustrated schematically. (Example) Figure 8 As shown, the first three-dimensional matrix representing the historical traffic of the target prediction region of the network has a length of k in the time dimension, that is, it includes the historical traffic in the k time periods from 1 to k. Figure 8 In the example shown, the sliding window length is 3. It should be noted that the sliding window length is schematically drawn as 3 for ease of illustration; in reality, the sliding window length can be any appropriate value (e.g., a value greater than 3). In this example, during data preprocessing, the first three-dimensional matrix can be further divided into multiple supervised learning datasets, including the sliding window three-dimensional matrix and sample labels. The time dimension of the sliding window three-dimensional matrix has the same length as the sliding window (3), and its spatial dimension is the same as the first three-dimensional matrix. The sample labels can be the actual historical flow values representing the flow values of each node in the next time period (i.e., the fourth time period) to be predicted.
[0077] For the current sliding window 3D matrix, spatial feature extraction, temporal feature extraction, predicted value generation, and parameter adjustment can be performed sequentially as described in detail above. After completing the flow prediction based on the current sliding window and adjusting the model parameters based on the comparison between the predicted results and historical values, the sliding window can slide forward one time period, and based on the adjusted model parameters, continue to perform spatial feature extraction, temporal feature extraction, predicted value generation, and parameter adjustment sequentially for the next sliding window 3D matrix. For example, as... Figure 8As shown, the flow value of each node in the fourth time period is predicted by a sliding window three-dimensional matrix that includes three time periods 1, 2, and 3 based on the time dimension. After the model parameters are adjusted by comparing the predicted value with the sample labels, the sliding window can be moved forward, and the new sliding window can include the second, third, and fourth time periods. The corresponding sample labels can be the real historical flow values of the flow values of each node in the fifth time period.
[0078] The sliding window is continuously advanced in the above manner, and predictions and parameter adjustments are made based on the new window until the model converges. It is understandable that after completing one round of window sliding, for example, after the sample labels corresponding to the current sliding window have become the actual historical flow values for each node in the k-th time period, if the model still has not converged, then a new round of sliding window-based training can be started, for example, by re-dividing the first three-dimensional matrix starting from the first time period. Figure 8 As shown, the first supervised learning set for the new round of training can be returned to the sliding window three-dimensional matrix that includes three time periods 1, 2, and 3 in the time dimension, as well as the sample labels representing the historical traffic values of each node for the fourth time period.
[0079] The following example, using urban network traffic prediction as a specific example, illustrates the training process of the machine learning model disclosed herein. In this example, it is assumed that the city is divided into a 100×100 grid, with each area in the grid considered as a node in the urban traffic network, and 50 days of traffic data are collected at 15-minute intervals.
[0080] The training process can begin by generating a three-dimensional matrix based on the collected traffic data. The original three-dimensional matrix can have dimensions of (4800, 100, 100), where the first dimension represents a time period of 15 minutes. Since 96 data points can be collected per day, totaling 4800 data points over 50 days, the first dimension can include 4800 discrete coordinates. The second and third dimensions represent nodes within a divided city grid. In this initial three-dimensional matrix, the network traffic of any node (i, j) in the city grid during time period t can be uniquely represented. This is represented as follows. During the generation of the three-dimensional matrix, the collected traffic data can be processed as detailed above, including outlier replacement and missing value imputation. A 15-minute timeframe does not match the network scheduling frequency (e.g., the network schedules in 1-hour timeframes). Furthermore, excessively short timeframes may lead to a large computational burden. Therefore, in this example, the original three-dimensional data is aggregated, combining four adjacent 15-minute data points into a single 1-hour data point. After aggregation, the three dimensions of the original three-dimensional data matrix become (1200, 100, 100). Next, the aggregated three-dimensional matrix can be normalized as detailed above to obtain the first three-dimensional matrix to be input into the machine learning model.
[0081] Next, the first three-dimensional matrix can be divided into multiple supervised learning datasets. Each supervised learning dataset includes a sliding window three-dimensional matrix and sample labels. In this example, we consider using traffic data from the previous 24 hours to predict the traffic data for the next hour; therefore, we can divide the dataset into 50 supervised learning datasets. The sliding window three-dimensional matrix of each supervised learning dataset has a size of (24, 100, 100), while the sample labels can be a matrix of size (100, 100) representing the historical traffic values of each node for the next hour.
[0082] Subsequently, the sliding window 3D matrix of the first supervised learning set can be input into the spatial feature extraction module according to this disclosure. Specifically, during spatial feature extraction, as described above, the spatial feature extraction module according to this disclosure slices the input sliding window 3D matrix of size (24, 100, 100) along the time dimension into 24 2D matrices of size (100, 100) representing the historical flow values of each node in each hour of the 24-hour period, and inputs these 24 2D matrices into a neural network for extracting spatial features, such as DenseNet.
[0083] According to this disclosure, in this example, the output of the neural network used to extract spatial features is also 24 two-dimensional matrices of size (100, 100), representing the spatial features of the historical flow values of each node in each hour of the 24-hour period. These 24 spatial feature two-dimensional matrices will be reassembled into a second three-dimensional matrix according to the original time dimension for time feature extraction; that is, the 24 two-dimensional matrices of size (100, 100) will be reassembled into a second three-dimensional matrix of size (24, 100, 100).
[0084] Next, the second three-dimensional matrix is input to the time feature extraction module according to this disclosure. This module slices the input (24, 100, 100) second three-dimensional matrix spatially into 100*100 (i.e., 10,000) sequences of length 24. In other words, the (24, 100, 100) second three-dimensional matrix can be expanded node-wise into 10,000 sequences of length 24, each sequence representing 24 spatial features corresponding to the historical flow values of each node over 24 hours. These 10,000 sequences are then input into a neural network for extracting time features, such as stacked LSTMs.
[0085] According to this disclosure, in this example, the output of the neural network used to extract temporal features is also 10,000 sequences of length 24, each representing the spatiotemporal features of the historical traffic values for each node. These sequences are then fed into a fully connected layer to be mapped to the traffic prediction results for the next hour. In this document, the prediction results can be a matrix of size (100, 100).
[0086] Subsequently, the prediction results can be compared with the sample labels. For example, the error function can be calculated, and the model parameters can be adjusted based on the calculation results. Then, using the adjusted parameters, the steps of spatial feature extraction, temporal feature extraction, prediction generation, and parameter adjustment can be repeated on the next supervised learning dataset (i.e., the supervised learning set corresponding to the next sliding window) until the model converges.
[0087] The architecture and training process of the machine learning model for predicting network traffic according to this disclosure have been described in detail with reference to the accompanying drawings and specific examples. See below for further details. Figure 9 This section describes the conceptual operational process of the method 90 for predicting network traffic according to this disclosure.
[0088] The method begins at step 900.
[0089] At step 902, the historical traffic data of the target prediction region of the network is preprocessed. For example, the target prediction region of the network may be a region covering the entire network. For example, the historical traffic data of the target prediction region may be pre-collected data or data downloaded from a public website. According to this disclosure, an input three-dimensional matrix can be generated based on the historical traffic data. The first dimension of the input three-dimensional matrix is the time dimension representing a time period, the second and third dimensions are the spatial dimensions jointly representing the spatial location in the target prediction region, and the value of each element in the input three-dimensional matrix is associated with the historical traffic value of the node at the corresponding spatial location during the corresponding time period. The specific preprocessing operations on the historical traffic data are similar to those described above regarding the training process. For example, the preprocessing operations may optionally include one or more of the following, depending on actual needs: replacing outliers in the historical traffic data with predetermined values; filling missing values in the historical traffic data with predetermined values; aggregating the historical traffic data; and normalizing the historical traffic data.
[0090] It should be noted that the dimensions of the input 3D matrix of the model generated through preprocessing operations during the actual prediction process must be the same in three dimensions as the dimensions of the 3D matrix generated through preprocessing operations during model training (e.g., the first 3D matrix described above). Specifically, if a sliding window is used during model training, then the dimensions of the input 3D matrix of the model generated through preprocessing operations during the actual prediction process must be the same in three dimensions as the dimensions of the sliding window 3D matrix.
[0091] In particular, if normalization is performed in the historical traffic data preprocessing operation, the scale on which the normalization is based (i.e., the maximum and minimum traffic values of each node) can be retained so that after obtaining the predicted normalized traffic values, the same scale can be used for inverse normalization to obtain the predicted actual traffic values.
[0092] In step 904, the three-dimensional matrix obtained in step 902 can be input into the machine learning model for predicting network traffic according to the present disclosure (i.e., the machine learning model trained according to the method described above) to obtain the traffic prediction values of each node in the target prediction area of the network in the next time period.
[0093] Optionally, if normalization has been performed in the historical traffic data preprocessing operation, method 90 may further include denormalizing the data output from the machine learning model for predicting network traffic according to this disclosure using the normalization scale saved at step 902 in order to obtain the actual traffic prediction value.
[0094] The method ends at step 906.
[0095] Figure 10 An example of a network traffic prediction result using the machine learning model according to this disclosure is illustrated. This prediction result corresponds to the urban network traffic prediction scenario described above. As explained above, in this prediction scenario, there are 100 × 100, or 10,000 nodes. Figure 10 The diagram shows a comparison of the actual and predicted traffic values for four nodes numbered 47, 444, 1099, and 3493. For example, nodes could be numbered sequentially from left to right and top to bottom, starting with the first node in the first row of a (100, 100) matrix. Figure 10 In the diagram, solid lines represent the actual traffic values of these nodes in each time period, while dashed lines represent the predicted traffic values of these nodes in each time period using the machine learning model according to this disclosure. Figure 10 As shown, network traffic can be accurately predicted using the machine learning model according to this disclosure.
[0096] The solutions disclosed herein have been described in detail with reference to the accompanying drawings. The machine learning model for predicting network traffic disclosed herein can obtain traffic prediction values for multiple nodes simultaneously, resulting in high prediction efficiency. Furthermore, the machine learning model for predicting network traffic disclosed herein can simultaneously consider the temporal correlation of historical traffic among nodes in the region to be predicted and the spatial correlation of historical traffic between nodes, thereby effectively improving prediction accuracy. Moreover, the machine learning model for predicting network traffic disclosed herein is not limited by network size and is particularly capable of capturing the spatial correlation of historical traffic between nodes in large-scale networks, making it especially suitable for traffic prediction in large-scale networks.
[0097] Figure 11 An exemplary configuration of a computing device 1200 capable of implementing embodiments of the present disclosure is shown.
[0098] Computing device 1200 is an example of a hardware device capable of applying the above aspects of this disclosure. Computing device 1200 can be any machine configured to perform processing and / or computation. Computing device 1200 can be, but is not limited to, a workstation, server, desktop computer, laptop computer, tablet computer, personal data assistant (PDA), smartphone, in-vehicle computer, or a combination thereof.
[0099] like Figure 11As shown, computing device 1200 may include one or more components that can be connected to or communicate with bus 1202 via one or more interfaces. Bus 2102 may include, but is not limited to, Industry Standard Architecture (ISA) bus, Micro Channel Architecture (MCA) bus, Enhanced ISA (EISA) bus, Video Electronics Standards Association (VESA) local bus, and Peripheral Component Interconnect (PCI) bus. Computing device 1200 may include, for example, one or more processors 1204, one or more input devices 1206, and one or more output devices 1208. The one or more processors 1204 may be any type of processor and may include, but is not limited to, one or more general-purpose processors or special-purpose processors (such as dedicated processing chips). Processor 1202 may, for example, be configured to execute reference... Figure 3 The described method for training a machine learning model to predict network traffic can also be configured to perform a reference. Figure 9 The method described is for predicting network traffic. Input device 1206 can be any type of input device capable of inputting information to a computing device, and may include, but is not limited to, a mouse, keyboard, touchscreen, microphone, and / or remote controller. Output device 1208 can be any type of device capable of presenting information, and may include, but is not limited to, a monitor, speaker, video / audio output terminal, vibrator, and / or printer.
[0100] The computing device 1200 may also include or be connected to a non-transitory storage device 1214, which may be any non-transitory storage device capable of storing data, and may include, but is not limited to, disk drives, optical storage devices, solid-state storage, floppy disks, flexible disks, hard disks, magnetic tapes or any other magnetic media, compressed disks or any other optical media, cache memory and / or any other storage chip or module, and / or any other medium from which a computer may read data, instructions and / or code. The computing device 1200 may also include random access memory (RAM) 1210 and read-only memory (ROM) 1212. ROM 1212 may store executable programs, utilities, or processes in a non-volatile manner. RAM 1210 provides volatile data storage and stores instructions related to the operation of the computing device 1200. The computing device 1200 may also include a network / bus interface 1216 coupled to a data link 1218. Network / bus interface 1216 can be any kind of device or system capable of enabling communication with external devices and / or networks, and may include, but is not limited to, modems, network cards, infrared communication devices, wireless communication devices and / or chipsets (such as Bluetooth). TMEquipment, 802.11 equipment, WiFi equipment, WiMax equipment, cellular communication facilities, etc.
[0101] This disclosure can be implemented as any combination of apparatus, system, integrated circuit, and computer program on a non-transitory computer-readable medium. One or more processors can be implemented as integrated circuits (ICs), application-specific integrated circuits (ASICs), or large-scale integrated circuits (LSIs), system LSIs, super LSIs, or ultra LSI components that perform some or all of the functions described in this disclosure.
[0102] This disclosure includes the use of software, application programs, computer programs, or algorithms. Software, application programs, computer programs, or algorithms may be stored on a non-transitory computer-readable medium to cause a computer, such as one or more processors, to perform the steps described above and in the accompanying drawings. For example, one or more memories may store the software or algorithm in executable instructions, and one or more processors may be associated with executing a set of instructions of the software or algorithm to provide various functionalities according to embodiments described in this disclosure.
[0103] Software and computer programs (also referred to as programs, software applications, applications, components, or code) include machine instructions for programmable processors and can be implemented in high-level procedural languages, object-oriented programming languages, functional programming languages, logic programming languages, assembly languages, or machine languages. The term "computer-readable medium" means any computer program product, apparatus, or device used to provide machine instructions or data to a programmable data processor, such as magnetic disks, optical disks, solid-state storage devices, memories, and programmable logic devices (PLDs), including computer-readable media that receive machine instructions as computer-readable signals.
[0104] For example, computer-readable media may include dynamic random access memory (DRAM), random access memory (RAM), read-only memory (ROM), electrically erasable read-only memory (EEPROM), compact disc read-only memory (CD-ROM) or other optical disc storage devices, magnetic disk storage devices or other magnetic storage devices, or any other medium that can be used to carry or store required computer-readable program code in the form of instructions or data structures, and that can be accessed by a general-purpose or special-purpose computer or a general-purpose or special-purpose processor. As used herein, a disk or disc includes compact discs (CD), laser discs, optical discs, digital versatile discs (DVD), floppy disks, and Blu-ray discs, wherein a disk typically copies data magnetically, while a disc copies data optically using a laser. Combinations of the above are also included within the scope of computer-readable media.
[0105] The subject matter of this disclosure is provided as examples of apparatus, systems, methods, and programs for performing the features described herein. However, other features or variations are contemplated in addition to those described above. It is anticipated that the components and functions of this disclosure can be implemented using any emerging techniques that may replace any of the above-described implementations.
[0106] Furthermore, the above description provides examples and does not limit the scope, applicability, or configuration set forth in the claims. Changes may be made to the function and arrangement of the elements discussed without departing from the spirit and scope of this disclosure. Various processes or components may be appropriately omitted, substituted, or added in various embodiments. For example, features described with respect to certain embodiments may be combined in other embodiments.
[0107] Furthermore, descriptions of various embodiments of this disclosure have been given for illustrative purposes and are not intended to be exhaustive or limiting to the disclosed embodiments. The above effects are merely illustrative, and the solutions of this disclosure may also have other technical effects. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is chosen to best explain the principles of the embodiments, their practical application, or technical improvements to technologies found in the market, or to enable others skilled in the art to understand the embodiments disclosed herein.
[0108] Furthermore, in the description of this disclosure, the terms “first,” “second,” “third,” etc., are used for descriptive purposes only and should not be construed as indicating or implying relative importance or order.
[0109] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in a flowchart or block diagram may represent a portion of a module, segment, or instruction containing one or more executable instructions for implementing the specified logical function(s). In some alternative implementations, the functions marked in the blocks may occur in a different order than those marked in the figures. For example, depending on the functions involved, two consecutive blocks may actually be executed substantially in parallel, or these blocks may sometimes be executed in reverse order. It will also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented by a dedicated hardware-based system that performs the specified function or behavior or executes a combination of dedicated hardware and computer instructions.
[0110] Those skilled in the art should also understand that the various operations illustrated in sequence in the embodiments of this disclosure do not necessarily have to be performed in the illustrated order. Those skilled in the art can adjust the order of operations as needed. They can also add more operations or omit some operations as needed.
Claims
1. A method for training a machine learning model for predicting network traffic, wherein, The machine learning model includes a spatial feature extraction module, a temporal feature extraction module, and a fully connected layer; the method includes: The historical traffic data of the target prediction area of the network is preprocessed, including generating a first three-dimensional matrix based on the historical traffic data. The first dimension of the first three-dimensional matrix is the time dimension representing the time period, and the second and third dimensions are the spatial dimensions that jointly represent the spatial location in the target prediction area. The value of each element in the first three-dimensional matrix is associated with the historical traffic value of the node at the corresponding spatial location during the corresponding time period. Spatial feature extraction, wherein spatial feature extraction is performed based on multiple two-dimensional matrices obtained by slicing the first three-dimensional matrix in the time dimension; Temporal feature extraction, which is based on multiple sequences obtained by slicing the second three-dimensional matrix in the spatial dimension. The second three-dimensional matrix is obtained by concatenating multiple output two-dimensional matrices from the spatial feature extraction step according to the original temporal dimension. Temporal feature extraction includes: The second three-dimensional matrix is sliced spatially to obtain multiple sequences, where each sequence represents multiple spatial features of each node corresponding to historical flow values over multiple time periods. The multiple sequences are input into the time feature extraction module to extract time features, wherein the time feature extraction model outputs multiple feature vectors corresponding to multiple nodes; Prediction generation includes inputting the multiple feature vectors into a fully connected layer to obtain the traffic prediction value for the next time period at each node; Parameter tuning includes comparing predicted traffic values with historical traffic values and adjusting model parameters based on the comparison results; and Based on the adjusted model parameters, spatial feature extraction, temporal feature extraction, prediction generation, and parameter adjustment are re-executed until the machine learning model converges.
2. The method of claim 1, wherein spatial feature extraction further comprises: The first three-dimensional matrix is sliced along the time dimension to obtain the plurality of two-dimensional matrices corresponding to each time period, and The plurality of two-dimensional matrices are input into the spatial feature extraction module to extract spatial features, wherein the spatial feature extraction module outputs a plurality of output two-dimensional matrices with the same size as the input plurality of two-dimensional matrices.
3. The method of claim 1 or 2, wherein preprocessing the historical traffic data further comprises one or more of the following: Replace outliers in historical traffic data with predetermined values; Fill in the missing values in the historical traffic data with predefined values; Aggregate historical traffic data; and Normalize historical traffic data.
4. The method of claim 3, wherein Aggregating historical traffic data includes: Aggregate historical traffic within multiple adjacent first time periods into traffic within a second time period of equal length to the multiple first time periods, and / or Normalizing historical traffic data includes: for each node, normalizing the traffic in each time period based on the difference between the maximum and minimum traffic values of that node in each time period.
5. The method of claim 2, wherein the spatial feature extraction module comprises a densely connected convolutional neural network, DenseNet.
6. The method of claim 5, wherein the sampling rate is set to 1 during the extraction of spatial features based on DenseNet.
7. The method of claim 5, wherein during the extraction of spatial features based on DenseNet, the size of the plurality of two-dimensional matrices output between the layers of DenseNet is kept the same as the size of the plurality of two-dimensional matrices input to the spatial feature extraction module.
8. The method of claim 1 or 2, wherein the time feature extraction module comprises multiple stacked Long Short-Term Memory (LSTM) networks.
9. The method of claim 8, wherein a subset of neurons from all neurons is used for temporal feature extraction.
10. The method of claim 1 or 2, wherein The parameter adjustment further includes calculating an error function based on the traffic prediction and historical traffic values for each node, and adjusting the model parameters based on the calculation results of the error function.
11. The method as claimed in claim 1 or 2, wherein: The preprocessing of historical traffic data further includes: dividing the first three-dimensional matrix into multiple supervised learning datasets. Each supervised learning dataset includes a sliding window three-dimensional matrix and sample labels. The time dimension of the sliding window three-dimensional matrix has the same length as the sliding window, and the spatial dimension has the same length as the first three-dimensional matrix. The sample labels represent the historical traffic values of each node to be predicted for the next time period. Spatial feature extraction, temporal feature extraction, predicted value generation, and parameter adjustment are performed by advancing the sliding window time-by-time and targeting the three-dimensional matrix of each sliding window.
12. A method for predicting network traffic, comprising: The historical traffic data of the target prediction area of the network is preprocessed, including generating an input three-dimensional matrix based on the historical traffic data. The first dimension of the input three-dimensional matrix is the time dimension representing the time period, and the second and third dimensions together represent the spatial dimension of the spatial location in the target prediction area. The value of each element in the input three-dimensional matrix is associated with the historical traffic value of the node at the corresponding spatial location during the corresponding time period. The input three-dimensional matrix is fed into a machine learning model trained according to any one of claims 1-11, thereby obtaining the traffic prediction value for the next time period at each node.
13. An apparatus for training a machine learning model for predicting network traffic, comprising: A memory that stores instructions; as well as A processor is configured to execute instructions stored in the memory to perform the method as described in any one of claims 1 to 11.
14. An apparatus for predicting network traffic, comprising: A memory that stores instructions; as well as The processor is configured to execute instructions stored in the memory to perform the method as described in claim 12.
15. A computer-readable storage medium comprising computer-executable instructions, which, when executed by one or more processors, cause the one or more processors to perform the method of any one of claims 1 to 11, or the one or more processors to perform the method of claim 12.