Wireless traffic prediction method and system based on mGPT
By applying the self-supervised learning and context learning mechanism of the GPT model in wireless traffic prediction, the problem of traditional models being unable to capture complex features and lack of adaptability is solved, and more accurate and efficient wireless traffic prediction is achieved.
Patent Information
- Application Number
- CN202510611490.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-13
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-05-13
AI Technical Summary
Traditional wireless traffic prediction models cannot effectively capture complex features in data, lack adaptability and accuracy, and make it difficult to make accurate predictions in dynamically changing environments.
The wireless traffic prediction method based on GPT is adopted, combined with the self-supervised learning and context learning mechanism of the GPT model, and the sliding window mechanism and periodic update mechanism are used to generate accurate prediction results for future traffic.
It significantly improves the accuracy and computing efficiency of wireless traffic prediction, enhances the adaptability of the model in a dynamic environment, reduces the need for large amounts of labeled data, and reduces the cost of data preparation.
Smart Images

Figure CN120151867A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical fields of communication networks and artificial intelligence, and relates to a wireless traffic data prediction method and system based on GPT (Generative Pretrained Transformer), which can capture complex features in data and is used for network management and planning in communication systems. Background Art
[0002] Traditional wireless traffic prediction models cannot capture complex features in wireless traffic data. In recent years, large language models have shown excellent capabilities in machine vision and text language. However, few studies have applied large language models to wireless traffic data prediction.
[0003] Wireless traffic prediction based on large language models can not only make full use of their advantages in processing large-scale data and capturing time series patterns, but also improve the adaptability of the model in a dynamically changing environment through context learning.
[0004] As a powerful generative pre-training model, GPT has been proven to have extremely strong generalization ability in a variety of tasks due to its excellent performance in the field of natural language processing. Applying it to wireless traffic prediction can automatically identify and capture potential patterns of traffic changes through its context understanding ability of historical data, and make accurate predictions.
[0005] However, in traditional models, only specific data can be used to train the model for specific tasks, lacking adaptability. Therefore, there is an urgent need for a model that can accurately predict data with a slight modification without having to retrain the model from scratch. Summary of the Invention
[0006] Aiming at the deficiencies of the prior art, the present invention provides a wireless traffic prediction method based on mGPT; which is used to solve the problem of limited non-linear modeling ability of traditional models.
[0007] The present invention combines the self-supervised learning ability of the GPT model and an efficient context learning mechanism to predict the service traffic in a wireless communication network. By introducing context information into wireless traffic data, the GPT model can automatically learn and generate future change trends of wireless traffic, significantly improving the prediction accuracy and calculation efficiency.
[0008] The present invention provides a method for predicting wireless traffic data based on an improved GPT. Based on historical data of wireless traffic, a sliding window mechanism is used to divide the data into a training set, a validation set, and a test set, ensuring that the temporal sequence and context information of the data are fully utilized. The GPT model is used for training, combining historical data and context information to generate prediction results for future traffic. A periodic update mechanism is adopted to fine-tune the model according to new historical data and prediction results, ensuring that the model adapts to changes in traffic patterns. The prediction results are evaluated, and metrics such as Mean-Square Error (MSE) are used to evaluate the prediction accuracy; the model is updated based on the evaluation results to continuously optimize the prediction results.
[0009] The technical solution of the present invention is as follows: A method for predicting wireless traffic based on mGPT, including: Preprocessing the wireless traffic data to be predicted and inputting it into a trained wireless traffic prediction model to achieve wireless traffic data prediction; The wireless traffic prediction model includes multiple layers of Block blocks, and each Block block includes a self-attention mechanism, a feed-forward network, layer normalization, and a residual connection; The self-attention mechanism includes a linear transformation layer that generates QKV, performs dimension reshaping and transposition, calculates attention through scaled dot product, and then performs projection; when the self-attention mechanism processes sequence data in the wireless traffic prediction model, different weights are assigned to each element in the sequence data; in the feed-forward network, the linear layer expands the dimension, and after activation by GELU and a projection layer, the dimension is restored; layer normalization defines scaling and offset parameters and uses a normalization operation; the residual connection directly adds the input to the layer output, and in the Block class, the attention after layer normalization and the results of the feed-forward network are respectively connected.
[0010] Preferably according to the present invention, the following operations are performed before training the wireless traffic prediction model, including: A. Obtain historical data of the base station, divide it into a training data set and a test data set, and on the training data set and the test data set, according to the sliding window mechanism, select the sliding window size , and use the sliding window for processing to generate a training sample data set and a test sample data set; B. For the training sample data set, obtain the maximum and minimum values of the traffic in the training sample data set, and use Min-Max normalization for the training sample data set, the validation sample data, and the test sample data, and normalize them to the range of [0, 1].
[0011] Preferably according to the present invention, the training process of the wireless traffic prediction model is as follows: (1) In the initial stage, the input tensor, that is, the training sample data set, passes through a linear transformation Each feature is transformed into a higher-dimensional embedding space to obtain ; (2) Add positional encoding ; Specifically, it refers to: ; (3) They are successively fed into multiple Block blocks, and the processing flow in the Block blocks is as follows: (I); (II); (III); (IV); (V); In formula (I), is the tensor after layer normalization, which is performed on the input tensor ; In formula (II), is the tensor after the self-attention mechanism; Formula (III) is the residual connection, is the input tensor and the tensor obtained by adding the tensor processed by the self-attention mechanism; Formula (IV) is the multi-layer perceptron and layer normalization, is obtained by passing it through layer normalization and then inputting it into a multi-layer perceptron (MLP); Formula (V) is another residual connection, which adds the tensor processed by the MLP and the tensor processed by the self-attention mechanism and residual connection; Repeat steps (1) to (3) until the following training end conditions are met to obtain a trained wireless traffic prediction model; the training end conditions are ① or ②: ①: The parameters of the wireless traffic prediction model converge, that is, the update change of the parameters is less than the set threshold; ②: The number of updates of the wireless traffic prediction model reaches the preset number of times.
[0012] According to the preference of the present invention, the wireless traffic prediction model updates its parameters according to historical data, and the specific steps include: A. Select an optimization algorithm; update the parameters of the wireless traffic prediction model based on the gradient descent strategy; B. Select the number of samples according to the batch size from the training dataset and calculate the gradient through backpropagation; C. Update the parameters of the wireless traffic prediction model according to the gradient information of the current batch of samples; adjust the weight parameters through an optimization algorithm (AdamW); D. Repeat steps B and C until the training end condition is met.
[0013] Further preferably, in step C, updating the parameters of the wireless traffic prediction model according to the gradient information of the current batch of samples; adjusting the weight parameters through an optimization algorithm (AdamW); includes: Calculate the first moment: by calculating the exponentially weighted moving average of the gradient; the formula is: (VI); Where, is the estimate of the first moment, is the gradient of the current batch, is the momentum hyperparameter; Calculate the second moment: the formula is: (VII); Where, is the estimate of the second moment, is another hyperparameter; Update the parameters: use the first moment and the second moment to update the parameters of the wireless traffic prediction model ; the formula is: (VIII); Where, is the learning rate, is a constant to prevent division by zero errors.
[0014] Preferably according to the present invention, after realizing the wireless traffic data prediction, the following operations are performed: Perform the inverse operation of normalization on the predicted value to obtain the true scale of the predicted value, and evaluate the prediction performance according to the corresponding evaluation index; after the evaluation is completed, store the newly predicted data of the wireless traffic prediction model into the historical database.
[0015] A computer device, including a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, the steps of a wireless traffic data prediction method based on mGPT are implemented.
[0016] A computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of a wireless traffic data prediction method based on mGPT are implemented.
[0017] A wireless traffic prediction system based on mGPT, including: A preprocessing module, configured to: preprocess the wireless traffic data to be predicted; The wireless traffic data prediction module is configured to: preprocess the wireless traffic data to be predicted and input it into a trained wireless traffic prediction model to implement wireless traffic data prediction; the wireless traffic prediction model includes multiple layers of Block blocks, and each Block block includes a self-attention mechanism, a feed-forward network, layer normalization, and a residual connection; the self-attention mechanism includes a linear transformation layer, which generates QKV, performs dimension reshaping and transposition, calculates attention through scaled dot product, and then performs projection; when the self-attention mechanism processes sequence data in the wireless traffic prediction model, different weights are assigned to each element in the sequence data; in the feed-forward network, the linear layer expands the dimension, and after being activated by GELU and restored by the projection layer; layer normalization defines scaling and offset parameters and uses normalization operations; the residual connection directly adds the input to the layer output, and in the Block class, the attention after layer normalization and the results of the feed-forward network are connected respectively.
[0018] The beneficial effects of the present invention are as follows: 1. In the embodiment of the present invention, by using the pre-trained GPT for wireless traffic prediction, the wireless traffic data of future time steps can be accurately predicted, and it has a certain generalization ability.
[0019] 2. Using the pre-trained GPT model for wireless traffic prediction: By using the pre-trained GPT model for wireless traffic prediction, the present invention can make full use of the deep feature learning ability of the large language model (Large Language Model) to accurately capture the potential laws in time series data and predict the wireless traffic data of future time steps.
[0020] 3. The present invention replaces the loss function in the original text generation task with the mean square error (MSE), and adjusts the output layer to a regression linear layer, optimizing the accuracy and stability of wireless traffic prediction. The MSE loss function is applicable to regression problems and can effectively reduce the error between the predicted value and the actual value, improving the prediction accuracy.
[0021] 4. Utilizing the few-shot learning ability of the GPT model, the present invention effectively reduces the need for a large amount of labeled data in wireless traffic prediction and reduces the cost of data preparation.
[0022] 5. By adopting the self-attention mechanism and deep non-linear transformation of the GPT model, the present invention can be generalized in various wireless traffic data environments, adapt to different prediction scenarios, such as short-term and long-term traffic predictions, and improve the adaptability and robustness of the model. Description of the Drawings
[0023] Figure 1 It is a schematic diagram of the architecture of the wireless traffic prediction model; Figure 2It is a comparison schematic diagram of the predicted value and the true value of a certain area. Detailed implementation mode
[0024] The present invention will be further defined below in conjunction with the accompanying drawings of the specification and embodiments, but is not limited thereto.
[0025] Embodiment 1 A wireless traffic prediction method based on mGPT (Modified Generative Pretrained Transformer) includes: Preprocess the wireless traffic data to be predicted and input it into the trained wireless traffic prediction model to realize the prediction of wireless traffic data; to improve the accuracy of wireless traffic data prediction.
[0026] As Figure 1 shown, the wireless traffic prediction model includes multiple layers of Block blocks, and each Block block includes a self-attention mechanism, a feed-forward network, layer normalization, and a residual connection; The self-attention mechanism includes a linear transformation layer, which generates QKV, performs dimension reshaping and transposition, calculates attention through scaled dot product, and then performs projection; when the wireless traffic prediction model processes sequence data, the self-attention mechanism assigns different weights to each element in the sequence data; enabling the wireless traffic prediction model to better focus on the importance of different parts in the input sequence. In the feed-forward network, the linear layer expands the dimension, and after being activated by GELU and projected by the projection layer, the dimension is restored; the feed-forward network is usually located after the self-attention mechanism, and further non-linear transformation is performed on the output of the self-attention mechanism to increase the expression ability and complexity of the model, and help the model learn more complex feature representations. Layer normalization defines scaling and offset parameters and uses normalization operations; layer normalization helps to stabilize the training process, accelerate the convergence speed, and reduce internal covariate shift. It normalizes the input feature dimension to a distribution with a mean of 0 and a variance of 1, making the training of the network more stable and efficient. The residual connection directly adds the input to the layer output, and in the Block class, the attention and feed-forward network results after layer normalization are connected respectively. The residual connection allows the gradient to flow more easily in the network, alleviates the problem of gradient disappearance, and helps to train deeper networks. It directly adds the input to the output of the network layer, enabling the network to learn the residual between the input and the output, thus making it easier to optimize and train deep networks.
[0027] Embodiment 2 A wireless traffic prediction method based on mGPT according to Embodiment 1, wherein the difference lies in: Before training the wireless traffic prediction model, the following operations are performed, including: A. Obtain the historical data of the base station, divide it into a training data set and a test data set according to a certain ratio, and on the training data set and the test data set, select the sliding window size according to the sliding window mechanism , and use the sliding window for processing to generate a training sample data set and a test sample data set; the data set used in the present invention comes from a public data set widely used in the field of cellular services, providing real telecom data from the Telecom Italia Big Data Challenge. The data time span is from 00:00 on November 1, 2013 to 23:50 on January 1, 2014. In order to construct a data set suitable for training, the present invention adopts a sliding window technology for the data of 88 districts in Milan. Each sliding window contains n (determined according to the periodicity of the data and the time span required for the prediction task) consecutive data of time units and slides with a step size of m time units. This configuration aims to ensure the continuity and comprehensive coverage of the data. The data processed by the sliding window is divided into a training set, a validation set and a test set according to the ratio of 7:2:1.
[0028] B. For the training sample data set, obtain the maximum and minimum values of the traffic in the training sample data set, and use Min-Max normalization for the training sample data set, the validation sample data and the test sample data, and normalize them to the range of [0, 1]. The formula is as follows:
[0029] ; where, is the normalized data, is the minimum value in the training set, X is the data at the current moment, is the maximum value of the training set.
[0030] The training process of the wireless traffic prediction model is as follows: (1) In the initial stage, the input tensor, that is, the training sample data set, passes through a linear transformation , and each feature is transformed into a higher-dimensional embedding space to obtain ; The linear transformation is defined as: ; where, is the weight matrix of the linear transformation; represents the input tensor, usually the training sample data set. is the input tensor after passing through the linear transformation and adding a bias , that is, the embedding representation. is the bias vector of the linear transformation. TE is token Embedding.
[0031] (2) Add positional encoding ; Specifically: ; When processing sequential data (such as time series), since the order of the input usually has an important impact on the result, positional encoding is used to add positional information to each position in the sequence. is the result of adding the embedded representation and the positional encoding. By adding the positional encoding to the embedded representation, the input not only contains the embedded information of the original features but also their positions in the sequence, which helps the model consider the importance of positions when processing sequential data.
[0032] (3) are sequentially fed into multiple Block blocks, which form the core of the wireless traffic prediction model. Each Block block processes the input tensor through a series of steps designed to refine and transform the data to deepen understanding and extract patterns;
[0033] The processing flow in the Block block is as follows: (I); (II); (III); (IV); (V); In formula (I), is the tensor after layer normalization, which is performed on the input tensor . Layer normalization calculates the mean and variance on the feature dimensions of each sample and normalizes them to a distribution with a mean of 0 and a variance of 1. It is different from batch normalization in that layer normalization is performed on the feature dimensions of each sample rather than on the batch dimension.
[0034] In formula (II), is the tensor after the self-attention mechanism; SA is the self-attention mechanism that allows the model to assign different attention weights to each element in the input sequence, enabling the model to focus on the importance of different parts of the input sequence.
[0035] Specifically: First, queries, keys, and values are generated. For the input through different weight matrices , , Q, K, and V are obtained. The attention scores are calculated through the scaled dot-product attention formula, where is the dimension of the key vector, which is used to scale the attention scores.
[0036] ; Finally, a weighted sum of V is calculated based on the calculated attention scores to obtain .
[0037] Equation (Ⅲ) is a residual connection, is the input tensor added to the tensor processed by the self-attention mechanism; it allows gradients to propagate more easily in the network, helps solve the vanishing gradient problem in deep networks, and makes the network easier to train.
[0038] Equation (Ⅳ) is a multi-layer perceptron (MLP) and layer normalization, is obtained by passing it through layer normalization and then inputting it into a multi-layer perceptron (MLP); the layer normalization is the same as in Equation (Ⅰ). In the multi-layer perceptron, first, a linear transformation is performed, and the input features are transformed through one or more linear layers. Then, non-linear activation is used, and a non-linear activation function is used to introduce non-linearity, enabling the model to learn more complex functional relationships, which helps capture non-linear patterns in wireless traffic data. The traffic in different time periods may show non-linear change trends, and the MLP can better fit this complex relationship.
[0039] Equation (Ⅴ) is another residual connection, which adds the tensor processed by the MLP to the tensor processed by the self-attention mechanism and with residual connection; Since the model consists of multiple blocks, i represents the processing in the i-th block.
[0040] The superscript LN refers to the result after layer normalization, SA refers to the result after the self-attention mechanism, and MLP refers to the tensor processed by the multi-layer perceptron.
[0041] Repeat steps (1) to (3) until the following training end conditions are met to obtain a trained wireless traffic prediction model; the training end conditions are ① or ②: ①: The parameters of the wireless traffic prediction model converge, that is, the update change of the parameters is less than the set threshold; ②: The number of updates of the wireless traffic prediction model reaches the preset number of times.
[0042] The wireless traffic prediction model updates its parameters based on historical data. The specific steps include: A. Select an optimization algorithm; in this step, the AdamW optimizer is selected, and the parameters of the wireless traffic prediction model are updated based on the gradient descent strategy; B. Select an appropriate number of samples (usually 32) according to the batch size from the training dataset, and calculate the gradients through backpropagation; Forward propagation: First, input the selected batch of data into the wireless traffic prediction model to calculate the predicted values. For wireless traffic prediction, the input is the data processed by the sliding window and normalized. The model performs a series of calculations based on the input and finally outputs the predicted wireless traffic values.
[0043] Calculate the loss: Compare the predicted values with the true values and calculate the loss function. The formula is , where is the true wireless traffic value, is the predicted wireless traffic value, and n is the batch size.
[0044] Backpropagation: Use the chain rule to calculate the gradients of the loss function with respect to each model parameter. Starting from the loss function, calculate the gradients backward according to the computational graph (formed by the forward propagation process of the model). For each operation, calculate the gradients according to the derivative rules of that operation.
[0045] C. Update the parameters of the wireless traffic prediction model according to the gradient information of the current batch of samples; adjust the weight parameters through an optimization algorithm (AdamW); use hyperparameters such as the learning rate and momentum to ensure the stability and convergence of the parameter update of the wireless traffic prediction model.
[0046] D. Repeat steps B and C until the training end condition is met.
[0047] In step C, update the parameters of the wireless traffic prediction model according to the gradient information of the current batch of samples; adjust the weight parameters through an optimization algorithm (AdamW); including: AdamW combines the ideas of momentum and adaptive learning rate and is especially suitable for deep learning models. AdamW can adaptively adjust the learning rate according to the historical gradients of each parameter during training, thus accelerating convergence and reducing oscillations.
[0048] The update formula of AdamW is roughly as follows: Calculate the first moment (momentum): Calculate the exponentially weighted moving average of the gradients; to reduce the fluctuations of the gradients. The formula is:
[0049] ; where is the estimate of the first moment, is the gradient of the current batch, is the momentum hyperparameter; usually close to 1 (e.g., 0.9).
[0050] Calculate the second moment (variance of the gradient): Adam also calculates the variance of the gradient, which is used to scale the learning rates of different parameters so that parameters with smaller gradients can also be updated sufficiently. The formula is:
[0051] ; where is the estimate of the second moment, is another hyperparameter; it is usually set to be close to 1 (e.g., 0.999).
[0052] Update the parameters: Use the first moment and the second moment to update the parameters of the wireless traffic prediction model ; AdamW uses weight decay (L2 regularization), that is, a constraint on the model weights is added during the update process to avoid overfitting of the model. The formula is:
[0053]
[0054] where is the learning rate, is a constant to prevent division by zero error. Usually it is 1e-8.
[0055] Since the Adam optimization algorithm converges quickly, this implementation uses this optimization algorithm.
[0056] The learning rate determines the step size of parameter update. A larger learning rate will make the parameter update faster, but may lead to non-convergence or oscillation near the minimum value; a smaller learning rate will make the update slow and require more iteration times.
[0057] Weight decay is a feature in AdamW, which is used to control the complexity of the model and prevent overfitting. Weight decay will penalize larger weights during the optimization process of the model, thus restricting the parameters from being too large.
[0058] During training, by adjusting hyperparameters such as the learning rate and momentum, ensure that the model converges stably during the training process.
[0059] After implementing the wireless traffic data prediction, perform the following operations: Perform the inverse operation of normalization on the predicted value to obtain the true scale of the predicted value, and evaluate the prediction performance according to the corresponding evaluation metrics; after the evaluation is completed, store the newly predicted data of the wireless traffic prediction model into the historical database.
[0060] The inverse operation of normalization is the process of inverse normalization: , where is the predicted value.
[0061] The evaluation metrics include MSE, MAE, and MAPE; including: ; ; ; Wherein, m is the number of predicted data, is the true value, is the predicted value.
[0062] For the wireless traffic prediction of different regions, due to the pre-training of the GPT model, only very little data is required to perform the prediction for a specific region.
[0063] Figure 2 is the prediction using the trained model on the test set divided in the BICOCCA region (one of the 88 regions of the Milan city dataset). The abscissa is the time step of the predicted data, and the ordinate is the wireless traffic data. From Figure 2 results, it can be seen that mGPT can accurately predict the traffic data of future time steps and can also make good predictions when dealing with emergencies. Therefore, the wireless traffic prediction scheme proposed by the present invention can effectively improve the prediction performance.
[0064] Embodiment 3 A computer device, comprising a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the steps of a wireless traffic prediction method based on mGPT described in Embodiment 1 or 2 are implemented.
[0065] Embodiment 4 A computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of a wireless traffic prediction method based on mGPT described in Embodiment 1 or 2 are implemented.
[0066] Embodiment 5 A wireless traffic prediction system based on mGPT, comprising: A preprocessing module, configured to: preprocess the wireless traffic data to be predicted; The wireless traffic data prediction module is configured to: preprocess the wireless traffic data to be predicted and input it into the trained wireless traffic prediction model to achieve wireless traffic data prediction; the wireless traffic prediction model includes multiple layers of Block blocks, and each Block block includes a self-attention mechanism, a feed-forward network, layer normalization, and a residual connection; the self-attention mechanism includes a linear transformation layer, which generates QKV, performs dimension reshaping and transposition, calculates attention through scaled dot product, and then performs projection; when the self-attention mechanism processes sequence data in the wireless traffic prediction model, it assigns different weights to each element in the sequence data; in the feed-forward network, the linear layer expands the dimension, and after being activated by GELU and restored to the dimension by the projection layer; layer normalization defines scaling and offset parameters and uses normalization operations; the residual connection directly adds the input to the layer output, and in the Block class, the attention after layer normalization and the results of the feed-forward network are connected respectively.
Claims
1. A wireless traffic prediction method based on mGPT, characterized in that: include: The wireless traffic data to be predicted is pre-processed and input into the trained wireless traffic prediction model to realize wireless traffic data prediction; The wireless traffic prediction model consists of multiple layers of blocks, each of which includes a self-attention mechanism, a feedforward network, layer normalization, and a residual connection. The self-attention mechanism includes a linear transformation layer, which generates QKV, reshapes and transposes the dimensions, calculates the attention through the scaled dot product, and then projects. When the wireless traffic prediction model processes sequence data, the self-attention mechanism assigns different weights to each element in the sequence data. In the feedforward network, the linear layer expands the dimension, and the dimension is restored through GELU activation and projection layer. Layer normalization defines scaling and offset parameters and uses normalization operations. The residual connection adds the input directly to the layer output, and connects the layer normalized attention and feedforward network results in the Block class.
2. According to the mGPT-based wireless traffic prediction method of claim 1, it is characterized in that: The following operations are performed before training the wireless traffic prediction model, including: A. Obtain the historical data of the base station, divide it into a training data set and a test data set, and select the sliding window size based on the sliding window mechanism on the training data set and the test data set. , use sliding window for processing to generate training sample data set and test sample data set; B. For the training sample data set, obtain the maximum and minimum values of the flow in the training sample data set, use Min-Max normalization on the validation sample data and the test sample data of the training sample data set, and normalize them to the range of [0, 1].
3. The wireless traffic prediction method based on mGPT according to claim 1 is characterized in that: The training process of the wireless traffic prediction model is as follows: (1) In the initial stage, the input tensor, i.e. the training sample dataset, is transformed linearly , transform each feature into a higher-dimensional embedding space and obtain ; (2) Adding position coding ; Specifically refers to: ; (3) It is sent to multiple layers of Block in sequence. The processing flow in the Block is as follows: (I); (II); (III); (IV); (V); In formula (I), is the normalized tensor of the layer, for the input tensor conduct; In formula (II), is the tensor after the self-attention mechanism; Formula (III) is the residual connection, is the input tensor And the tensor processed by the self-attention mechanism Added together; Formula (IV) is a multi-layer perceptron with layer normalization. yes After layer normalization, it is then input into a multi-layer perceptron (MLP); Formula (V) is another residual connection, which converts the tensor processed by MLP into The tensor processed by the self-attention mechanism and residual connection Addition; Repeat steps (1) to (3) until the following training end conditions are met to obtain a trained wireless traffic prediction model; the training end conditions are ① or ②: ①: The parameters of the wireless traffic prediction model converge, that is, the update change of the parameters is less than the set threshold; ②: The wireless traffic prediction model is updated a preset number of times.
4. The method for wireless traffic prediction based on mGPT according to claim 1, characterized in that: The wireless traffic prediction model updates parameters based on historical data. The specific steps include: A. Select optimization algorithm; update the wireless traffic prediction model parameters based on gradient descent strategy; B. Select the number of samples from the training dataset based on the batch size and calculate the gradient through backpropagation; C. Update the parameters of the wireless traffic prediction model according to the gradient information of the current batch of samples; adjust the weight parameters through the optimization algorithm (AdamW); D. Repeat steps B and C until the training end condition is met.
5. The method for wireless traffic prediction based on mGPT according to claim 1, characterized in that: In step C, the parameters of the wireless traffic prediction model are updated according to the gradient information of the current batch of samples; the weight parameters are adjusted by the optimization algorithm; including: Calculate the first-order moment: by calculating the exponentially weighted moving average of the gradient; the formula is: (WE); in, is an estimate of the first moment, is the gradient of the current batch, is the momentum hyperparameter; Calculate the second moment: The formula is: (VII); in, is an estimate of the second-order moment, is another hyperparameter; Update parameters: Use the first-order moment and second-order moment to update the parameters of the wireless traffic prediction model ; The formula is: (VIII); in, is the learning rate, is a constant to prevent division by zero errors.
6. The method for wireless traffic prediction based on mGPT according to claim 1, characterized in that: After wireless traffic data prediction is implemented, perform the following operations: Perform the inverse operation of normalization on the predicted value to obtain the true scale of the predicted value, and evaluate the prediction performance according to the corresponding evaluation index; after the evaluation is completed, store the newly predicted data of the wireless traffic prediction model into the historical database.
7. A wireless traffic prediction system based on mGPT, characterized in that: include: The preprocessing module is configured to: preprocess the wireless traffic data to be predicted; The wireless traffic data prediction module is configured to: pre-process the wireless traffic data to be predicted and input it into the trained wireless traffic prediction model to realize wireless traffic data prediction; The wireless traffic prediction model includes multiple layers of Block blocks, each of which includes a self-attention mechanism, a feedforward network, layer normalization, and a residual connection. The self-attention mechanism includes a linear transformation layer, which generates QKV, reshapes and transposes the dimensions, calculates the attention by scaling the dot product, and then projects. When the wireless traffic prediction model processes sequence data, the self-attention mechanism assigns different weights to each element in the sequence data. In the feedforward network, the linear layer expands the dimension, and the dimension is restored by GELU activation and projection layer. Layer normalization defines scaling and offset parameters and uses normalization operations. The residual connection adds the input directly to the layer output, and connects the layer normalized attention and feedforward network results in the Block class.
Citation Information
Patent Citations
Wireless service traffic prediction method and device based on self-attention convolutional network, and medium
CN112910711A
Time sequence prediction method and device, equipment and storage medium
CN116993185A
Network traffic prediction method and device
CN118158105A
Network traffic generation method and device, electronic equipment and storage medium
CN118984281A
Parameter updating method and device based on low-rank decomposition, equipment and storage medium
CN119272897A