Network traffic prediction method and device
By combining segmented processing of historical network traffic data with a multi-head attention mechanism, the performance bottleneck of network traffic prediction in existing technologies when inputting extremely long sequences is solved, achieving more efficient and accurate traffic prediction.
Patent Information
- Application Number
- CN202510864074.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-25
- Publication Date
- 2025-09-09
AI Technical Summary
Existing network traffic prediction technologies face performance bottlenecks when processing extremely long sequence inputs, making it difficult to effectively utilize long-term or even global sequence information, resulting in high computational complexity and poor prediction performance.
The historical network traffic data is processed in segments, and the query weight matrix and multi-head attention mechanism are combined to retrieve relevant historical information and target memory parameters. The multi-head attention mechanism of group query is used for prediction, and the parameters of the long-term information retrieval module are dynamically updated to improve the model's prediction rationality and stability under extremely long input sequences.
It effectively reduces computational complexity, improves the timeliness and accuracy of prediction results, and enhances the model's application stability and prediction rationality under extremely long input sequences.
Smart Images

Figure CN120614261A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and in particular to a network traffic prediction method and device. Background Art
[0002] With the rapid development of information technologies such as mobile communications and the internet in recent years, network traffic forecasting has become increasingly important and a key issue in network management and operations. Accurately forecasting network traffic can help network administrators rationally plan and schedule network resources, improving network performance and reliability.
[0003] Existing technical solutions can achieve relatively good practical results, but they also have certain technical drawbacks. As network traffic becomes increasingly complex in the future, these solutions are mostly limited to simple short-term queue storage of pre-order input data features, lacking the ability to leverage long-term or even global sequence information. This leads to performance bottlenecks when processing extremely long sequence inputs, and the time series processing models they use will struggle to support such inputs. Summary of the Invention
[0004] In view of the above problems, embodiments of the present application provide a network traffic prediction method and apparatus that overcome the above problems or at least partially solve the above problems.
[0005] In a first aspect, an embodiment of the present application provides a network traffic prediction method, the method comprising:
[0006] Inputting historical network traffic data into a traffic prediction model, segmenting the historical network traffic data to obtain multiple segments of historical network traffic data;
[0007] Retrieving historical information related to first historical network traffic data according to a query weight matrix of the traffic prediction model, and retrieving a target memory parameter related to the first historical network traffic data from a plurality of preset memory parameters, wherein the first historical network traffic data is a last segment of the plurality of segments of historical network traffic data;
[0008] The first historical network traffic data, the historical information and the target memory parameter are used to predict target network traffic data at a future moment after the current moment.
[0009] Optionally, the retrieving historical information related to the first historical network traffic data according to the query weight matrix of the traffic prediction model includes:
[0010] Calculating a query vector corresponding to a first piece of historical network traffic data among the multiple pieces of historical network traffic data according to a query weight matrix of the traffic prediction model;
[0011] Based on the query vector corresponding to the first historical network traffic data, historical information related to the first historical network traffic data is retrieved.
[0012] Optionally, the predicting target network traffic data at a future moment after the current moment using the first historical network traffic data, the historical information, and the target memory parameter includes:
[0013] splicing the first historical network traffic data, the historical information, and the target memory parameter to obtain spliced data;
[0014] Based on the spliced data, target network traffic data at a future time after the current time is predicted.
[0015] Optionally, predicting target network traffic data at a future time after the current time based on the spliced data includes:
[0016] A multi-head attention mechanism for grouped queries is used to calculate the pre-output traffic data corresponding to the spliced data. In the multi-head attention mechanism, queries within the same group share the same key and value during the calculation process.
[0017] Calculating a target network parameter expression corresponding to the first historical network traffic data according to a target network parameter expression corresponding to second historical network traffic data, where the second historical network traffic data is historical network traffic data of a previous period of the first historical network traffic data;
[0018] According to the target network parameter expression corresponding to the first historical network traffic data and the pre-output traffic data, the target network traffic data at a future moment after the current moment is predicted.
[0019] Optionally, calculating the target network parameter expression corresponding to the first historical network traffic data according to the target network parameter expression corresponding to the second historical network traffic data includes:
[0020] Calculating a surprising update gradient of the traffic prediction model according to the target network parameter expression corresponding to the first historical network traffic data and the second historical network traffic data;
[0021] Calculating the cumulative amount of surprise corresponding to the first historical network traffic data according to the cumulative amount of surprise corresponding to the second historical network traffic data and the update gradient of the surprise of the traffic prediction model;
[0022] The target network parameter expression corresponding to the first historical network traffic data is calculated according to the target network parameter expression corresponding to the second historical network traffic data and the cumulative amount of surprise corresponding to the first historical network traffic data.
[0023] Optionally, the target network parameter expression corresponding to the first historical network traffic data is calculated by the following formula:
[0024] M t =(1-α t )M t-1 +μ t
[0025]
[0026] Among them, M t An expression representing a target network parameter corresponding to the first historical network traffic data;
[0027] M t-1 An expression representing a target network parameter corresponding to the second historical network traffic data;
[0028] α t represents the weight attenuation coefficient;
[0029] μ t represents the cumulative amount of surprise corresponding to the first historical network traffic data;
[0030] μ t-1 represents the cumulative amount of surprise corresponding to the second historical network traffic data;
[0031] η t represents the momentum coefficient;
[0032] θ t represents the learning rate;
[0033] The updated gradient representing the surprise of the traffic prediction model;
[0034] S t Indicates the first historical network traffic data;
[0035] W K and W V denote the key weight matrix and value weight matrix respectively;
[0036] Represents the gradient symbol;
[0037] Represents the square of the L2 norm.
[0038] Optionally, predicting target network traffic data at a future moment after the current moment based on a target network parameter expression corresponding to the first historical network traffic data and the pre-output traffic data includes:
[0039] Calculating a traffic prediction result based on a target network parameter expression corresponding to the first historical network traffic data and the pre-output traffic data;
[0040] The pre-output traffic data and the traffic prediction result are weightedly calculated to obtain the target network traffic data.
[0041] Optionally, the target network traffic data is calculated using the following formula:
[0042] O t =α O y t +(1-α O )M t (y t )
[0043] Among them, O t Indicates target network traffic data;
[0044] α O represents the prediction weight coefficient;
[0045] y t Indicates pre-output flow data;
[0046] M t In a second aspect, an embodiment of the present application further provides a network traffic prediction device, the device comprising:
[0047] A segmentation module is used to input historical network traffic data into the traffic prediction model, and perform segmentation processing on the historical network traffic data to obtain multiple segments of historical network traffic data;
[0048] a retrieval module, configured to retrieve historical information related to first historical network traffic data according to a query weight matrix of the traffic prediction model, and retrieve a target memory parameter related to the first historical network traffic data from a plurality of preset memory parameters, wherein the first historical network traffic data is a last segment of the plurality of segments of historical network traffic data;
[0049] The prediction module is used to predict the target network traffic data at a future moment after the current moment using the first historical network traffic data, the historical information and the target memory parameter.
[0050] In a third aspect, an embodiment of the present application further provides an electronic device, including a memory, a transceiver, and a processor:
[0051] A memory for storing a computer program; a transceiver for transmitting and receiving data under the control of a processor; and a processor for reading the computer program in the memory and executing the method as described in the first aspect above.
[0052] In a fourth aspect, an embodiment of the present application further provides a processor-readable storage medium, which stores a computer program, and the computer program is used to enable the processor to execute the method described in the first aspect above.
[0053] In the above embodiment of the present application, historical network traffic data is input into the traffic prediction model, and the historical network traffic data is segmented to obtain multiple segments of historical network traffic data. The block operation is performed on the long time series input (i.e., historical network traffic data). Compared with the existing scheme that uses the complete sequence for traffic prediction, it pays more attention to local information, can effectively reduce the computational complexity, and improve the efficiency of reasoning calculation. In addition, according to the query weight matrix of the traffic prediction model, historical information related to the first historical network traffic data is retrieved, and the target memory parameter related to the first historical network traffic data in multiple preset memory parameters is retrieved. The first historical network traffic data is the last segment of the multiple segments of historical network traffic data. The first historical network traffic data, the historical information and the target memory parameter are used to predict the target network traffic data at the future moment after the current moment. The above scheme predicts the target network traffic data at the future moment by jointly predicting the target network traffic data by the first network traffic data and the historical information and target memory parameter related to it. Compared with the existing scheme that performs traffic prediction based only on past time series data, it can improve the prediction rationality and application stability of the model under extremely long input sequences, and further improve the timeliness and accuracy of the prediction results. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments of the present application. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0055] Figure 1 A flowchart of a network traffic prediction method provided in an embodiment of the present application;
[0056] Figure 2 A data processing flow chart of the temporal memory model provided in an embodiment of the present application;
[0057] Figure 3 A flowchart of network traffic prediction and alarm provided in an embodiment of the present application;
[0058] Figure 4 A schematic diagram of weighted embedding of multi-dimensional spatiotemporal information provided in an embodiment of the present application;
[0059] Figure 5 A specific flow chart of the network traffic prediction method provided in an embodiment of the present application;
[0060] Figure 6 A structural block diagram of a network traffic prediction device provided in an embodiment of the present application;
[0061] Figure 7 This is a structural block diagram of the electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0062] The following will be combined with the accompanying drawings in the embodiments of the present application to clearly describe the technical solutions in the embodiments of the present application. Obviously, the embodiments described are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field are within the scope of protection of this application.
[0063] The terms "first," "second," and the like in the specification and claims of this application are used to distinguish similar objects, and are not used to describe a specific order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate, so that the embodiments of this application can be implemented in an order other than that illustrated or described herein, and that the objects distinguished by "first," "second," and the like are generally of the same type, and do not limit the number of objects; for example, the first object can be one or more. In addition, the term "and / or" in the specification and claims refers to at least one of the connected objects, and the character " / " generally indicates that the objects connected are in an "or" relationship.
[0064] Deep learning is a branch of artificial intelligence that typically uses artificial neural networks (especially deep neural networks) to model and learn complex data representations. Leveraging the powerful computing capabilities of today's hardware, such as graphics processing units (GPUs) and tensor processing units (TPUs), deep learning can automatically extract features from large amounts of data through multi-layered nonlinear transformations. It has been successfully applied to tasks such as classification, regression, and generation.
[0065] Currently, deep learning technology is being applied to network traffic prediction and detection, using data learning to construct recurrent time series networks. Given that traffic prediction requires processing long time series data, existing patented technologies are mostly based on Long Short-Term Memory (LSTM) networks or Transformer neural network models from deep learning. In other application areas, some patents propose improvements to the Transformer model by adding memory mechanisms to enhance performance when processing long time series inputs.
[0066] Although these existing technical solutions can achieve relatively good practical application results, they still have certain technical defects. As network traffic becomes increasingly complex in the future, the time series processing models they use will be unable to support the input of extremely long sequences. Specifically:
[0067] Existing solutions often have high time and space complexity. Although some solutions attempt to linearize the model, they face the problem of information loss and perform poorly when processing long contexts because they need to compress historical data into fixed-size vectors or matrices.
[0068] On the other hand, the time series learning models used in existing solutions are trained and inferred on a time-step basis, making it difficult to fully utilize the parallel computing capabilities of modern hardware. Therefore, related model solutions for network traffic prediction still have considerable room for improvement in hardware efficiency. In application areas other than traffic prediction, some solutions have proposed introducing memory mechanisms based on time series models. However, most of these solutions are limited to simple short-term queue storage of previous input data features, lack the utilization of long-term or even global sequence information, and face performance bottlenecks when processing extremely long sequence inputs.
[0069] Therefore, the present application provides a network traffic prediction method that can perform block operations on long time series inputs, so that it pays more attention to local information. It also predicts the target network traffic data at future moments through the first network traffic data and its related historical information and target memory parameters. It can improve the prediction rationality and application stability of the model under extremely long input sequences, and improve the timeliness and accuracy of the prediction results.
[0070] The network traffic prediction method provided by the embodiment of the present application is described in detail below with reference to the accompanying drawings through specific embodiments and their application scenarios.
[0071] Specifically, the present invention provides a method for predicting network traffic. Figure 1 As shown, the following steps may be specifically included:
[0072] Step 101: input historical network traffic data into a traffic prediction model, and perform segmentation processing on the historical network traffic data to obtain multiple segments of historical network traffic data.
[0073] The traffic prediction model includes a time series memory model constructed by integrating time series and memory structures, such as Figure 2 As shown in the figure, the temporal memory model includes: a window sequence processing module, a long-term information retrieval module, and a persistent memory storage module, thereby forming a hierarchical memory mode for long temporal inputs. This will effectively reduce the time and space complexity of the technical solution, achieve efficient parallel computing through adaptive block processing of temporal data, and ultimately support the encoding abstraction and output inference prediction of extremely long network traffic input sequences.
[0074] Among them, the window sequence processing module is used to process the short-term input sequence of current historical network traffic data and predict the output network traffic; the long-term information retrieval module is used to dynamically store and retrieve long-term historical information; and the persistent memory storage module is used to store task-related knowledge, the parameters of which are independent of the input historical network traffic data.
[0075] The window sequence processing module performs hardware adaptive segmentation processing on the historical network traffic data (i.e., the input sequence). Specifically, the input sequence is divided into multiple segments (e.g., t segments) of historical network traffic data with a fixed size. The multiple segments of historical network traffic data have the same fixed window length, and the tth segment of historical network traffic data is recorded as S t Exemplarily, the window sequence processing module is used to segment historical network traffic data containing Internet Protocol (IP) network traffic information into a plurality of historical network traffic data of a fixed size of 256.
[0076] The window length can be adaptively adjusted based on the available memory capacity of the hardware device where the traffic prediction model is running to ensure actual inference and operation efficiency:
[0077] w=α·Mem+β
[0078] Where w represents the window length;
[0079] Mem represents the available memory capacity of the hardware device where the traffic prediction model runs;
[0080] α and β are different configurable parameters.
[0081] The following describes the process of obtaining historical network traffic data:
[0082] like Figure 3As shown, historical raw network traffic data related to network traffic is obtained (i.e., data preparation), and the raw network traffic data is preprocessed, that is, a deep decomposition architecture is used to decompose the data according to cycles and trends, and the raw network traffic data is decomposed into: long-cycle trend sequences, short-cycle trend sequences, and seasonal cycle trend sequences.
[0083] For each trend sequence, long time series data containing information such as inflow, outflow, and bandwidth utilization are embedded and coded according to the Greenwich time node. In addition, it is necessary to further consider multi-dimensional spatiotemporal information such as geographic location, hour, day, week, month, and whether it is a holiday, and perform additional spatiotemporal information embedding coding. Finally, a multi-dimensional embedding vector E = [e1, e2, ..., e n ], where 1 to n represent different dimensions.
[0084] In order to more accurately weigh the importance of each dimension embedding vector, a neural network is used to calculate each dimension embedding e i (i is any integer from 1 to n) attention score:
[0085]
[0086] Among them, W pre and b pre are different learnable neural network parameters and are updated via the back-gradient propagation algorithm;
[0087] v pre is a weight vector;
[0088] tanh refers to the hyperbolic tangent function.
[0089] The attention score a of each embedding i,score Normalized by Softmax function, the corresponding attention weight a is obtained i .
[0090] like Figure 4 As shown in the figure, attention weights are used to perform weighted calculations on the embedding vectors of each dimension. Weighted calculations are performed on the embedding vectors of each dimension at each moment to obtain historical network traffic data for subsequent traffic prediction model processing. For example, trend sequence embedding is weighted with location information embedding and time information embedding.
[0091] like Figure 3The following example illustrates the data preprocessing process: Using a deep decomposition architecture, we decompose the long-term raw network traffic data, consisting of customer IP traffic data over a wide time period, into periods and trends. This decomposition is then performed based on long-term trend sequences, short-term trend sequences, and seasonal trend sequences. These decomposed trend sequences are then encoded according to time nodes and embedded using attention-weighted embedding, integrating multiple information such as time, day, hour, week, month, holidays, and location information, to generate historical network traffic data.
[0092] The above scheme preprocesses the acquired raw network traffic data and converts it into a weighted embedding expression of multi-dimensional spatiotemporal information. That is, by combining the construction of spatiotemporal dimensions and the distribution of weights, it realizes the conversion of raw network traffic data into a weighted embedding vector that can fully reflect its spatiotemporal characteristics and importance, laying a solid foundation for subsequent machine learning, deep learning and other tasks, and forming an input suitable for subsequent deep learning models.
[0093] Step 102: Retrieve historical information related to the first historical network traffic data according to the query weight matrix of the traffic prediction model, and retrieve target memory parameters related to the first historical network traffic data from multiple preset memory parameters, where the first historical network traffic data is the last segment of historical network traffic data among the multiple segments of historical network traffic data.
[0094] like Figure 2 As shown, the query matrix in the window sequence processing module is obtained, and the long-term information retrieval module outputs historical information related to the query according to the query matrix in the window sequence processing module (i.e., the historical retrieval step). In order to avoid information redundancy caused by the full amount of persistent memory parameters, the persistent memory storage module receives the first historical network traffic data as input, and then calculates the cosine similarity based on the multiple preset memory parameters P stored in itself (i.e., the memory parameter retrieval step), recalls and outputs the top N most relevant target memory parameters P ′ =[p1,p2,…,p N ] (i.e., the Top N recall step).
[0095] In one embodiment, the recall parameter N can be set to 128, and the persistent memory storage module can be used to calculate and sample the most relevant set of memory parameters P according to the first historical network traffic data. ′ .
[0096] It should be noted that, in the process of segmenting the historical network traffic data, the historical network traffic data is segmented in chronological order, each segment is a segment of historical network traffic data, and the last segment of historical network traffic data is the first historical network traffic data.
[0097] Step 103 : Using the first historical network traffic data, the historical information, and the target memory parameters, predict target network traffic data at a future moment after the current moment.
[0098] like Figure 3 As shown, based on the first historical network traffic data, historical information, and target memory parameters, the network traffic at the future moment is inferred and predicted to obtain the target network traffic data (i.e., the model inference predicted traffic). The above scheme predicts the target network traffic data at the future moment through the first network traffic data and its related historical information and target memory parameters. Compared with the existing scheme that only predicts traffic based on past time series data, it can improve the prediction rationality and application stability of the model under extremely long input sequences, and further improve the timeliness and accuracy of the prediction results.
[0099] In one embodiment, if Figure 3 As shown, the target network traffic data at the predicted future moment is compared with the pre-set abnormal monitoring alarm threshold. If the traffic size of the target network traffic data is greater than the abnormal monitoring alarm threshold, the early warning mechanism is activated and the network administrator is notified in advance. Specifically, the network administrator can be notified by emitting sound and light, SMS reminders and early warnings; otherwise, the current process ends, thereby realizing abnormal monitoring of traffic flow.
[0100] In the above embodiment of the present application, historical network traffic data is input into the traffic prediction model, and the historical network traffic data is segmented to obtain multiple segments of historical network traffic data. The block operation is performed on the long time series input (i.e., historical network traffic data). Compared with the existing scheme that uses the complete sequence for traffic prediction, it pays more attention to local information, can effectively reduce the computational complexity, and improve the efficiency of reasoning calculation. In addition, according to the query weight matrix of the traffic prediction model, historical information related to the first historical network traffic data is retrieved, and the target memory parameter related to the first historical network traffic data in multiple preset memory parameters is retrieved. The first historical network traffic data is the last segment of the multiple segments of historical network traffic data. The first historical network traffic data, the historical information and the target memory parameter are used to predict the target network traffic data at the future moment after the current moment. The above scheme predicts the target network traffic data at the future moment by jointly predicting the target network traffic data by the first network traffic data and the historical information and target memory parameter related to it. Compared with the existing scheme that performs traffic prediction based only on past time series data, it can improve the prediction rationality and application stability of the model under extremely long input sequences, and further improve the timeliness and accuracy of the prediction results.
[0101] In an optional specific embodiment, step 102 retrieves historical information related to the first historical network traffic data according to the query weight matrix of the traffic prediction model, including:
[0102] Calculating a query vector corresponding to a first piece of historical network traffic data among the multiple pieces of historical network traffic data according to a query weight matrix of the traffic prediction model;
[0103] Based on the query vector corresponding to the first historical network traffic data, historical information related to the first historical network traffic data is retrieved.
[0104] The long-term information retrieval module receives the first historical network traffic data as input and first generates a query vector using the query weight matrix in the window sequence processing module:
[0105] q t =S t W Q
[0106] Among them, q t is the query vector; S t is the first historical network traffic data; W Q is the query weight matrix.
[0107] Then, the query vector is used as the input of the long-term information retrieval module, which calculates and outputs the historical information embedding expression related to the query vector:
[0108] h t =M t-1 (q t )
[0109] Among them, h t Indicates the embedded expression of historical information; q t Expression query vector; M t-1 It represents the target network parameter expression corresponding to the second historical network traffic data in the long-term information retrieval module, where the second historical network traffic data is the previous historical network traffic data of the first historical network traffic data.
[0110] In an optional specific embodiment, the step 103 uses the first historical network traffic data, the historical information, and the target memory parameter to predict target network traffic data at a future moment after the current moment, including:
[0111] splicing the first historical network traffic data, the historical information, and the target memory parameter to obtain spliced data;
[0112] Based on the spliced data, target network traffic data at a future time after the current time is predicted.
[0113] The first historical network traffic data S t 、Historical information t , target memory parameter P ′ The data is spliced together to form a new input sequence (i.e., spliced data). The network traffic at the future moment is inferred and predicted through the new input sequence to obtain the target network traffic data.
[0114] For example, the specific splicing method may be:
[0115]
[0116] The above formula represents the first historical network traffic data S t 、Historical information t , target memory parameter P ′ Perform horizontal splicing to obtain a new input sequence
[0117] In an optional specific embodiment, the step of predicting target network traffic data at a future time after the current time based on the spliced data specifically includes:
[0118] A multi-head attention mechanism for grouped queries is used to calculate the pre-output traffic data corresponding to the spliced data. In the multi-head attention mechanism, queries within the same group share the same key and value during the calculation process.
[0119] Calculating a target network parameter expression corresponding to the first historical network traffic data according to a target network parameter expression corresponding to second historical network traffic data, where the second historical network traffic data is historical network traffic data of a previous period of the first historical network traffic data;
[0120] According to the target network parameter expression corresponding to the first historical network traffic data and the pre-output traffic data, the target network traffic data at a future moment after the current moment is predicted.
[0121] like Figure 2 As shown in the figure, to fully learn the multi-dimensional features of the new input sequence (i.e., the concatenated data), a multi-head attention (Grouped Query Attention, GQA) mechanism is used to calculate the pre-output traffic data. The number of queries in a group is a configurable variable (e.g., it can be set to 8, meaning each group contains 8 queries). Unlike traditional multi-head attention mechanisms, queries within the same group in GQA share the same key and value during the calculation process.
[0122] Obtain the target network parameter expression corresponding to the second historical network traffic data of the previous segment of the first historical network traffic data, and calculate the target network parameter expression corresponding to the first historical network traffic data based on the target network parameter expression corresponding to the second historical network traffic data. Then, based on the pre-output traffic data and the target network parameter expression corresponding to the first historical network traffic data, predict the target network traffic data at a future moment and use it as the output of the traffic prediction model.
[0123] In an optional specific embodiment, the step of calculating the target network parameter expression corresponding to the first historical network traffic data based on the target network parameter expression corresponding to the second historical network traffic data specifically includes:
[0124] Calculating a surprising update gradient of the traffic prediction model according to the target network parameter expression corresponding to the first historical network traffic data and the second historical network traffic data;
[0125] Calculating the cumulative amount of surprise corresponding to the first historical network traffic data according to the cumulative amount of surprise corresponding to the second historical network traffic data and the update gradient of the surprise of the traffic prediction model;
[0126] The target network parameter expression corresponding to the first historical network traffic data is calculated according to the target network parameter expression corresponding to the second historical network traffic data and the cumulative amount of surprise corresponding to the first historical network traffic data.
[0127] In one embodiment, first, based on the target network parameter expression corresponding to the first historical network traffic data and the second historical network traffic data, the surprising update gradient of the traffic prediction model is calculated. Specifically, the calculation can be performed using the following formula:
[0128]
[0129] in, The updated gradient representing the surprise of the traffic prediction model;
[0130] M t-1 An expression representing a target network parameter corresponding to the second historical network traffic data;
[0131] S t Indicates the first historical network traffic data;
[0132] W K and W V denote the key weight matrix and value weight matrix respectively;
[0133] Represents the gradient symbol;
[0134] Represents the square of the L2 norm.
[0135] Surprise update gradients are a method for adjusting parameter updates based on the model's degree of surprise at new information. By quantifying the model's degree of surprise, or "surprise," the method dynamically adjusts the direction and magnitude of model parameter updates. When a model encounters new information that is inconsistent with its expectations, this "surprise" guides the model parameter updates in a direction that reduces this inconsistency, thereby improving the model's adaptability and performance, helping the model adapt to new data more quickly and enhancing learning efficiency and performance.
[0136] In the above formula, first, S t and W K Multiply to get the first result, and then input the first result into M t-1 The second result is obtained, and then S t and W V Multiply to get the third result, calculate the square of the L2 norm of the difference between the second result and the third result, and then perform partial derivative operations to get the surprising update gradient of the traffic prediction model.
[0137] This approach differs from existing solutions in that the long-term information retrieval module in the traffic prediction model also updates during inference, dynamically adjusting the memory state. This allows the model to capture new information in the input data while forgetting unimportant historical information. The traffic prediction model's surprise update gradient measures the difference between the current input and historical memory. A smaller value indicates that similar information is already stored and does not require repeated learning. Furthermore, the historical information output by the long-term information retrieval module is more relevant to the initial historical network traffic data.
[0138] In one embodiment, the cumulative amount of surprise corresponding to the first historical network traffic data is calculated based on the cumulative amount of surprise corresponding to the second historical network traffic data and the update gradient of the surprise of the traffic prediction model. Specifically, the calculation can be performed using the following formula:
[0139]
[0140] Among them, μ t represents the cumulative amount of surprise corresponding to the first historical network traffic data;
[0141] μ t-1 represents the cumulative amount of surprise corresponding to the second historical network traffic data;
[0142] η t represents the momentum coefficient;
[0143] θ t represents the learning rate;
[0144] The updated gradient representing the surprise of the traffic prediction model.
[0145] The cumulative amount of surprise quantifies the sum of all "surprise" events encountered by the model during its lifetime. This cumulative amount reflects the model's learning process from historical experience, helping the model make more accurate and reasonable predictions when faced with new, unseen data. The accumulated amount of surprise can be considered a quantitative indicator of the model's knowledge growth. It guides the model in striking a balance between exploring new strategies and leveraging known ones, thereby improving the model's long-term adaptability.
[0146] In the above formula, η t and μ t-1 Multiply to get the fourth result, and set θ t and The fifth result is obtained by multiplying them, and the accumulated amount of surprise corresponding to the first historical network traffic data can be obtained by subtracting the fifth result from the fourth result.
[0147] In one embodiment, the target network parameter expression corresponding to the first historical network traffic data is calculated based on the target network parameter expression corresponding to the second historical network traffic data and the cumulative amount of surprise corresponding to the first historical network traffic data. Specifically, the calculation can be performed using the following formula:
[0148] M t =(1-α t )M t-1 +μ t
[0149] Among them, M t An expression representing a target network parameter corresponding to the first historical network traffic data;
[0150] M t-1 An expression representing a target network parameter corresponding to the second historical network traffic data;
[0151] α t Is a weight decay coefficient used to control the forgetting mechanism;
[0152] μ t Indicates the cumulative amount of surprise corresponding to the first historical network traffic data.
[0153] In the above formula, subtract α from 1 t The sixth result is obtained by multiplying the difference between the two and the target network parameter expression corresponding to the second historical network traffic data. t The target network parameter expression corresponding to the first historical network traffic data can be obtained by adding them together.
[0154] In an optional specific embodiment, the step of predicting target network traffic data at a future moment after the current moment based on the target network parameter expression corresponding to the first historical network traffic data and the pre-output traffic data specifically includes:
[0155] Calculating a traffic prediction result based on a target network parameter expression corresponding to the first historical network traffic data and the pre-output traffic data;
[0156] The pre-output traffic data and the traffic prediction result are weightedly calculated to obtain the target network traffic data.
[0157] In one embodiment, if Figure 2 As shown, based on the target network parameter expression corresponding to the first historical network traffic data and the pre-output traffic data, the traffic prediction result is calculated, that is, the pre-output traffic data is input into the target network parameter expression corresponding to the first historical network traffic data, and the traffic prediction result (i.e., the result of the secondary search) is calculated. Then, the pre-output traffic data and the traffic prediction result are weighted to obtain the predicted target network traffic data (i.e., the prediction result). Specifically, the target network traffic data can be calculated using the following formula:
[0158] O t =α O y t +(1-α O )M t (y t )
[0159] Among them, O t Indicates target network traffic data;
[0160] α O represents the prediction weight coefficient;
[0161] y t Indicates pre-output flow data;
[0162] M t An expression representing a target network parameter corresponding to the first historical network traffic data;
[0163] M t (y t ) represents the traffic prediction result.
[0164] In the above formula, α O with y t Multiply to get the seventh result, subtract α from 1 O The difference between t (y t) to obtain the eighth result, and the target network traffic data can be obtained by adding the seventh result to the eighth result.
[0165] The above process of calculating the target network parameter expression corresponding to the first historical network traffic data continuously updates the network parameter expression of the long-term information retrieval module in the traffic prediction model during inference, and synchronously weights the results of the secondary retrieval based on the pre-output traffic data. Compared with the existing scheme that only updates parameters during the training phase, it can achieve dynamic adjustment of the model memory and rapid adaptation to new input sequence data, and enhance the real-time and accuracy of the prediction.
[0166] like Figure 5 As shown, the above flow prediction process is described below through a specific embodiment:
[0167] Step 501: Input historical network traffic data;
[0168] Step 502: Segment the historical network traffic data.
[0169] Step 503: Calculate the query vector corresponding to the first historical network traffic data according to the query weight matrix.
[0170] Step 504: Retrieve historical information related to the first historical network traffic data based on the query vector corresponding to the first historical network traffic data.
[0171] Step 505: Retrieve a target memory parameter related to the first historical network traffic data from a plurality of preset memory parameters.
[0172] Step 506: Splice the first historical network traffic data, historical information, and target memory parameters to obtain spliced data.
[0173] Step 507: Use the multi-head attention mechanism of group query to calculate the pre-output traffic data corresponding to the spliced data.
[0174] Step 508: Calculate the target network parameter expression corresponding to the first historical network traffic data based on the target network parameter expression corresponding to the second historical network traffic data.
[0175] Step 509: predicting target network traffic data at a future time based on the target network parameter expression corresponding to the first historical network traffic data and the pre-output traffic data.
[0176] The following describes the training process of the traffic prediction model:
[0177] If the traffic prediction model needs to be trained, the traffic prediction model training step is entered. The prediction accuracy of the traffic prediction model is improved by optimizing the parameters of the proposed temporal memory model.
[0178] When the neural network parameters in the traffic prediction module need to be trained, the following steps are performed to update the parameters:
[0179] First, calculate the global loss function
[0180]
[0181] Among them, N batch Indicates the number of samples in the current training batch, i ranges from 1 to N batch Any sample in ;
[0182] Represents the real network traffic data of the i-th sample.
[0183] Parameters of the window sequence processing module, i.e., the query weight matrix W of the attention mechanism Q , key weight matrix W K And the value weight matrix W V . Updated via backpropagation and gradient descent:
[0184]
[0185] Among them, η c Indicates the learning rate of the core module;
[0186] Represents the gradient symbol;
[0187] Loss represents the global loss function;
[0188] Indicates the partial derivative operation on Loss.
[0189] The above update formula shows that each weight matrix (W Q 、W K and W V ) minus η c and The difference of the product of is the updated weight matrix.
[0190] The parameter update of the long-term information retrieval module is slightly different from that in the inference phase. Specifically, the parameter update of the long-term information retrieval module is also calculated based on the global loss function Loss:
[0191]
[0192] Among them, M i Represents the target network parameter expression corresponding to the i-th sample data;
[0193] Mi-1 Represents the target network parameter expression corresponding to the i-1th sample data;
[0194] α t Is a weight decay coefficient used to control the forgetting mechanism;
[0195] μ i Indicates the cumulative amount of surprise corresponding to the i-th sample data;
[0196] μ i-1 Indicates the cumulative amount of surprise corresponding to the i-1th sample data;
[0197] η t represents the momentum coefficient;
[0198] θ t represents the learning rate;
[0199] Represents the surprising update gradient of the traffic prediction model obtained by performing partial derivative operations on Loss.
[0200] The parameters of the persistent memory storage module are updated through backpropagation and gradient descent:
[0201]
[0202] Among them, η p represents the learning rate of the persistent memory module; P represents the memory parameter; Represents the surprising update gradient of the traffic prediction model obtained by performing partial derivative operations on Loss.
[0203] The above update formula shows that the parameter P of the persistent memory storage module is subtracted from η p and The difference of the product is the parameter of the updated persistent memory storage module.
[0204] Update the learnable neural network parameters:
[0205]
[0206] Among them, W pre and b pre are different learnable neural network parameters; η pre represents the learning rate of the embedding importance network; Represents the surprising update gradient of the traffic prediction model obtained by performing partial derivative operations on Loss.
[0207] The above update formula shows that the learnable neural network parameters minus η pre and The difference of the product is the updated learnable neural network parameters.
[0208] To sum up, the above embodiments of the present application, in order to solve the problem of model application effect, comprehensively utilize the multi-dimensional information of short-term window sequence (i.e., the first historical network traffic data), historical retrieval content (i.e., historical information), and recalled persistent memory parameters (i.e., target memory parameters) to infer future time slot traffic, and dynamically update the long-term information retrieval module parameters in the reasoning stage, and design the weighted inclusion of secondary retrieval content (i.e., traffic prediction results) in the predicted output (i.e., target network traffic data), so that the model can more effectively capture the long-term dependencies in the input data and maintain the application stability of the model under extremely long input sequences.
[0209] To address model inference efficiency issues, a new architecture that integrates a time series model with a memory model is constructed. The window sequence processing module processes the data of the current segment, focusing on local information. The long-term information retrieval module retrieves historical features across blocks, resolving the contextual disconnection caused by block partitioning. The persistent memory storage module fixes learned global knowledge to enable full-scale retrieval. By designing a hardware-adaptive block partitioning mechanism, a limited-time window attention calculation mechanism, and a group query multi-head attention cache mechanism, it can support the processing of extremely long-term spatiotemporal embedded information and significantly improve the high computational complexity and low inference efficiency faced by existing technical solutions.
[0210] The above describes the network traffic prediction method provided by the embodiment of the present application. The following will describe the network traffic prediction device provided by the embodiment of the present application with reference to the accompanying drawings.
[0211] like Figure 6 As shown, the embodiment of the present application further provides a network traffic prediction device 600, which includes:
[0212] Segmentation module 601, used to input historical network traffic data into the traffic prediction model, segment the historical network traffic data, and obtain multiple segments of historical network traffic data;
[0213] a retrieval module 602 configured to retrieve historical information related to first historical network traffic data based on a query weight matrix of the traffic prediction model, and retrieve a target memory parameter related to the first historical network traffic data from a plurality of preset memory parameters, wherein the first historical network traffic data is a last segment of the plurality of segments of historical network traffic data;
[0214] The prediction module 603 is configured to use the first historical network traffic data, the historical information, and the target memory parameters to predict target network traffic data at a future moment after the current moment.
[0215] Optionally, when the retrieval module 602 retrieves historical information related to the first historical network traffic data according to the query weight matrix of the traffic prediction model, it is specifically configured to:
[0216] Calculating a query vector corresponding to a first piece of historical network traffic data among the multiple pieces of historical network traffic data according to a query weight matrix of the traffic prediction model;
[0217] Based on the query vector corresponding to the first historical network traffic data, historical information related to the first historical network traffic data is retrieved.
[0218] Optionally, the prediction module 603 is specifically configured to:
[0219] splicing the first historical network traffic data, the historical information, and the target memory parameter to obtain spliced data;
[0220] Based on the spliced data, target network traffic data at a future time after the current time is predicted.
[0221] Optionally, when predicting target network traffic data at a future time after the current time based on the spliced data, the prediction module 603 is specifically configured to:
[0222] A multi-head attention mechanism for grouped queries is used to calculate the pre-output traffic data corresponding to the spliced data. In the multi-head attention mechanism, queries within the same group share the same key and value during the calculation process.
[0223] Calculating a target network parameter expression corresponding to the first historical network traffic data according to a target network parameter expression corresponding to second historical network traffic data, where the second historical network traffic data is historical network traffic data of a previous period of the first historical network traffic data;
[0224] According to the target network parameter expression corresponding to the first historical network traffic data and the pre-output traffic data, the target network traffic data at a future moment after the current moment is predicted.
[0225] Optionally, when the prediction module 603 calculates the target network parameter expression corresponding to the first historical network traffic data based on the target network parameter expression corresponding to the second historical network traffic data, it is specifically configured to:
[0226] Calculating a surprising update gradient of the traffic prediction model according to the target network parameter expression corresponding to the first historical network traffic data and the second historical network traffic data;
[0227] Calculating the cumulative amount of surprise corresponding to the first historical network traffic data according to the cumulative amount of surprise corresponding to the second historical network traffic data and the update gradient of the surprise of the traffic prediction model;
[0228] The target network parameter expression corresponding to the first historical network traffic data is calculated according to the target network parameter expression corresponding to the second historical network traffic data and the cumulative amount of surprise corresponding to the first historical network traffic data.
[0229] Optionally, the target network parameter expression corresponding to the first historical network traffic data is calculated by the following formula:
[0230] M t =(1-α t )M t-1 +μ t
[0231]
[0232] Among them, M t An expression representing a target network parameter corresponding to the first historical network traffic data;
[0233] M t-1 An expression representing a target network parameter corresponding to the second historical network traffic data;
[0234] α t represents the weight attenuation coefficient;
[0235] μ t represents the cumulative amount of surprise corresponding to the first historical network traffic data;
[0236] μ t-1 represents the cumulative amount of surprise corresponding to the second historical network traffic data;
[0237] η t represents the momentum coefficient;
[0238] θ t represents the learning rate;
[0239] The updated gradient representing the surprise of the traffic prediction model;
[0240] S t Indicates the first historical network traffic data;
[0241] W K and W V denote the key weight matrix and value weight matrix respectively;
[0242] Represents the gradient symbol;
[0243] Represents the square of the L2 norm.
[0244] Optionally, when predicting target network traffic data at a future moment after the current moment based on the target network parameter expression corresponding to the first historical network traffic data and the pre-output traffic data, the prediction module 603 is specifically configured to:
[0245] Calculating a traffic prediction result based on a target network parameter expression corresponding to the first historical network traffic data and the pre-output traffic data;
[0246] The pre-output traffic data and the traffic prediction result are weightedly calculated to obtain the target network traffic data.
[0247] Optionally, the target network traffic data is calculated using the following formula:
[0248] O t =α O y t +(1-α O )M t (y t )
[0249] Among them, O t Indicates target network traffic data;
[0250] α O represents the prediction weight coefficient;
[0251] y t Indicates pre-output flow data;
[0252] M t Represents the target network parameter expression corresponding to the first historical network traffic data.
[0253] It should be noted here that the above-mentioned network traffic prediction device provided in the embodiment of the present application can implement all the method steps implemented in the above-mentioned network traffic prediction method embodiment, and can achieve the same technical effect. The parts and beneficial effects of this embodiment that are the same as those in the method embodiment will not be described in detail here.
[0254] It should be noted that the division of units in the embodiments of the present application is schematic and is merely a logical functional division. In actual implementation, other division methods may be used. Furthermore, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0255] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a processor-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) or a processor to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0256] like Figure 7 As shown, an embodiment of the present application further provides an electronic device, including a memory 720, a transceiver 710, and a processor 700:
[0257] Memory 720, for storing computer programs;
[0258] a transceiver 710 for transmitting and receiving data under the control of the processor;
[0259] The processor 700 is configured to read the computer program in the memory and execute the steps of the network traffic prediction method as described in any of the above embodiments.
[0260] Among them, Figure 7 In the embodiment, the bus architecture may include any number of interconnected buses and bridges, specifically various circuits of one or more processors represented by processor 700 and memory represented by memory 720. The bus architecture may also link together various other circuits such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art and therefore will not be described further herein. The bus interface provides an interface. The transceiver 710 may be a plurality of components, i.e., a transmitter and a receiver, providing a unit for communicating with various other devices on a transmission medium, such as a wireless channel, a wired channel, an optical cable, and the like. The processor 700 is responsible for managing the bus architecture and general processing, and the memory 720 may store data used by the processor 700 when performing operations.
[0261] The processor 700 may be a central processing unit (CPU), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), or a complex programmable logic device (CPLD). The processor may also adopt a multi-core architecture.
[0262] The processor calls the computer program stored in the memory to execute the network traffic prediction method provided by the embodiment of the present application according to the obtained executable instructions. The processor and the memory can also be arranged physically separately.
[0263] It should be noted here that the above-mentioned electronic device provided in the embodiment of the present application can implement all the method steps implemented in the above-mentioned network traffic prediction method embodiment, and can achieve the same technical effect. The parts and beneficial effects of this embodiment that are the same as the method embodiment will not be described in detail here.
[0264] An embodiment of the present application further provides a processor-readable storage medium, wherein the processor-readable storage medium stores a computer program, and the computer program is used to enable the processor to execute the above-mentioned network traffic prediction method.
[0265] The processor-readable storage medium can be any available medium or data storage device that can be accessed by the processor, including but not limited to magnetic storage (such as floppy disks, hard disks, magnetic tapes, magneto-optical disks (MO)), optical storage (such as CDs, DVDs, BDs, HVDs, etc.), and semiconductor storage (such as ROMs, EPROMs, EEPROMs, non-volatile memories (NANDFLASH), solid-state drives (SSDs)), etc.
[0266] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage and optical storage, etc.) that contain computer-usable program code.
[0267] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer-executable instructions. These computer-executable instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the steps in the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0268] These processor-executable instructions may also be stored in a processor-readable memory that can direct a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the processor-readable memory produce an article of manufacture comprising an instruction device that implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0269] These processor-executable instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are performed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for executing on the computer or other programmable device to implement the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0270] Obviously, those skilled in the art may make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if such changes and modifications of the present application fall within the scope of the claims of the present application and their equivalents, the present application is intended to include such changes and modifications.
Claims
1. A network traffic prediction method, characterized in that: The method comprises: Inputting historical network traffic data into a traffic prediction model, segmenting the historical network traffic data to obtain multiple segments of historical network traffic data; Retrieving historical information related to first historical network traffic data according to a query weight matrix of the traffic prediction model, and retrieving a target memory parameter related to the first historical network traffic data from a plurality of preset memory parameters, wherein the first historical network traffic data is a last segment of the plurality of segments of historical network traffic data; The first historical network traffic data, the historical information and the target memory parameter are used to predict target network traffic data at a future moment after the current moment.
2. The method according to claim 1, characterized in that The retrieving historical information related to the first historical network traffic data according to the query weight matrix of the traffic prediction model includes: Calculating a query vector corresponding to a first piece of historical network traffic data among the plurality of segments of historical network traffic data according to a query weight matrix of the traffic prediction model; Based on the query vector corresponding to the first historical network traffic data, historical information related to the first historical network traffic data is retrieved.
3. The method according to claim 1, characterized in that The step of using the first historical network traffic data, the historical information, and the target memory parameter to predict target network traffic data at a future moment after the current moment includes: splicing the first historical network traffic data, the historical information, and the target memory parameter to obtain spliced data; Based on the spliced data, target network traffic data at a future time after the current time is predicted.
4. The method according to claim 3, characterized in that The step of predicting target network traffic data at a future time after the current time based on the spliced data includes: A multi-head attention mechanism for grouped queries is used to calculate the pre-output traffic data corresponding to the spliced data. In the multi-head attention mechanism, queries within the same group share the same key and value during the calculation process. Calculating a target network parameter expression corresponding to the first historical network traffic data according to a target network parameter expression corresponding to second historical network traffic data, where the second historical network traffic data is historical network traffic data of a previous period of the first historical network traffic data; According to the target network parameter expression corresponding to the first historical network traffic data and the pre-output traffic data, the target network traffic data at a future moment after the current moment is predicted.
5. The method according to claim 4, characterized in that The calculating, based on the target network parameter expression corresponding to the second historical network traffic data, the target network parameter expression corresponding to the first historical network traffic data includes: Calculating a surprising update gradient of the traffic prediction model according to the target network parameter expression corresponding to the first historical network traffic data and the second historical network traffic data; Calculating the cumulative amount of surprise corresponding to the first historical network traffic data according to the cumulative amount of surprise corresponding to the second historical network traffic data and the update gradient of the surprise of the traffic prediction model; The target network parameter expression corresponding to the first historical network traffic data is calculated according to the target network parameter expression corresponding to the second historical network traffic data and the cumulative amount of surprise corresponding to the first historical network traffic data.
6. The method according to claim 5, characterized in that The target network parameter expression corresponding to the first historical network traffic data is specifically calculated using the following formula: M t =(1-a t )M t-1 +m t Among them, M t An expression representing a target network parameter corresponding to the first historical network traffic data; M t-1 An expression representing a target network parameter corresponding to the second historical network traffic data; α t represents the weight attenuation coefficient; μ t represents the cumulative amount of surprise corresponding to the first historical network traffic data; μ t-1 represents the cumulative amount of surprise corresponding to the second historical network traffic data; η t represents the momentum coefficient; θ t represents the learning rate; The updated gradient representing the surprise of the traffic prediction model; S t Indicates the first historical network traffic data; W K and W V denote the key weight matrix and value weight matrix respectively; Represents the gradient symbol; Represents the square of the L2 norm.
7. The method according to any one of claims 4 to 6, characterized in that The predicting, based on the target network parameter expression corresponding to the first historical network traffic data and the pre-output traffic data, target network traffic data at a future moment after the current moment includes: Calculating a traffic prediction result based on a target network parameter expression corresponding to the first historical network traffic data and the pre-output traffic data; The pre-output traffic data and the traffic prediction result are weightedly calculated to obtain the target network traffic data.
8. The method according to claim 7, characterized in that The target network traffic data is specifically calculated using the following formula: The t =a O y t +(1-a O )M t (y t ) Among them, O t Indicates target network traffic data; α O represents the prediction weight coefficient; y t Indicates pre-output flow data; M t Represents the target network parameter expression corresponding to the first historical network traffic data.
9. A network traffic prediction device, characterized in that: The device comprises: A segmentation module is used to input historical network traffic data into the traffic prediction model, and perform segmentation processing on the historical network traffic data to obtain multiple segments of historical network traffic data; a retrieval module, configured to retrieve historical information related to first historical network traffic data according to a query weight matrix of the traffic prediction model, and retrieve a target memory parameter related to the first historical network traffic data from a plurality of preset memory parameters, wherein the first historical network traffic data is a last segment of the plurality of segments of historical network traffic data; The prediction module is used to predict the target network traffic data at a future moment after the current moment using the first historical network traffic data, the historical information and the target memory parameter.
10. An electronic device, characterized in that: Including memory, transceiver, processor: Memory for storing computer programs; a transceiver, configured to transmit and receive data under the control of the processor; A processor is configured to read the computer program in the memory and execute the network traffic prediction method according to any one of claims 1 to 8.
11. A processor-readable storage medium, characterized in that: The processor-readable storage medium stores a computer program, and the computer program is used to enable the processor to execute the network traffic prediction method according to any one of claims 1 to 8.