Communication network resource optimization method and system based on artificial intelligence

By adopting an artificial intelligence-based method in the communication network, combining time series decomposition and long-term memory network, processing historical network traffic data and performing dynamic calibration, the problems of low prediction accuracy and inability to dynamically adjust the model in the prior art are solved, and more accurate and reliable network traffic prediction is achieved.

CN120224463AActive Publication Date: 2025-06-27BEIJING XUNFENG TIMES SOFTWARE DEVELOPMENT CO LTD

Patent Information

Application Number
CN202510356291.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-25
Publication Date
2025-06-27
Estimated Expiration
2045-03-25

AI Technical Summary

Technical Problem

The existing communication network resource optimization technology is difficult to take into account the changing characteristics of network traffic over different time periods at the same time, resulting in the unsatisfactory prediction accuracy, the correlation between network equipment performance indicators and traffic changes is not fully considered, the impact of network status on traffic prediction and the lack of an effective calibration mechanism for prediction results is not implemented, and the prediction model cannot be dynamically adjusted and optimized based on the actual network operation.

Method used

Using an artificial intelligence-based method, by collecting historical network traffic data, network equipment performance data and user business demand data, the data is decomposed into trend terms, periodic terms and residual terms using a time series decomposition algorithm, and input these terms into the long and short-term memory network for processing, combining the attention mechanism to weight the importance of performance indicators, and using a multi-scale prediction mechanism and dynamic calibration method to correct the prediction value.

Benefits of technology

It improves the accuracy and reliability of network traffic prediction, enhances the model's ability to adapt to changes in the network environment, makes the prediction results more in line with the actual network operation, effectively reduces prediction deviations, and improves the model's prediction performance under different time periods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120224463A_ABST
    Figure CN120224463A_ABST
Patent Text Reader

Abstract

The invention provides a communication network resource optimization method and system based on artificial intelligence, and relates to the technical field of communication optimization, and the method comprises the steps: collecting historical network data, carrying out the feature extraction through time series decomposition and an LSTM algorithm, carrying out the weight calculation of a performance index through combining an attention mechanism, and predicting the network flow through a multi-scale prediction mechanism. And dynamic calibration is carried out based on historical errors, so that the network flow change trend can be accurately predicted, the prediction precision is improved, reasonable distribution and optimal configuration of network resources are realized, and the network operation efficiency is effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of communication optimization, and particularly to a communication network resource optimization method and system based on artificial intelligence. Background Art

[0002] With the rapid development of 5G networks and the in-depth promotion of digital transformation, the types of services and data traffic carried by communication networks have shown explosive growth. Efficient network resource management and optimization are of great significance for ensuring network service quality and improving user experience.

[0003] Traditional communication network resource optimization methods mainly rely on manual experience and fixed rules, and it is difficult to adapt to complex and changeable network environments and dynamically changing service requirements. Currently, methods based on statistical analysis are generally used to predict network traffic and optimize resources, using simple time series models or regression analysis to model and analyze historical data.

[0004] Existing network resource optimization technologies still have problems such as only focusing on predictions at a single time scale, being difficult to simultaneously consider the change characteristics of network traffic in different time periods, resulting in unsatisfactory prediction accuracy, not fully considering the correlation between network device performance indicators and traffic changes, ignoring the impact of network status on traffic prediction, and lacking an effective prediction result calibration mechanism, and being unable to dynamically adjust and optimize the prediction model according to the actual network operation situation.

[0005] Therefore, there is an urgent need for a solution to solve the problems existing in the prior art. Summary of the Invention

[0006] Embodiments of the present invention provide a communication network resource optimization method and system based on artificial intelligence, which can at least solve some problems existing in the prior art.

[0007] In the first aspect of the embodiments of the present invention, a communication network resource optimization method based on artificial intelligence is provided, including:

[0008] Collect historical network traffic data, network device performance data, and user service demand data in the communication network, decompose the historical network traffic data into a trend term, a periodic term, and a residual term through a time series decomposition algorithm, and calculate the bandwidth occupancy rate, processor occupancy rate, memory occupancy rate, and transmission delay in the network device performance data.

[0009] Form an input vector from the trend term, the periodic term, and the residual term, add the input vector to a long short-term memory network, perform gating operations on the input vector through a forgetting gate, an input gate, and an output gate, update the cell state to obtain a hidden state, and perform a linear transformation on the hidden state to obtain a time series encoding.

[0010] The bandwidth occupancy rate, processor occupancy rate, memory occupancy rate, and transmission delay are combined to form a performance metric vector. An attention mechanism is introduced to calculate the importance weights for the performance metric vector, and the temporal encoding and importance weights are fused to obtain the fused features.

[0011] A multi-scale prediction mechanism is used to process the fused features, predicting the network traffic values in the short term, medium term, and long term respectively. Adaptive weights are set based on the prediction errors at each time scale, and the prediction results at multiple time scales are weighted and combined to output the initial network traffic prediction value.

[0012] A correction model is constructed based on the statistical characteristics of historical prediction errors. The exponential smoothing method is used to dynamically calibrate the initial network traffic prediction value, and the prediction value is corrected in combination with the mutation detection results of network performance metrics to obtain the network traffic prediction result.

[0013] In an alternative embodiment,

[0014] Historical network traffic data, network device performance data, and user service demand data in the communication network are collected. The historical network traffic data is decomposed into a trend term, a periodic term, and a residual term through a time series decomposition algorithm. Calculating the bandwidth occupancy rate, processor occupancy rate, memory occupancy rate, and transmission delay in the network device performance data includes:

[0015] Collect historical network traffic data in the communication network within a preset sampling period. The historical network traffic data is the number of bytes of traffic on the network link per unit time. At the same time, collect network device performance data, which reflects the working state of the network link, and collect user service demand data.

[0016] The historical network traffic data is input into the time series decomposition algorithm to extract the fluctuation law of the historical network traffic data over time, obtaining a trend term reflecting the long-term change trend, a periodic term reflecting the periodic change, and a residual term reflecting the random fluctuation.

[0017] Statistical calculations are performed on the network device performance data, and the network device performance data is averaged by sliding within a preset time window to obtain the bandwidth occupancy rate, processor occupancy rate, memory occupancy rate, and transmission delay.

[0018] In an alternative embodiment,

[0019] The trend term, periodic term, and residual term are combined to form an input vector, which is added to the long short-term memory network. Gate control operations are performed on the input vector through the forget gate, input gate, and output gate to update the cell state to obtain the hidden state, and a linear transformation is performed on the hidden state to obtain the temporal encoding, including:

[0020] Sequentially splice the trend term, periodic term, and residual term in the time dimension to form an input vector, where the input vector contains historical values at different time steps;

[0021] Based on the input vector and the pre-acquired historical state information, randomly forget historical information through the forget gate, calculate the current update information through the input gate and generate a candidate cell state, combine the retained part of the historical information with the current update information to update the cell state, and perform a gating operation on the cell state through the output gate to obtain a hidden state;

[0022] Perform a linear transformation on the hidden state to obtain a temporal encoding.

[0023] In an alternative embodiment,

[0024] Form a performance metric vector from the bandwidth occupancy rate, processor occupancy rate, memory occupancy rate, and transmission delay. Introduce an attention mechanism to calculate the importance weights for the performance metric vector, and perform feature fusion on the temporal encoding and the importance weights to obtain a fused feature, including:

[0025] Collect the performance parameters of the network device at multiple consecutive time points, where the performance parameters include the bandwidth occupancy rate, processor occupancy rate, memory occupancy rate, and transmission delay, and form a performance metric vector from the performance parameters collected at each time point;

[0026] Based on the attention mechanism, input the performance metric vector into the query matrix, key matrix, and value matrix to obtain a query vector, a key vector, and a value vector. Based on the dot product operation of the query vector and the key vector, obtain the correlation score between the performance metrics, and perform softmax normalization on the correlation score to obtain the importance weights, where the importance weights represent the influence degree of each performance metric on the network performance;

[0027] Perform weighted superposition on the temporal encoding and the importance weights to obtain a fused feature.

[0028] In an alternative embodiment,

[0029] Process the fused feature using a multi-scale prediction mechanism to respectively predict the network traffic values in the short term, medium term, and long term. Set adaptive weights based on the prediction errors at each time scale, and perform weighted combination on the prediction results at multiple time scales to output an initial network traffic prediction value, including:

[0030] Use a double-layer long short-term memory network structure to respectively construct prediction processing units for three time scales of short term, medium term, and long term. The double-layer long short-term memory network structure processes the fused feature and outputs the initial network traffic prediction values at each time scale;

[0031] Continuously collect network traffic data at each time scale within a preset sampling time window, calculate the change rate of the network traffic data at each time scale, calculate the temporal correlation degree between adjacent time scales and the fluctuation degree of each time scale according to the change rate; set a corresponding temperature coefficient for each time scale based on the fluctuation degree, and the greater the fluctuation degree, the smaller the temperature coefficient;

[0032] Slice historical samples into multiple consecutive training sequence segments at a preset time interval, calculate the difference between the true value of network traffic and the initial network traffic prediction value at the corresponding time scale within each training sequence segment, perform cumulative summation operation on the difference and calculate the average value to generate the prediction error at each time scale; multiply the prediction error by the temperature coefficient of the corresponding time scale and then take the opposite number, substitute it into the exponential function for operation to generate an exponential mapping value, calculate the ratio of the exponential mapping value to the sum of the exponential mapping values of three time scales to determine the initial adaptive weight corresponding to each time scale;

[0033] Extract the law characteristics of the temporal correlation degree changing with time between each time scale according to the graph structure network, establish the constraint relationship between the prediction values of adjacent time scales according to the law characteristics, correct the initial network traffic prediction value based on the constraint relationship and the federated learning framework to generate a corrected prediction value considering temporal correlation, and adjust the initial adaptive weight based on the corrected prediction value to obtain a corrected weight considering temporal correlation, where the corrected weight corresponding to the time scale with a larger fluctuation degree increases relatively;

[0034] Perform weighted summation operation on the corrected prediction value and the corresponding corrected weight, and output the final network traffic prediction value.

[0035] In an optional implementation manner,

[0036] Extract the law characteristics of the temporal correlation degree changing with time between each time scale according to the graph structure network, establish the constraint relationship between the prediction values of adjacent time scales according to the law characteristics, and correct the initial network traffic prediction value based on the constraint relationship and the federated learning framework to generate a corrected prediction value considering temporal correlation, including:

[0037] Construct a dynamic heterogeneous graph network structure, connect network traffic nodes at different time scales through dynamic edges, where the node attributes include traffic values, statistical features and change trends, and the edge attributes are initialized as the basic temporal correlation degree between adjacent time scales; adopt a spatio-temporal attention mechanism to dynamically update the attributes of nodes and edges, and output the updated node state vectors and edge association strengths;

[0038] Construct a time-series correlation degree extraction model based on the updated node state vector and edge association strength, capture the change characteristics of network traffic at different time scales through a sliding time window, and combine the counterfactual reasoning method to identify the dominant factors of the time-series correlation degree, and extract the regular characteristics of the change of the time-series correlation degree over time;

[0039] Construct a conditional time-series pattern graph according to the extracted regular characteristics, map the time-series correlation patterns under different conditions to the nodes in the graph, and establish node connections based on the evolutionary relationship between the patterns. Use the conditional time-series pattern graph to describe the constraint relationship between the predicted values at adjacent time scales, adopt a distributed federated learning framework to integrate the constraint relationship information of multiple network regions, adjust the calculation parameters of the time-series correlation degree in real time based on the local prediction deviation, and continuously optimize the extraction method of the regular characteristics using the global prediction performance index, construct an adaptive probability graph model, and dynamically associate the update frequency of the model parameters with the fluctuation degree of network traffic;

[0040] Construct an energy function according to the node state distribution and transition probability in the adaptive probability graph model, and combine the constraint relationship provided by the conditional time-series pattern graph. The constraint relationship strength, conditional time-series law, and probability distribution are used as components of the energy function, where the weight coefficient of each component is adaptively adjusted according to the change of the prediction accuracy. Iteratively optimize the energy function to correct the initial network traffic prediction value until all time-series constraint conditions are met, and output the corrected prediction value considering time-series correlation.

[0041] In an alternative embodiment,

[0042] Construct a correction model according to the statistical characteristics of historical prediction errors, dynamically calibrate the initial network traffic prediction value using the exponential smoothing method, and combine the mutation detection results of network performance indicators to correct the prediction value. The network traffic prediction results include:

[0043] Obtain a historical prediction error sequence, which is obtained by subtracting the predicted network traffic value from the actual network traffic value. Calculate the error mean and error standard deviation based on the historical prediction error sequence, construct the autocorrelation function of the error sequence, and calculate the exponentially weighted moving variance, where the exponentially weighted moving variance is determined by the weighted sum of the square difference between the current moment error and the mean and the exponentially weighted moving variance of the previous moment;

[0044] Construct a correction model based on the error mean, the error standard deviation, the autocorrelation function, and the exponentially weighted moving variance. The correction model includes an error distribution correction term, a time-series correlation correction term, and a fluctuation correction term. The error distribution correction term, the time-series correlation correction term, and the fluctuation correction term correspond to different adaptive weight coefficients, and the adaptive weight coefficients are updated by the gradient descent method;

[0045] According to the exponential smoothing method, a historical information storage matrix is constructed, and the initial network traffic prediction value is dynamically calibrated through an adaptive smoothing factor and an attention mechanism to obtain a first corrected prediction value. Among them, the adaptive smoothing factor is determined by the weighted sum of the smoothing factor at the previous moment and the absolute value of the historical prediction error at the current moment, and the weighting coefficient is dynamically adjusted through the change amount of the exponentially weighted moving variance;

[0046] A multi-dimensional network performance index vector is constructed. The multi-dimensional network performance index vector includes the values of multiple network performance indexes at the current moment. The cumulative summation operation is performed based on the time series difference of the multi-dimensional network performance index vector to obtain a mutation detection value, and the mutation detection value is compared with the difference between the mutation detection value at the previous moment and the detection threshold;

[0047] When the mutation detection value is greater than the mutation determination threshold, calculate the relative change rate of each network performance index in the multi-dimensional network performance index vector, multiply the relative change rate of each network performance index by the corresponding weight coefficient and sum to obtain the mutation degree, and output the product of the first corrected prediction value and the mutation degree as the network traffic prediction result; when the mutation detection value is not greater than the mutation determination threshold, output the first corrected prediction value as the network traffic prediction result.

[0048] In an alternative embodiment,

[0049] According to the exponential smoothing method, constructing a historical information storage matrix and dynamically calibrating the initial network traffic prediction value through an adaptive smoothing factor and an attention mechanism includes:

[0050] The network traffic time series is divided into multiple time scale layers according to the time scale, and an independent adaptive smoothing factor is constructed for each time scale layer;

[0051] Calculate the volatility based on the change trend of the network traffic time series, dynamically adjust the historical data window size based on the volatility, and the historical data window size is determined by the product of the basic window size and the volatility. Obtain the smoothing results of each time scale layer within the historical data window range;

[0052] Construct a feature matrix from the smoothing results according to the time scale layer, calculate the attention weight based on the feature matrix and the query vector at the current moment, the attention weight represents the importance of different time scale features to the current prediction, and adaptively fuse the features of different time scales using the attention weight;

[0053] Construct a historical information storage matrix, which includes historical network traffic values, the prediction error, and the adaptive smoothing factor. Use the attention mechanism to extract historical information related to the current state from the historical information storage matrix, and determine the dynamic learning rate based on the change amount of the prediction error. The dynamic learning rate decays as the change amount of the prediction error increases.

[0054] Introduce a residual connection structure to the initial network traffic prediction value. Perform weighted summation on the output of the correction function at each time scale and the adaptive smoothing factor according to the attention weights to obtain a residual term, and add the residual term to the initial network traffic prediction value to obtain the first corrected prediction value.

[0055] In the second aspect of the embodiments of the present invention, a communication network resource optimization system based on artificial intelligence is provided, including:

[0056] The first unit is used to collect historical network traffic data, network device performance data, and user service demand data in the communication network. Decompose the historical network traffic data into a trend term, a periodic term, and a residual term through the time series decomposition algorithm, and calculate the bandwidth occupancy rate, processor occupancy rate, memory occupancy rate, and transmission delay in the network device performance data.

[0057] The second unit is used to form an input vector from the trend term, the periodic term, and the residual term, add the input vector to the long short-term memory network, perform gating operations on the input vector through the forget gate, input gate, and output gate, update the unit state to obtain a hidden state, and perform a linear transformation on the hidden state to obtain a time series encoding.

[0058] The third unit is used to form a performance index vector from the bandwidth occupancy rate, processor occupancy rate, memory occupancy rate, and transmission delay, introduce the attention mechanism to calculate the importance weights for the performance index vector, and perform feature fusion on the time series encoding and the importance weights to obtain a fusion feature.

[0059] The fourth unit is used to process the fusion feature using a multi-scale prediction mechanism, respectively predict the network traffic values in the short term, medium term, and long term, set adaptive weights based on the prediction errors at each time scale, and perform weighted combination on the prediction results at multiple time scales to output the initial network traffic prediction value.

[0060] The fifth unit is used to construct a correction model according to the statistical characteristics of the historical prediction error, dynamically calibrate the initial network traffic prediction value using the exponential smoothing method, and combine the mutation detection results of the network performance indicators to correct the prediction value to obtain the network traffic prediction result.

[0061] In the present invention, historical network traffic data is processed by combining time series decomposition and long short-term memory networks, fully considering the temporal characteristics and long-term dependencies of the data, improving the accuracy and reliability of network traffic prediction, providing a more accurate decision-making basis for the optimal allocation of network resources, introducing an attention mechanism to weight the importance of network performance indicators, and fusing with temporal characteristics to achieve a comprehensive perception and feature extraction of the network state, enhancing the adaptability of the model to changes in the network environment, making the prediction results more in line with the actual network operation situation, adopting a multi-scale prediction mechanism and a dynamic calibration method, adaptively combining the prediction results of different time scales through adaptive weights, and correcting them in combination with the statistical characteristics of historical errors, effectively reducing the prediction deviation and improving the prediction performance of the model under different time periods, providing more accurate prediction support for the intelligent scheduling and optimization of network resources. BRIEF DESCRIPTION OF THE DRAWINGS

[0062] Figure 1 is a schematic flowchart of an artificial intelligence-based communication network resource optimization method according to an embodiment of the present invention;

[0063] Figure 2 is a comparison chart of prediction accuracies under different load conditions corresponding to an artificial intelligence-based communication network resource optimization method according to an embodiment of the present invention;

[0064] Figure 3 is a prediction error data chart under different fluctuation scenarios corresponding to an artificial intelligence-based communication network resource optimization method according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0065] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0066] The technical solutions of the present invention will be described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments.

[0067] Figure 1 is a schematic flowchart of an artificial intelligence-based communication network resource optimization method according to an embodiment of the present invention, as Figure 1 shown, the method includes:

[0068] Collect historical network traffic data, network device performance data, and user service demand data in the communication network. Decompose the historical network traffic data into a trend term, a periodic term, and a residual term through a time series decomposition algorithm, and calculate the bandwidth occupancy rate, processor occupancy rate, memory occupancy rate, and transmission delay in the network device performance data;

[0069] Form an input vector from the trend term, the periodic term, and the residual term, add the input vector to a long short-term memory network, perform gating operations on the input vector through forget gates, input gates, and output gates, update the cell state to obtain a hidden state, and perform a linear transformation on the hidden state to obtain a time series encoding;

[0070] Form a performance metric vector from the bandwidth occupancy rate, processor occupancy rate, memory occupancy rate, and transmission delay, introduce an attention mechanism to calculate the importance weights for the performance metric vector, and perform feature fusion on the time series encoding and the importance weights to obtain a fused feature;

[0071] Process the fused feature using a multi-scale prediction mechanism, predict the network traffic values in the short term, medium term, and long term respectively, set adaptive weights based on the prediction errors at each time scale, and perform weighted combination on the prediction results at multiple time scales to output an initial network traffic prediction value;

[0072] Construct a correction model based on the statistical characteristics of historical prediction errors, dynamically calibrate the initial network traffic prediction value using the exponential smoothing method, and combine the mutation detection results of network performance metrics to correct the prediction value to obtain a network traffic prediction result.

[0073] In an alternative implementation,

[0074] Collect historical network traffic data, network device performance data, and user service demand data in the communication network. Decompose the historical network traffic data into a trend term, a periodic term, and a residual term through a time series decomposition algorithm, and calculate the bandwidth occupancy rate, processor occupancy rate, memory occupancy rate, and transmission delay in the network device performance data, including:

[0075] Collect historical network traffic data in the communication network within a preset sampling period. The historical network traffic data is the number of traffic bytes on the network link per unit time. At the same time, collect network device performance data, which reflects the working state of the network link, and collect user service demand data;

[0076] Input the historical network traffic data into the time series decomposition algorithm, extract the fluctuation law of the historical network traffic data over time, and obtain a trend term reflecting the long-term change trend, a periodic term reflecting the periodic change, and a residual term reflecting the random fluctuation;

[0077] Statistically calculate the performance data of the network device, perform a moving average on the network device performance data according to a preset time window, and obtain the bandwidth occupancy rate, processor occupancy rate, memory occupancy rate, and transmission delay.

[0078] Deploy traffic collection devices in the network, and set the sampling period to 5 minutes. The collector will record the number of traffic bytes passing through the network link at each sampling time point. For example, for a 10Gbps network link, the collector records the actual traffic data on this link every 5 minutes, including the number of bytes and packets in both the upstream and downstream directions. These raw traffic data will be stored in the database in real time to form continuous time series data.

[0079] Collect the performance data of the network device. Access the network device through the SNMP protocol to obtain its operating status parameters. The collector sends an SNMP query request to the device every 5 minutes to read performance metrics such as the CPU usage rate, memory occupancy, and port status of the device. These performance data reflect the current working status and load conditions of the network link.

[0080] Collect the business requirement data of users, including information such as bandwidth application records submitted by users, service types, and quality of service requirements. These data are used to understand the sources and changing rules of network load.

[0081] Process the collected historical network traffic data using the time series decomposition algorithm. First, arrange the continuous traffic data in chronological order, and use the seasonal decomposition method to decompose the data into three components: the trend term reflects the long-term change trend of the traffic, such as whether it is generally increasing or decreasing; the cycle term reflects the periodic change rule of the traffic, such as the daily peaks and valleys, the differences between weekdays and weekends in a week, etc.; the residual term reflects the random short-term fluctuations.

[0082] For the network device performance data, perform statistical calculations using a sliding time window. Set a fixed-size time window (such as 30 minutes), and calculate the average value of each performance metric within the window. The window slides over time, continuously updating the calculation results.

[0083] Exemplarily, monitor a 10Gbps backbone link. The traffic collection device records data every 5 minutes: at 8:00, the traffic is recorded as 3.2Gbps, at 8:05, it is 3.5Gbps, and at 8:10, it is 3.8Gbps. The collected data for 24 consecutive hours shows after time series decomposition: the trend term shows a trend of increasing by 100Mbps per hour; the cycle term shows that the peak periods are from 9:00 to 11:00 and from 14:00 to 16:00 every day, with the traffic reaching 5Gbps, and the low period is from 2:00 to 5:00 in the early morning, with the traffic dropping to 1Gbps; the residual term shows that there are random fluctuations in the traffic of ±500Mbps.

[0084] The performance data statistics results show that within a 30-minute sliding window, the average bandwidth occupancy rate of the link is 35%, the average processor occupancy rate is 45%, the memory occupancy rate is 60%, and the average transmission delay is 15 milliseconds. The user service demand data shows that this link mainly carries the data center synchronization service (2 Gbps) and the video transmission service (3 Gbps).

[0085] In this embodiment, by performing time series decomposition on historical traffic data, the changing pattern of network traffic is accurately grasped, providing an important basis for network planning and optimization, improving the utilization efficiency of network resources. The sliding time window method is used to statistically analyze performance indicators, comprehensively understand the operating status of network devices, timely discover performance bottlenecks, ensure the stable operation of the network. Combining with user service demand data, the on-demand allocation and dynamic adjustment of network resources are realized, improving the quality of user experience and reducing the network operation and maintenance costs.

[0086] In an alternative embodiment

[0087] Form an input vector from the trend term, periodic term, and residual term, add the input vector to a long short-term memory network, and perform gating operations on the input vector through the forget gate, input gate, and output gate to update the cell state to obtain a hidden state, and perform a linear transformation on the hidden state to obtain a time series encoding, including:

[0088] Sequentially splice the trend term, periodic term, and residual term in the time dimension to form an input vector, and the input vector contains historical values at different time steps;

[0089] Based on the input vector and the pre-acquired historical state information, randomly forget historical information through the forget gate, calculate the current update information through the input gate and generate a candidate cell state, combine the retained part of the historical information with the current update information to update the cell state, and perform a gating operation on the cell state through the output gate to obtain a hidden state;

[0090] Perform a linear transformation on the hidden state to obtain a time series encoding.

[0091] Obtain the trend term, periodic term, and residual term. Taking the daily average temperature data in a certain area within one year as an example, through time series decomposition, the trend term reflecting the long-term change trend, the periodic term reflecting the seasonal change, and the residual term reflecting the random fluctuation can be obtained.

[0092] The obtained trend term, periodic term, and residual term are combined into an input vector in chronological order. Taking 30 days as a time window, each time step contains the trend value, periodic value, and residual value corresponding to the date. For example, the input vector on the 1st day contains the trend value of 20 degrees, periodic value of 2 degrees, and residual value of 0.5 degrees on that day; the input vector on the 2nd day contains the trend value of 19.8 degrees, periodic value of 1.8 degrees, and residual value of -0.3 degrees on that day, and so on to construct an input sequence containing 30 time steps.

[0093] Use a long short-term memory network to process the input vector. The network contains three gated units: a forget gate, an input gate, and an output gate. The forget gate calculates the proportion of historical information to be forgotten based on the current input vector and the previous hidden state. Taking temperature prediction as an example, if the current temperature differs greatly from the historical temperature, the forget gate will reduce the retention degree of historical information.

[0094] The input gate calculates the information to be updated at the current moment. Combining the current input vector and the historical hidden state, a candidate cell state is generated. The input gate determines how much new information needs to be written into the memory cell. For example, when a temperature mutation is detected, the input gate will increase the proportion of new information written.

[0095] Combine the historical information filtered by the forget gate with the new information updated by the input gate to update the cell state. The output gate controls the degree of information output based on the updated cell state to obtain the hidden state at the current moment. Finally, perform a linear transformation on the hidden state to obtain an encoded vector containing temporal features.

[0096] Exemplarily, at a certain moment, the value of the trend term is 3.5 Gbps, the value of the periodic term is 1.2 Gbps, and the value of the residual term is -0.3 Gbps. Take the data of the nearest 6 time steps for splicing to form an input vector with a length of 18 (6 time steps × 3 components).

[0097] The input vector enters the LSTM network. The cell state at the previous moment is [0.8, 0.6, 0.4], and the hidden state is [0.7, 0.5, 0.3]. The forget gate calculates the forget coefficient [0.4, 0.3, 0.5], indicating the proportion of historical information retained. The input gate calculates the updated information [0.6, 0.5, 0.4] and generates a candidate cell state [0.9, 0.7, 0.5]. After combination and update, the new cell state is [0.85, 0.65, 0.45]. After the output gate performs a gating operation, the hidden state is [0.75, 0.55, 0.35]. The hidden state is mapped to a two-dimensional temporal encoding [0.65, 0.45] through a linear transformation.

[0098] In this embodiment, multiple characteristic components of data are extracted through time series decomposition, which can capture the variation law of data more comprehensively and improve the accuracy of time series feature extraction. The gating mechanism of the long short-term memory network is adopted, which can adaptively adjust the retention degree of historical information and the update degree of new information, effectively process long-term dependence relationships, combine the trend term, cycle term and residual term for modeling, and perform non-linear feature extraction through a neural network, so that richer time series patterns can be learned and the expression ability of the model can be enhanced.

[0099] In an alternative embodiment,

[0100] The bandwidth occupancy rate, processor occupancy rate, memory occupancy rate and transmission delay are combined to form a performance metric vector. The attention mechanism is introduced to calculate the importance weights for the performance metric vector, and the time series encoding and the importance weights are fused to obtain the fused features, including:

[0101] Collect the performance parameters of the network device at multiple consecutive time points. The performance parameters include the bandwidth occupancy rate, processor occupancy rate, memory occupancy rate and transmission delay, and the performance parameters collected at each time point are combined to form a performance metric vector;

[0102] Based on the attention mechanism, the performance metric vector is input into the query matrix, key matrix and value matrix to obtain the query vector, key vector and value vector. The correlation score between the performance metrics is obtained based on the dot product operation of the query vector and the key vector, and the softmax normalization process is performed on the correlation score to obtain the importance weights, which represent the influence degree of each performance metric on the network performance;

[0103] The time series encoding and the importance weights are weighted and superimposed to obtain the fused features.

[0104] Four performance parameters of the network device are collected at consecutive time points: the bandwidth occupancy rate reflects the link usage situation, the processor occupancy rate represents the device computing load, the memory occupancy rate reflects the resource consumption status, and the transmission delay shows the packet transmission quality. The four parameters collected at each time point are arranged in order to form a performance metric vector, and the sampled data at each time point forms a set of elements in the vector.

[0105] The attention mechanism is adopted to process the performance metric vector. The performance metric vector is respectively input into three transformation matrices: the query matrix is used to extract the target features to be analyzed, the key matrix is used to extract the reference features, and the value matrix is used to extract the actual feature values. After transformation, a query vector, a key vector, and a value vector are obtained. Calculate the dot product of the query vector and the key vector to obtain a correlation score representing the degree of association between different performance metrics. Perform softmax normalization on the correlation score to map the score between 0 and 1, obtaining the importance weights. The importance weights reflect the influence degree of each performance metric on the overall network performance.

[0106] Perform a weighted superposition operation on the obtained temporal encoding and the importance weights. The temporal encoding contains the temporal features of the network traffic, and the importance weights reflect the influence degree of the performance metrics. The fused features obtained by superimposing the two reflect both the temporal variation law and the performance influencing factors.

[0107] Exemplarily, four performance parameters are collected at a certain moment: the bandwidth occupancy rate is 35%, the processor occupancy rate is 45%, the memory occupancy rate is 60%, and the transmission delay is 15 ms. The data of 6 time points are continuously sampled to form a 24-dimensional performance metric vector (6 time points × 4 parameters).

[0108] Input the performance metric vector into the attention mechanism, and obtain through three transformation matrices: an 8-dimensional query vector [0.4, 0.5, 0.3, 0.6, 0.4, 0.2, 0.5, 0.3], an 8-dimensional key vector [0.3, 0.6, 0.4, 0.5, 0.3, 0.4, 0.6, 0.2], and an 8-dimensional value vector [0.5, 0.4, 0.6, 0.3, 0.5, 0.3, 0.4, 0.6]. Calculate the dot product of the query vector and the key vector to obtain the correlation score [0.8, 0.6, 0.7, 0.5]. After softmax normalization, the importance weights [0.35, 0.20, 0.30, 0.15] are obtained, indicating that the bandwidth occupancy rate and the memory occupancy rate have a greater impact on the network performance.

[0109] Perform a weighted superposition on the obtained two-dimensional temporal encoding [0.65, 0.45] and the four-dimensional importance weights [0.35, 0.20, 0.30, 0.15] to obtain the fused features [0.55, 0.40] that reflect the temporal features and performance influence.

[0110] In this embodiment, the importance weights of performance indicators are calculated through the attention mechanism, and the key indicators that have a greater impact on network performance are accurately identified, improving the accuracy of network performance evaluation. The temporal encoding is used to retain the characteristic information of the performance indicators changing over time, enabling the fused features to reflect the dynamic change law of network performance and enhancing the temporal expression ability of the features. The importance weights and the temporal encoding are fused in terms of features, and the obtained fused features have stronger expression ability, which can simultaneously depict the importance degree of performance indicators and the temporal change characteristics, providing more effective feature support for subsequent network performance prediction and fault diagnosis.

[0111] In an alternative implementation manner,

[0112] The multi-scale prediction mechanism is adopted to process the fused features, and the network traffic values in the short term, medium term, and long term are predicted respectively. Adaptive weights are set based on the prediction errors at each time scale, and the prediction results at multiple time scales are weighted and combined to output the initial network traffic prediction value, including:

[0113] The double-layer long short-term memory network structure is used to construct the prediction processing units for the three time scales of short term, medium term, and long term respectively. The double-layer long short-term memory network structure processes the fused features and outputs the initial network traffic prediction values at each time scale;

[0114] The network traffic data at each time scale is continuously collected within a preset sampling time window, the change rate of the network traffic data at each time scale is calculated, and the temporal correlation degree between adjacent time scales and the fluctuation degree of each time scale are calculated according to the change rate; A corresponding temperature coefficient is set for each time scale based on the fluctuation degree, and the greater the fluctuation degree, the smaller the temperature coefficient;

[0115] The historical samples are sliced into multiple consecutive training sequence segments at a preset time interval. The difference between the true value of the network traffic and the initial network traffic prediction value at the corresponding time scale is calculated within each training sequence segment, and the cumulative sum operation is performed on the difference and the average value is calculated to generate the prediction errors at each time scale; The prediction error is multiplied by the temperature coefficient at the corresponding time scale and then the opposite number is taken, and the exponential function is substituted for operation to generate an exponential mapping value. The ratio of the exponential mapping value to the sum of the exponential mapping values of the three time scales is calculated to determine the initial adaptive weights corresponding to each time scale;

[0116] Extract the law characteristics of the temporal correlation degree between each time scale changing with time according to the graph structure network, establish the constraint relationship between the predicted values of adjacent time scales according to the law characteristics, correct the initial network traffic predicted value based on the constraint relationship and the federated learning framework, generate the corrected predicted value considering the temporal correlation, and adjust the initial adaptive weight based on the corrected predicted value to obtain the corrected weight considering the temporal correlation, where the corrected weight corresponding to the time scale with a larger degree of fluctuation is relatively increased;

[0117] Perform a weighted summation operation on the corrected predicted value and the corresponding corrected weight, and output the final network traffic predicted value.

[0118] Construct a two-layer LSTM network structure, and each processing unit of each time scale contains two LSTM layers. The number of hidden units in the first layer of LSTM is set to 128, which is used to extract the temporal patterns in the input features; the number of hidden units in the second layer of LSTM is set to 64, which focuses on the sequence prediction task. The short-term prediction unit processes the data of the most recent 24 hours, the medium-term prediction unit processes the data of the most recent 7 days, and the long-term prediction unit processes the data of the most recent 4 weeks. The input of each processing unit includes the temporal encoding and performance weight information in the fused features. The LSTM network updates the cell state and the hidden state at each time step through forward propagation, and the final output layer uses a fully connected layer to map the hidden state to the predicted value;

[0119] Within a 30-minute sampling window, traffic data is collected every 5 minutes. The short-term scale calculates the change rate between adjacent sampling points: (current traffic - previous traffic) / sampling interval; the medium-term scale calculates the change rate between adjacent hours; the long-term scale calculates the change rate between adjacent days. For the correlation degree between adjacent time scales, the Pearson correlation coefficient is used for calculation: the covariance of the data sequences of the two time scales is calculated after standardization, and then divided by the product of the standard deviations. The degree of fluctuation is obtained by calculating the standard deviation of the traffic sequence within the sliding window, and the window sizes are 2 hours for the short term, 2 days for the medium term, and 2 weeks for the long term. The temperature coefficient is set using the exponential decay formula: exp(-degree of fluctuation);

[0120] The 30-day historical samples are sliced at fixed intervals. For short-term prediction, a 1-hour interval is used to generate 720 training sequence segments; for medium-term prediction, a 6-hour interval is used to generate 120 training sequence segments; for long-term prediction, a 24-hour interval is used to generate 30 training sequence segments. Within each sequence segment, the prediction error is cumulatively calculated: the difference between the predicted value and the true value at each time point is summed, and then divided by the length of the sequence segment to obtain the average error. Multiply the average error by the temperature coefficient and take the negative value, and perform a non-linear mapping through the exp function to obtain the exponential mapping value. The exponential mapping values of the three time scales are added to obtain the normalized denominator, and each exponential mapping value is divided by the denominator to obtain the initial adaptive weight;

[0121] Construct a three-node graph network, where the nodes represent the predicted values at three time scales, and the weights of the edges are the temporal correlation degrees. Each node contains the current predicted value and the characteristics of the historical prediction sequence. The dynamic correlation features between nodes are extracted through the graph attention layer to generate the edge feature matrix. Based on the edge feature matrix, a constraint equation is established: the difference in predicted values between adjacent nodes should be inversely proportional to their correlation degrees. In the federated learning framework, the prediction models at each time scale act as federated members, sharing constraint information but not directly exchanging data. By iteratively optimizing the predicted values that violate the constraints, the corrected predicted values are generated. The calculation of the corrected weights takes into account the original weights and the degree of fluctuation: corrected weight = original weight * (1 + normalized degree of fluctuation);

[0122] Perform a weighted summation operation on the corrected predicted values and corrected weights at the three time scales. The weight values reflect the credibility of the prediction results at each time scale, and the predicted values reflect the corrected results considering the temporal correlation. Weighted summation can balance the prediction biases at different time scales to obtain a more accurate final predicted value.

[0123] Exemplarily, the collected traffic data shows that the standard deviation of the short-term scale is 0.8 Gbps, the medium-term is 0.5 Gbps, and the long-term is 0.3 Gbps. Accordingly, the temperature coefficients are set as: 0.4 for the short-term, 0.6 for the medium-term, and 0.8 for the long-term.

[0124] The double-layer LSTM outputs the initial predicted values at three time scales: 4.2 Gbps for the short-term, 4.0 Gbps for the medium-term, and 3.8 Gbps for the long-term. The actual traffic value is 4.1 Gbps, and the predicted errors are calculated as: 0.1 Gbps for the short-term, -0.1 Gbps for the medium-term, and -0.3 Gbps for the long-term.

[0125] Multiply the predicted errors by the temperature coefficients and take the negative, then substitute them into the exponential function for calculation: exp(-0.04) = 0.96 for the short-term, exp(0.06) = 1.06 for the medium-term, and exp(0.24) = 1.27 for the long-term. After normalization, the initial adaptive weights are obtained: 0.29 for the short-term, 0.32 for the medium-term, and 0.39 for the long-term.

[0126] The graph structure network extracts the short-term - medium-term correlation degree of 0.7 and the medium-term - long-term correlation degree of 0.6. Based on the constraint relationship, the corrected predicted values are: 4.15 Gbps for the short-term, 4.05 Gbps for the medium-term, and 3.9 Gbps for the long-term. The adjusted corrected weights are: 0.35 for the short-term, 0.33 for the medium-term, and 0.32 for the long-term.

[0127] The weighted summation gives the traffic prediction value: 4.15×0.35 + 4.05×0.33 + 3.9×0.32 = 4.04 Gbps.

[0128] In this embodiment, through the multi-scale prediction mechanism, the characteristics of network traffic data at different time scales are fully utilized, improving the prediction accuracy and robustness. The adaptive weight allocation scheme is adopted to dynamically adjust the weights according to the prediction errors and data fluctuation degrees of each time scale, making the prediction results more accurate and reliable. The graph structure network and the federated learning framework are introduced to effectively extract the temporal correlation features and achieve distributed collaborative optimization, enhancing the generalization ability and practicality of the model.

[0129] In an alternative embodiment,

[0130] According to the graph structure network, extract the regular features of the temporal correlation degree changing with time between each time scale, establish the constraint relationship between the predicted values of adjacent time scales according to the regular features, and correct the initial network traffic predicted value based on the constraint relationship and the federated learning framework to generate the corrected predicted value considering the temporal correlation, including:

[0131] Construct a dynamic heterogeneous graph network structure, connect the network traffic nodes of different time scales through dynamic edges, where the node attributes include traffic values, statistical features, and change trends, and the edge attributes are initialized as the basic temporal correlation degree between adjacent time scales; adopt the spatio-temporal attention mechanism to dynamically update the attributes of nodes and edges, and output the updated node state vectors and edge association strengths;

[0132] Based on the updated node state vectors and edge association strengths, construct a temporal correlation degree extraction model, capture the change features of network traffic at different time scales through a sliding time window, and combine the counterfactual reasoning method to identify the dominant factors of the temporal correlation degree, and extract the regular features of the temporal correlation degree changing with time;

[0133] Construct a conditional temporal pattern graph according to the extracted regular features, map the temporal correlation patterns under different conditions to the nodes in the graph, and establish node connections based on the evolution relationship between the patterns. Use the conditional temporal pattern graph to describe the constraint relationship between the predicted values of adjacent time scales, adopt a distributed federated learning framework to integrate the constraint relationship information of multiple network regions, adjust the calculation parameters of the temporal correlation degree in real time based on the local prediction deviation, and continuously optimize the extraction method of the regular features using the global prediction performance index. Construct an adaptive probability graph model, and dynamically associate the update frequency of the model parameters with the fluctuation degree of the network traffic;

[0134] Construct an energy function based on the node state distribution and transition probability in the adaptive probability graph model, combined with the constraint relationships provided by the conditional temporal pattern graph. The constraint relationship strength, conditional temporal pattern, and probability distribution are used as components of the energy function. Among them, the weight coefficients of each component are adaptively adjusted according to the change of prediction accuracy. The initial network traffic prediction value is corrected by iteratively optimizing the energy function until all temporal constraint conditions are met, and the corrected prediction value considering temporal correlation is output.

[0135] Construct a dynamic heterogeneous graph network structure, organize network traffic data at different time scales in the form of a graph network, and construct the network traffic data at each time point as a node in the graph. Each node contains the actual traffic value, statistical feature information, and change trend information at that time point. The statistical features are calculated through a sliding time window, including statistics such as the average value, standard deviation, skewness, and kurtosis of the data distribution within the window range; the change trend information is obtained by analyzing the traffic changes between adjacent time points and is used to characterize the change state of the traffic. For adjacent nodes in time, the system establishes dynamic connection edges, and the initial attributes of the edges are determined by calculating the correlation coefficient of the traffic data at adjacent time points.

[0136] To achieve the dynamic update of node and edge attributes, design a spatio-temporal attention mechanism. Calculate the degree of association between each node and its surrounding nodes in the spatial dimension, and respectively focus on different types of spatial association features through the multi-head attention mechanism. In the time dimension, introduce a temporal attention layer to capture temporal position information, enabling the model to understand the dependence relationship of the data in the time dimension.

[0137] Based on the updated node state and edge association information, construct a temporal correlation extraction model. Adopt a dynamically adjusted sliding window mechanism to analyze the trend, periodicity, and burstiness characteristics of network traffic within each window. Identify the key factors affecting temporal correlation through counterfactual reasoning: The system constructs a control experiment, observes the degree of change in temporal correlation by changing a single factor, and thus evaluates the importance of each factor.

[0138] Construct a conditional temporal pattern graph based on the extracted regular features, define conditional dimensions including network load level, time period, etc., and create corresponding pattern nodes for each conditional combination. Establish the connection relationship between pattern nodes by analyzing the evolution law of patterns in historical data. Adopt a distributed federated learning framework to integrate the constraint relationship information of multiple network regions. Each region maintains a local model and exchanges parameters with the central node regularly to achieve knowledge sharing while protecting data privacy.

[0139] Build an adaptive probability graph model. First, establish a multi-dimensional state space, including a traffic level dimension (divide traffic into multiple load levels), a change trend dimension (indicating the rising, falling, or stable state of traffic), and a time attribute dimension (distinguish weekdays, weekends, holidays, etc.). The core components of the model include a state transition module, an observation mapping module, and an adaptive adjustment module.

[0140] The state transition module is responsible for learning and updating the transition rules between states. By analyzing the change patterns of state sequences in historical data, initial state transition relationships are established. As new data is continuously collected, the system dynamically updates these transition relationships, and the importance weights of old and new data are considered during the update process to ensure that the model can adapt to changes in the network environment.

[0141] The observation mapping module establishes the correspondence between states and actual traffic values, statistically analyzes the distribution of traffic values that may occur in each state, and forms an observation probability mapping. The mapping relationship is continuously optimized as new data accumulates, making the prediction results more accurate.

[0142] The adaptive adjustment module is responsible for dynamic adjustments at three levels: First, the adjustment of the parameter update frequency. The system dynamically changes the update frequency of model parameters according to the degree of traffic fluctuations. Second, the adjustment of state transition probabilities. The transition probabilities are updated through the weighted fusion of old and new data. Third, the optimization of the state space structure. Regularly evaluate and adjust the granularity of state division, merge low-usage states or subdivide high-frequency usage states.

[0143] Build an energy function that integrates multiple types of information, including constraint relation strength, conditional temporal rules, and probability distribution. The weight coefficients are adaptively adjusted according to changes in prediction accuracy. The initial prediction value is corrected by iteratively optimizing the energy function through the gradient descent algorithm until all temporal constraint conditions are met, and a corrected prediction result considering temporal correlation is output. During the optimization process, continuously monitor the prediction accuracy metrics and adjust the model parameters according to the performance feedback.

[0144] Exemplarily, the initial attributes of the three types of nodes at a certain moment are as follows: The short-term node [4.2 Gbps, (4.0, 0.3, 4.5), (0.1, 0.8)] represents the current traffic of 4.2 Gbps, mean value of 4.0, variance of 0.3, peak value of 4.5, growth rate of 0.1, and periodic intensity of 0.8; the medium-term node [4.0 Gbps, (3.8, 0.2, 4.2), (0.05, 0.6)]; the long-term node [3.8 Gbps, (3.6, 0.1, 4.0), (0.02, 0.4)]. The initial edge correlation degree between short-term and medium-term is 0.7, and between medium-term and long-term is 0.6. After being updated by the attention mechanism, the node state vectors become: short-term [4.15, 4.0, 0.12], medium-term [4.05, 3.9, 0.06], long-term [3.85, 3.7, 0.03], and the edge correlation intensity is updated to 0.75 and 0.65.

[0145] Within a 6-hour sliding window, the calculated short-term traffic amplitude is 1.2 Gbps, frequency is 0.2 times per hour, and peak time is 10:00; the medium-term amplitude is 0.8 Gbps, frequency is 0.1 times per hour, and peak is 11:00; the long-term amplitude is 0.5 Gbps, frequency is 0.05 times per hour, and peak is 12:00. Counterfactual analysis shows that a 1 Gbps increase in short-term traffic leads to a 0.6 Gbps change in the medium-term, and a 1 Gbps increase in the medium-term leads to a 0.4 Gbps change in the long-term. The feature importance scores are: amplitude 0.8, frequency 0.6, and phase 0.4.

[0146] The conditional time series patterns include: peak mode (trigger condition: short-term traffic > 4.5 Gbps, correlation intensity 0.85, duration 2 hours), stable mode (3.5 - 4.5 Gbps, intensity 0.7, duration 6 hours), and trough mode (< 3.5 Gbps, intensity 0.6, duration 4 hours). The mode conversion probabilities are: peak to stable 0.8, stable to trough 0.6, and trough to stable 0.7. When the current traffic standard deviation is 0.8 Gbps, the model updates the parameters every 10 minutes.

[0147] During the optimization process of the energy function, the initial weight coefficients are all around 0.33. The current constraint term value is -0.65, the regular term probability is 0.8, and the probability term density is 0.7. After 50 iterations of optimization, the corrected predicted values are obtained: short-term 4.12 Gbps, medium-term 4.08 Gbps, long-term 3.95 Gbps, and the weights are updated to: constraint term 0.35, regular term 0.35, probability term 0.30, and the final energy value drops to -0.2 to meet the convergence condition.

[0148] In this embodiment, by introducing a dynamic heterogeneous graph network structure and a spatio-temporal attention mechanism, the modeling ability of multi-scale temporal correlations is significantly improved. By using counterfactual reasoning and the method of conditional temporal pattern graphs, the depth of understanding of the laws of network traffic changes is enhanced. Based on the design of a federated learning framework and an adaptive probability model, the model's perception and response capabilities to network state changes are improved. Through the optimization mechanism of a unified energy function, precise correction and adaptive adjustment of the prediction results are achieved;

[0149] In the prior art, network traffic prediction usually uses a single-time-scale prediction model or simply combines the prediction results of multiple time scales by weighting, ignoring the complex temporal correlation relationships between different time scales and making it difficult to accurately capture the dynamic evolution characteristics of network traffic at multiple time scales. Most of the modeling of temporal correlations in the prior art is based on static correlation coefficients and cannot adapt to the changes in correlation strength caused by network traffic fluctuations, resulting in a significant decline in the accuracy of prediction results during periods of severe fluctuations. In addition, existing methods often use a fixed weight system for fusing prediction results and lack the adaptive ability to network state changes;

[0150] In this embodiment, by constructing a heterogeneous graph structure containing nodes of multiple time scales and combining a spatio-temporal attention mechanism to dynamically update node attributes and edge correlation strengths, the modeling ability of the temporal characteristics of network traffic is effectively improved. The counterfactual reasoning method is used to analyze the dominant factors of temporal correlations and extract correlation laws, enabling the prediction model to better understand the mutual influence mechanism between different time scales. A conditional temporal pattern graph is constructed to describe the constraint relationship between predicted values of adjacent time scales through a probabilistic graph model. The update frequency of model parameters is dynamically correlated with the degree of traffic fluctuations, ensuring the fast response ability during periods of severe fluctuations. In summary, while maintaining the prediction accuracy, this embodiment significantly improves the model's adaptability to severe fluctuations in network traffic, enhances the temporal consistency of prediction results, better meets the requirements for the accuracy and stability of traffic prediction in the actual network environment, and provides a more reliable decision-making basis for network resource scheduling and performance optimization.

[0151] Figure 2 This is a comparison graph of prediction accuracies under different load conditions corresponding to the communication network resource optimization method based on artificial intelligence according to the embodiments of the present invention, as Figure 2The figure shows the comparison of the prediction accuracies of three different prediction methods under different network load conditions. As the network load changes from low to high, the prediction accuracies of all methods show a downward trend, but the proposed technical solution maintains the highest accuracy under various load conditions. Under low load (30%), the accuracy of the traditional neural network is 82.3%, the graph neural network (GNN) model reaches 86.8%, and the proposed technical solution is as high as 91.2%, which is 4.4 percentage points higher than the graph neural network model and 8.9 percentage points higher than the traditional neural network. Under medium load (60%), the accuracies of the three methods are 77.4% (traditional neural network), 83.2% (graph neural network model), and 88.7% (proposed technical solution) respectively. Under high load (80%), the accuracy further drops to 71.2% (traditional neural network), 79.6% (graph neural network model), and 86.3% (proposed technical solution). The most challenging is the peak load (95%) condition, where the accuracy of the traditional neural network is only 65.1%, the graph neural network model is 74.1%, and the proposed technical solution still remains at a relatively high level of 82.9%. It is worth noting that the proposed technical solution has the smallest performance decay when the load increases. The accuracy only drops by 8.3 percentage points from low load to peak load, while the traditional neural network and the graph neural network model drop by 17.2 and 12.7 percentage points respectively, indicating that the proposed solution has stronger stability and robustness under high load conditions.

[0152] In an alternative embodiment,

[0153] A correction model is constructed based on the statistical characteristics of historical prediction errors, the initial network traffic prediction value is dynamically calibrated using the exponential smoothing method, and the prediction value is corrected in combination with the mutation detection result of the network performance index. The obtained network traffic prediction result includes:

[0154] Obtain the historical prediction error sequence, which is obtained by subtracting the predicted network traffic value from the actual network traffic value. Calculate the error mean and error standard deviation based on the historical prediction error sequence, construct the autocorrelation function of the error sequence, and calculate the exponentially weighted moving variance, where the exponentially weighted moving variance is determined by the weighted sum of the squared difference between the error at the current moment and the mean and the exponentially weighted moving variance at the previous moment;

[0155] Construct a correction model based on the error mean, the error standard deviation, the autocorrelation function, and the exponentially weighted moving variance. The correction model includes an error distribution correction term, a time series correlation correction term, and a fluctuation correction term. The error distribution correction term, the time series correlation correction term, and the fluctuation correction term correspond to different adaptive weight coefficients, and the adaptive weight coefficients are updated by the gradient descent method;

[0156] According to the exponential smoothing method, a historical information storage matrix is constructed. The initial network traffic prediction value is dynamically calibrated through an adaptive smoothing factor and an attention mechanism to obtain a first corrected prediction value. Among them, the adaptive smoothing factor is determined by the weighted sum of the smoothing factor at the previous moment and the absolute value of the historical prediction error at the current moment, and the weighting coefficient is dynamically adjusted through the change amount of the exponentially weighted moving variance;

[0157] A multi-dimensional network performance index vector is constructed. The multi-dimensional network performance index vector contains the values of multiple network performance indexes at the current moment. The cumulative sum operation is performed based on the time series difference of the multi-dimensional network performance index vector to obtain a mutation detection value, and the mutation detection value is compared with the difference between the mutation detection value at the previous moment and the detection threshold;

[0158] When the mutation detection value is greater than the mutation determination threshold, calculate the relative change rate of each network performance index in the multi-dimensional network performance index vector, multiply the relative change rate of each network performance index by the corresponding weight coefficient and sum to obtain the mutation degree, and take the product of the first corrected prediction value and the mutation degree as the network traffic prediction result to output; when the mutation detection value is not greater than the mutation determination threshold, take the first corrected prediction value as the network traffic prediction result to output.

[0159] The historical prediction error sequence is obtained by calculating the difference between the actual network traffic value and the predicted network traffic value. Statistical analysis is performed on the error sequence, the error mean is calculated as a measure of the system prediction deviation, and the error standard deviation is calculated to characterize the prediction fluctuation degree. The autocorrelation function of the error sequence is constructed to analyze the correlation between error values at different time intervals and capture the time series dependence characteristics of the prediction error. The exponentially weighted moving variance is calculated, and the squared difference between the error at the current moment and the mean is weighted and combined with the exponentially weighted moving variance at the previous moment, and the weighting coefficient decays with time to highlight the influence of recent error fluctuations.

[0160] A three-term correction model is constructed based on the error statistical characteristics. The error distribution correction term uses the error mean and standard deviation to characterize the distribution characteristics of the prediction error; the time series correlation correction term models the time series dependence relationship of the error sequence based on the autocorrelation function; the fluctuation correction term reflects the dynamic fluctuation characteristics of the prediction error through the exponentially weighted moving variance. Initial weight coefficients are assigned to the three correction terms, and the gradient descent method is used to dynamically update the weight coefficients based on the improvement degree of the prediction performance, so that the correction effect is adaptively adjusted with the change of the prediction error characteristics.

[0161] Construct a historical information storage matrix to save past prediction data, and apply the exponential smoothing method to calibrate the initial prediction value. The adaptive smoothing factor is determined by the weighted combination of the previous moment's smoothing factor and the absolute value of the current moment's historical prediction error. The weighting coefficient is dynamically adjusted according to the change of the exponentially weighted moving variance. When the error fluctuation intensifies, the weight of the current error is increased. Combine the attention mechanism to calculate the influence degree of historical data on the current prediction, and dynamically calibrate the initial prediction value to obtain the first corrected prediction value.

[0162] Construct a multi-dimensional index vector containing the current values of multiple network performance indicators. Calculate the time series difference of the index vector and perform cumulative summation to obtain the mutation detection value. Compare the detection value with the difference between the previous moment's detection value and the preset threshold to determine whether a mutation has occurred. When the detection value exceeds the mutation determination threshold, calculate the relative change rate of each performance indicator, multiply the change rate by the corresponding weight coefficient and sum to obtain the mutation degree, and adjust the first corrected prediction value with the mutation degree and then output the final prediction result. When the detection value does not exceed the threshold, directly output the first corrected prediction value as the final prediction result.

[0163] Exemplarily, a historical prediction error sequence for the last 24 hours is calculated, with an error mean of 0.2 Gbps and an error standard deviation of 0.5 Gbps. The autocorrelation function shows that the correlation coefficient at an interval of 1 hour is 0.6, 0.4 at 2 hours, and 0.2 at 3 hours. The squared difference between the current moment's error and the mean is 0.09, the previous moment's exponentially weighted moving variance is 0.16, the weighting coefficient is 0.3, and the current exponentially weighted moving variance is calculated to be 0.14.

[0164] The initial weights of the three correction models are all 0.33. After gradient descent update, the weights are: the error distribution correction term 0.35, the time series correlation correction term 0.4, and the fluctuation correction term 0.25. The historical information storage matrix saves the prediction data for the last 12 hours. The previous moment's smoothing factor is 0.7, the absolute value of the current error is 0.3, and the change amount of the exponentially weighted moving variance is -0.02. Based on this, the current smoothing factor is determined to be 0.65.

[0165] The influence weight distribution of historical data is calculated through the attention mechanism, and the initial prediction value of 4.5 Gbps is calibrated to obtain the first corrected prediction value of 4.3 Gbps. The multi-dimensional performance index vector includes: bandwidth occupancy rate 75%, processor occupancy rate 60%, memory occupancy rate 50%, and transmission delay 25 ms. The calculated mutation detection value is 85, which is greater than the mutation determination threshold of 80. The relative change rates of each index are 0.2, 0.15, 0.1, and 0.25 respectively, and the corresponding weights are 0.3, 0.2, 0.2, and 0.3. The calculated mutation degree is 1.2. Finally, the first corrected prediction value of 4.3 Gbps is multiplied by the mutation degree of 1.2, and the prediction result of 5.16 Gbps is output.

[0166] In this embodiment, by comprehensively analyzing the historical prediction error sequence, considering the distribution characteristics, temporal correlation and dynamic fluctuation characteristics of the errors, the exponential smoothing method and the attention mechanism are used to calibrate the predicted value. Through the dynamic update of the adaptive smoothing factor, the influence degree of historical information can be adjusted in real time according to the fluctuation condition of the prediction error, improving the response speed of the model to the change of network state, enhancing the timeliness of the prediction result, calculating the mutation degree based on the relative change rate of each performance index, making targeted adjustment to the prediction result, improving the prediction accuracy of the model during the period of drastic change of network state, and realizing the precise correction of the network traffic prediction result, providing an effective solution for improving the accuracy and reliability of network traffic prediction.

[0167] In an alternative embodiment,

[0168] According to the exponential smoothing method, constructing a historical information storage matrix, and dynamically calibrating the initial network traffic prediction value through the adaptive smoothing factor and the attention mechanism to obtain the first corrected prediction value, including:

[0169] The network traffic time series is divided into multiple time scale layers according to the time scale, and an independent adaptive smoothing factor is constructed for each time scale layer;

[0170] Calculating the volatility according to the change trend of the network traffic time series, dynamically adjusting the historical data window size based on the volatility, where the historical data window size is determined by the product of the basic window size and the volatility, and obtaining the smoothing results of each time scale layer within the historical data window range;

[0171] Constructing a feature matrix from the smoothing results according to the time scale layer, calculating the attention weight based on the feature matrix and the query vector at the current moment, where the attention weight represents the importance degree of different time scale features to the current prediction, and adaptively fusing the features of different time scales by using the attention weight;

[0172] Constructing a historical information storage matrix, which contains historical network traffic values, the prediction error and the adaptive smoothing factor, extracting historical information related to the current state from the historical information storage matrix by using the attention mechanism, and determining the dynamic learning rate based on the change amount of the prediction error, where the dynamic learning rate decays as the change amount of the prediction error increases;

[0173] Introduce a residual connection structure to the initial network traffic prediction value, and perform a weighted sum of the output of the correction function for each time scale and the adaptive smoothing factor according to the attention weight to obtain a residual term. Add the residual term to the initial network traffic prediction value to obtain a first corrected prediction value.

[0174] Perform a multi-level time scale division on the network traffic time series, dividing it into three scale layers: hourly, daily, and weekly. An adaptive smoothing factor is independently configured for each time scale layer to smooth the historical data corresponding to the scale. The initial value of the smoothing factor is set according to the span of the time scale. The larger the scale span, the smaller the initial smoothing factor, to reflect the degree of dependence of different time scales on historical data.

[0175] Calculate the volatility of the network traffic time series, which is measured by the ratio of the change amplitude of the traffic values at adjacent time points to the mean. Dynamically adjust the historical data window size based on the calculated volatility. The window size is equal to the product of a preset base window size and the current volatility. When the traffic fluctuates greatly, increase the window to obtain more historical information; when the fluctuation is small, decrease the window to highlight the influence of recent data. Within the determined window range, use the adaptive smoothing factors of each time scale layer to smooth the historical data to obtain the smoothed results of different time scales.

[0176] Organize the smoothed results of each time scale into a feature matrix, where the rows of the matrix represent different time scales and the columns represent time steps. Construct a query vector for the current moment, which includes the current traffic value and recent change trend information. Calculate the attention weight based on the query vector and the feature matrix. Adopt the scaled dot-product attention mechanism to calculate the similarity between the query vector and each time scale feature in the feature matrix, and normalize it through the softmax function to obtain the attention weight. The weight value reflects the importance of different time scale features for the current prediction.

[0177] Construct a historical information storage matrix to save historical traffic values, prediction errors, and adaptive smoothing factors. Use the attention mechanism to calculate the similarity between the current state and the historical state, and extract relevant historical information. Determine the dynamic learning rate according to the change amount of the prediction error. When the change amount of the prediction error increases, reduce the learning rate to reduce the correction amplitude and improve the stability of the prediction.

[0178] Introduce a residual connection structure to correct the initial prediction value. First, calculate the output of the correction function according to the features of each time scale, multiply the correction output by the corresponding adaptive smoothing factor, and then perform a weighted sum with the attention weight to obtain a residual term. Finally, add the residual term to the initial prediction value to obtain a corrected prediction value considering the features of multiple time scales.

[0179] Exemplarily, the network traffic time series is divided into three scale levels: hourly, daily, and weekly, and the initial smoothing factors are set to 0.8, 0.6, and 0.4 respectively. The current volatility is calculated to be 1.5, and the base window size is set to 24 hours, resulting in an actual window size of 36 hours. The three time scales are smoothed within the 36-hour window. The hourly smoothing result is [4.2, 4.0, 3.8,...] Gbps, the daily is [4.1, 3.9, 3.7,...] Gbps, and the weekly is [4.0, 3.8, 3.6,...] Gbps.

[0180] A feature matrix is constructed with a size of 3×36, representing the smoothed values of the three time scales at 36 time steps. The current query vector is [4.5, 0.2, 0.1], representing the current traffic of 4.5 Gbps, the recent growth rate of 0.2, and the acceleration of 0.1. The attention weights are calculated as: hourly 0.5, daily 0.3, and weekly 0.2.

[0181] The historical information storage matrix records the data of the last 72 hours, including the traffic value, prediction error, and smoothing factor at each time point. The current change in the prediction error is 0.3 Gbps, and based on this, the dynamic learning rate is calculated to be 0.6. Historical information with a high similarity to the current state is extracted for reference.

[0182] The initial prediction value is 4.8 Gbps, and the correction function outputs for each time scale are -0.3, -0.2, and -0.1 respectively. The correction outputs are multiplied by the corresponding smoothing factors: -0.24, -0.12, and -0.04, and then weighted and summed with the attention weights to obtain the residual term -0.16. Finally, the residual term is added to the initial prediction value to obtain the corrected prediction value of 4.64 Gbps.

[0183] In this embodiment, through the multi-layer division of the time series and the configuration of independent smoothing factors, the differential processing of features at different time scales is realized. Based on the historical data window dynamically adjusted by volatility, the data sampling range can adaptively expand and contract according to the traffic change situation, improving the pertinence and effectiveness of historical data utilization. The attention mechanism is used to adaptively fuse the features of multiple time scales. By calculating the correlation degree between the features of different scales and the current state, the dynamic evaluation and selective fusion of feature importance are realized. The historical information storage matrix and the dynamic learning rate mechanism are introduced, enabling the correction process to make full use of historical experience and flexibly adjust the correction intensity according to the change in the prediction error;

[0184] Existing network traffic prediction correction methods usually process historical data using fixed time windows and unified smoothing factors, and are unable to adapt to the dynamic change characteristics of network traffic at different time scales. At the same time, they often simply superimpose or average the characteristics of different time scales, ignoring the differences in the importance of the characteristics of each time scale for the current prediction. In addition, the correction parameters in the existing technologies usually adopt fixed configurations and lack the ability to adaptively adjust to changes in network states, resulting in poor correction effects during periods of drastic traffic fluctuations.

[0185] In this embodiment, through the dynamic fusion of multi-scale features and adaptive parameter adjustment, the correction process can better adapt to the change characteristics of network traffic. The introduced dynamic window mechanism and residual connection structure effectively improve the response ability to abnormal fluctuations and correction stability. While ensuring the correction effect, it significantly enhances the self-adaptability and robustness of the correction process, can better meet the requirements for traffic prediction accuracy in the actual network environment, and provides more reliable technical support for the efficient scheduling and optimized management of network resources.

[0186] Figure 3 It is the prediction error data graph corresponding to the communication network resource optimization method based on artificial intelligence in different fluctuation scenarios in the embodiment of the present invention. As Figure 3 shown, it presents the comparison of the mean absolute percentage error of each prediction method in different fluctuation scenarios. The data clearly shows that this technical solution exhibits the lowest prediction error in all fluctuation scenarios. In the low-fluctuation scenario (volatility < 0.5), the mean absolute percentage error of this technical solution is 3.42%, which is 41.7% lower than 5.87% of the traditional exponential smoothing method and 30.5% lower than 4.92% of the SARIMA model. As the traffic volatility increases, the advantages of this technical solution become more obvious. Especially in the high-fluctuation scenario (1.5 ≤ volatility < 3.0), the mean absolute percentage error of this technical solution is 8.76%, which is 31.9% lower than 12.87% of the GRU network and 23.3% lower than 11.42% of the TCN model. In the extreme-fluctuation scenario (volatility ≥ 3.0), the mean absolute percentage error of this technical solution is 17.52%, which is 46.6% lower than 32.78% of the exponential smoothing method and 24.1% lower than 23.08% of the closest TCN model. Judging from the mean absolute percentage error, this technical solution is 8.64%, which is 22.0% lower than the best-performing TCN model (11.08%) in the existing technologies. This fully demonstrates the effectiveness of the adaptive smoothing factor and dynamic window adjustment mechanism in this technical solution when dealing with different volatility scenarios, especially having significant prediction stability advantages in high-fluctuation and extreme-fluctuation scenarios.

[0187] A communication network resource optimization system based on artificial intelligence, comprising:

[0188] The first unit is used to collect historical network traffic data, network device performance data, and user service demand data in the communication network. The historical network traffic data is decomposed into a trend term, a periodic term, and a residual term through a time series decomposition algorithm, and the bandwidth occupancy rate, processor occupancy rate, memory occupancy rate, and transmission delay in the network device performance data are calculated;

[0189] The second unit is used to form an input vector from the trend term, the periodic term, and the residual term, add the input vector to a long short-term memory network, perform gating operations on the input vector through a forget gate, an input gate, and an output gate, update the unit state to obtain a hidden state, and perform a linear transformation on the hidden state to obtain a time series encoding;

[0190] The third unit is used to form a performance metric vector from the bandwidth occupancy rate, the processor occupancy rate, the memory occupancy rate, and the transmission delay, introduce an attention mechanism to calculate the importance weights for the performance metric vector, and perform feature fusion on the time series encoding and the importance weights to obtain a fusion feature;

[0191] The fourth unit is used to process the fusion feature by adopting a multi-scale prediction mechanism, respectively predict the network traffic values in the short term, medium term, and long term, set adaptive weights based on the prediction errors at each time scale, and perform weighted combination on the prediction results at multiple time scales to output an initial network traffic prediction value;

[0192] The fifth unit is used to construct a correction model according to the statistical characteristics of the historical prediction errors, dynamically calibrate the initial network traffic prediction value by using the exponential smoothing method, and combine the mutation detection results of the network performance metrics to correct the prediction value to obtain the network traffic prediction result.

[0193] The present invention can be a method, an apparatus, a system, and / or a computer program product. The computer program product may include a computer-readable storage medium having thereon computer-readable program instructions for performing various aspects of the present invention.

[0194] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A communication network resource optimization method based on artificial intelligence, characterized in that: include: Collect historical network traffic data, network equipment performance data and user service demand data in the communication network, decompose the historical network traffic data into trend items, period items and residual items through the time series decomposition algorithm, and calculate the bandwidth occupancy rate, processor occupancy rate, memory occupancy rate and transmission delay in the network equipment performance data; The trend term, the cycle term and the residual term form an input vector, the input vector is added to the long short-term memory network, the input vector is gated by a forget gate, an input gate and an output gate, the unit state is updated to obtain a hidden state, and the hidden state is linearly transformed to obtain a temporal code; Bandwidth occupancy, processor occupancy, memory occupancy and transmission delay are combined into a performance indicator vector. The attention mechanism is introduced to calculate the performance indicator vector to obtain the importance weight. The time series coding and the importance weight are fused to obtain the fused feature. A multi-scale prediction mechanism is used to process the fusion features, and the short-term, medium-term and long-term network traffic values ​​are predicted respectively. Adaptive weights are set based on the prediction errors at each time scale, and the prediction results of multiple time scales are weightedly combined to output the initial network traffic prediction value. A correction model is constructed based on the statistical characteristics of historical prediction errors. The initial network traffic prediction value is dynamically calibrated using the exponential smoothing method. The prediction value is corrected based on the mutation detection results of network performance indicators to obtain the network traffic prediction result.

2. The method according to claim 1, characterized in that Collect historical network traffic data, network equipment performance data and user service demand data in the communication network, decompose the historical network traffic data into trend items, period items and residual items through the time series decomposition algorithm, and calculate the bandwidth occupancy rate, processor occupancy rate, memory occupancy rate and transmission delay in the network equipment performance data, including: Collect historical network traffic data in the communication network within a preset sampling period, the historical network traffic data being the number of traffic bytes on the network link per unit time, and network equipment performance data, the network equipment performance data reflecting the working status of the network link, and user service demand data; Inputting the historical network traffic data into a time series decomposition algorithm, extracting the fluctuation pattern of the historical network traffic data over time, and obtaining a trend term reflecting a long-term change trend, a periodic term reflecting a periodic change, and a residual term reflecting a random fluctuation; The network device performance data is statistically calculated and the network device performance data is slidingly averaged according to a preset time window to obtain bandwidth occupancy, processor occupancy, memory occupancy and transmission delay.

3. The method according to claim 1, characterized in that The trend term, the cycle term and the residual term form an input vector, the input vector is added to the long short-term memory network, the input vector is gated by the forget gate, the input gate and the output gate, the unit state is updated to obtain the hidden state, and the hidden state is linearly transformed to obtain the temporal encoding, including: The trend term, the period term and the residual term are sequentially concatenated in the time dimension to form an input vector, wherein the input vector includes historical values ​​at different time steps; Based on the input vector and the pre-acquired historical state information, randomly forget the historical information through a forget gate, calculate the current update information through an input gate and generate a candidate unit state, combine the retained part of the historical information with the current update information to update the unit state, and perform a gated operation on the unit state through an output gate to obtain a hidden state; The hidden state is linearly transformed to obtain a temporal code.

4. The method according to claim 1, characterized in that: Bandwidth occupancy, processor occupancy, memory occupancy and transmission delay are combined into a performance indicator vector. The attention mechanism is introduced to calculate the performance indicator vector to obtain the importance weight. The time series coding and the importance weight are fused to obtain the fusion features including: Collecting performance parameters of the network device at multiple consecutive time points, the performance parameters including bandwidth occupancy, processor occupancy, memory occupancy and transmission delay, and forming a performance indicator vector with the performance parameters collected at each time point; Based on the attention mechanism, the performance indicator vector is input into the query matrix, the key matrix and the value matrix to obtain the query vector, the key vector and the value vector. The correlation score between the performance indicators is obtained based on the dot product operation of the query vector and the key vector. The correlation score is subjected to softmax normalization processing to obtain the importance weight, which represents the influence of each performance indicator on the network performance. The temporal coding and the importance weight are weighted and superimposed to obtain a fusion feature.

5. The method according to claim 1, characterized in that A multi-scale prediction mechanism is used to process the fusion features, and the short-term, medium-term and long-term network traffic values ​​are predicted respectively. Adaptive weights are set based on the prediction errors at each time scale, and the prediction results of multiple time scales are weighted and combined to output the initial network traffic prediction values, including: A double-layer long short-term memory network structure is used to construct prediction processing units for three time scales: short-term, medium-term and long-term. The double-layer long short-term memory network structure processes the fusion features and outputs the initial network traffic prediction value at each time scale. Continuously collect network traffic data of each time scale within a preset sampling time window, calculate the change rate of the network traffic data at each time scale, and calculate the temporal correlation between adjacent time scales and the fluctuation degree of each time scale according to the change rate; set a corresponding temperature coefficient for each time scale based on the fluctuation degree, and the greater the fluctuation degree, the smaller the temperature coefficient; Divide the historical samples into a plurality of continuous training sequence segments according to a preset time interval, calculate the difference between the actual value of the network traffic and the initial network traffic prediction value at the corresponding time scale in each training sequence segment, perform cumulative summation operation on the difference and calculate the average value to generate the prediction error at each time scale; multiply the prediction error by the temperature coefficient of the corresponding time scale, take the inverse number, substitute it into the exponential function for operation to generate the exponential mapping value, calculate the ratio of the exponential mapping value to the sum of the exponential mapping values ​​of the three time scales, and determine the initial adaptive weight corresponding to each time scale; According to the graph structure network, the regular characteristics of the temporal correlation between each time scale changing with time are extracted, and the constraint relationship between the prediction values ​​of adjacent time scales is established according to the regular characteristics. The initial network traffic prediction value is corrected based on the constraint relationship and the federated learning framework to generate a corrected prediction value considering the temporal correlation, and the initial adaptive weight is adjusted based on the corrected prediction value to obtain a corrected weight considering the temporal correlation, wherein the corrected weight corresponding to the time scale with a larger degree of fluctuation is relatively increased; The corrected prediction value is weighted and summed with the corresponding corrected weight to output the final network traffic prediction value.

6. The method according to claim 5, characterized in that According to the graph structure network, the regular characteristics of the temporal correlation between each time scale over time are extracted, and the constraint relationship between the prediction values ​​of adjacent time scales is established according to the regular characteristics. Based on the constraint relationship and the federated learning framework, the initial network traffic prediction value is corrected, and the corrected prediction value considering the temporal correlation is generated, including: A dynamic heterogeneous graph network structure is constructed to connect network traffic nodes of different time scales through dynamic edges, where node attributes include traffic value, statistical characteristics and change trend, and edge attributes are initialized as the basic temporal correlation between adjacent time scales; the spatiotemporal attention mechanism is used to dynamically update the attributes of nodes and edges, and the updated node state vector and edge correlation strength are output; Based on the updated node state vector and edge correlation strength, a temporal correlation extraction model is constructed to capture the changing characteristics of network traffic at different time scales through a sliding time window, and the dominant factors of temporal correlation are identified by combining the counterfactual reasoning method to extract the regular characteristics of the temporal correlation changing over time. A conditional time series pattern graph is constructed based on the extracted regular features, the time series association patterns under different conditions are mapped into nodes in the graph, and node connections are established based on the evolutionary relationship between the patterns. The conditional time series pattern graph is used to characterize the constraint relationship between the prediction values ​​of adjacent time scales, and a distributed federated learning framework is used to integrate the constraint relationship information of multiple network areas. The calculation parameters of the time series correlation degree are adjusted in real time based on the local prediction deviation, and the global prediction performance index is used to continuously optimize the extraction method of regular features, and an adaptive probabilistic graph model is constructed to dynamically associate the update frequency of the model parameters with the fluctuation degree of the network traffic. According to the node state distribution and transition probability in the adaptive probabilistic graph model, an energy function is constructed in combination with the constraint relationship provided by the conditional timing pattern graph. The constraint relationship strength, conditional timing law and probability distribution are used as components of the energy function. The weight coefficient of each component is adaptively adjusted with the change of prediction accuracy. The initial network traffic prediction value is corrected by iteratively optimizing the energy function until all timing constraints are met, and the corrected prediction value considering timing correlation is output.

7. The method according to claim 1, characterized in that According to the statistical characteristics of historical prediction errors, a correction model is constructed, and the initial network traffic prediction value is dynamically calibrated using the exponential smoothing method. The prediction value is corrected in combination with the mutation detection results of network performance indicators. The network traffic prediction results include: Obtain a historical prediction error sequence, wherein the historical prediction error sequence is obtained by subtracting the predicted network traffic value from the actual network traffic value, calculate the error mean and error standard deviation based on the historical prediction error sequence, construct an autocorrelation function of the error sequence, and calculate an exponentially weighted moving variance, wherein the exponentially weighted moving variance is determined by the weighted sum of the square difference between the error and the mean at the current moment and the exponentially weighted moving variance at the previous moment; A correction model is constructed based on the error mean, the error standard deviation, the autocorrelation function and the exponentially weighted moving variance, wherein the correction model includes an error distribution correction term, a time series correlation correction term and a fluctuation correction term, wherein the error distribution correction term, the time series correlation correction term and the fluctuation correction term correspond to different adaptive weight coefficients respectively, and the adaptive weight coefficients are updated by a gradient descent method; According to the exponential smoothing method, a historical information storage matrix is ​​constructed, and the initial network traffic prediction value is dynamically calibrated by an adaptive smoothing factor and an attention mechanism to obtain a first revised prediction value, wherein the adaptive smoothing factor is determined by the weighted sum of the smoothing factor at the previous moment and the absolute value of the historical prediction error at the current moment, and the weighting coefficient is dynamically adjusted by the change in the exponential weighted moving variance; Constructing a multidimensional network performance indicator vector, the multidimensional network performance indicator vector includes the values ​​of multiple network performance indicators at the current moment, performing a cumulative sum operation based on the time series difference of the multidimensional network performance indicator vector to obtain a mutation detection value, and comparing the mutation detection value with the difference between the mutation detection value at the previous moment and the detection threshold; When the mutation detection value is greater than the mutation determination threshold, the relative change rate of each network performance indicator in the multidimensional network performance indicator vector is calculated, the relative change rate of each network performance indicator is multiplied by the corresponding weight coefficient and the sum is obtained to obtain the mutation degree, and the product of the first revised prediction value and the mutation degree is output as the network traffic prediction result; when the mutation detection value is not greater than the mutation determination threshold, the first revised prediction value is output as the network traffic prediction result.

8. The method according to claim 7, characterized in that According to the exponential smoothing method, a historical information storage matrix is ​​constructed, and the initial network traffic prediction value is dynamically calibrated through the adaptive smoothing factor and attention mechanism to obtain the first revised prediction value, including: The network traffic time series is divided into multiple layers according to the time scale to obtain multiple time scale layers, and an independent adaptive smoothing factor is constructed for each time scale layer; Calculate the volatility according to the change trend of the network traffic time series, dynamically adjust the size of the historical data window based on the volatility, the size of the historical data window is determined by the product of the basic window size and the volatility, and obtain the smoothing results of each time scale layer within the historical data window; Constructing a feature matrix according to the time scale layer of the smoothing result, calculating an attention weight based on the feature matrix and the query vector at the current moment, wherein the attention weight represents the importance of features of different time scales to the current prediction, and using the attention weight to adaptively fuse features of different time scales; Constructing a historical information storage matrix, the historical information storage matrix includes historical network traffic values, the prediction error and the adaptive smoothing factor, using an attention mechanism to extract historical information related to the current state from the historical information storage matrix, and determining a dynamic learning rate based on a change in the prediction error, wherein the dynamic learning rate decays as the change in the prediction error increases; A residual connection structure is introduced into the initial network traffic prediction value, and the residual term is obtained by weighted summing the correction function output of each time scale and the adaptive smoothing factor according to the attention weight, and the residual term is added to the initial network traffic prediction value to obtain a first corrected prediction value.

9. A communication network resource optimization system based on artificial intelligence, used to implement the method according to any one of claims 1 to 7, characterized in that: include: The first unit is used to collect historical network traffic data, network equipment performance data and user service demand data in the communication network, decompose the historical network traffic data into trend items, period items and residual items through the time series decomposition algorithm, and calculate the bandwidth occupancy rate, processor occupancy rate, memory occupancy rate and transmission delay in the network equipment performance data; The second unit is used to form an input vector by combining the trend term, the cycle term and the residual term, add the input vector to the long short-term memory network, perform a gating operation on the input vector through a forget gate, an input gate and an output gate, update the unit state to obtain a hidden state, and perform a linear transformation on the hidden state to obtain a temporal code; The third unit is used to combine bandwidth occupancy, processor occupancy, memory occupancy and transmission delay into a performance indicator vector, introduce an attention mechanism to calculate the performance indicator vector to obtain importance weights, and perform feature fusion of the time series coding and the importance weights to obtain fusion features; The fourth unit is used to process the fusion features using a multi-scale prediction mechanism to predict the short-term, medium-term and long-term network traffic values ​​respectively, set adaptive weights based on the prediction errors at each time scale, and perform weighted combination of the prediction results at multiple time scales to output the initial network traffic prediction value; The fifth unit is used to build a correction model based on the statistical characteristics of historical prediction errors, dynamically calibrate the initial network traffic prediction value using the exponential smoothing method, and correct the prediction value in combination with the mutation detection results of the network performance indicators to obtain the network traffic prediction result.

Citation Information

Patent Citations

  • Adaptive link switching method and system based on context awareness, equipment and medium

    CN116842440A

  • Resource Allocation Control Based on Connected Devices

    US20190334805A1

Cited By

  • Intelligent factory monitoring method and system based on multi-sensor fusion

    CN120469321A

  • Computing task scheduling method and system based on time sequence diagram network resource state prediction

    CN120492131A

  • Big data analysis system and method based on artificial intelligence

    CN120596551A

  • Communication flow distribution method and system based on network load awareness

    CN120785889A

  • Network load-aware communication traffic distribution method and system

    CN120785889B