Training Method and Device for Network Delay Prediction Model and Network State Detection Method

By using the LSTM module of the multi-head attention module in the network delay prediction model and adding network characteristic constraint loss, the problem of insufficient delay prediction accuracy in the existing technology in complex network environments is solved, and a higher precision delay prediction is achieved.

CN120034449BActive Publication Date: 2025-06-20CIVIL AVIATION UNIV OF CHINA
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510519635.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-24
Publication Date
2025-06-20
Estimated Expiration
2045-04-24

AI Technical Summary

Technical Problem

The prior art has a large estimation error in delay prediction under low signal-to-noise ratio or homofrequency interference conditions, making it difficult to achieve high-precision delay prediction in complex network environments.

Method used

The LSTM module with a multi-head attention module is used as the network delay prediction model, and the constraint loss based on network load, burst traffic and protocol reception window is added during the model training process to improve prediction accuracy.

Benefits of technology

Through this method, the impact of complex network activities on prediction accuracy is alleviated, making the accuracy of delay prediction more accurate, and better optimizing data scheduling and improving service quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120034449B_ABST
    Figure CN120034449B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of computer technology applications, and particularly to a training method and device for a network delay prediction model and a network state detection method, including: constructing an initial network delay prediction model; obtaining a historical time series data set, which is a data set obtained by sorting multiple network traffic data of a target network link collected within a set time period according to time series; using a sliding time window to divide the historical time series data set into multiple sample time series data; using the numerical values of network transmission parameters in the sample time series data as input information and the numerical values of network delays in the sample time series data as label information, and training the initial network delay prediction model with multiple sample time series data to obtain a trained network delay prediction model. Among them, during the training process, the total loss of the network delay prediction model includes a data fitting loss and a network characteristic constraint loss. The present invention can improve the accuracy of network delay prediction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology applications, and particularly to a method for training a network delay prediction model, an apparatus, and a network status detection method. Background Art

[0002] Delay prediction refers to a technology for estimating the time delay in the data transmission process through algorithms or models. In a communication network, the delay is affected by factors such as noise, interference, and bandwidth fluctuations. Accurately predicting the delay can optimize data scheduling and improve service quality. For example, in a civil aviation data communication network, it is necessary to ensure smooth communication between the data airport and the data center and avoid unnecessary losses. The core of delay prediction lies in dynamically modeling and predicting the delay by analyzing historical data and real-time signal characteristics through mathematical modeling or intelligent learning methods, and solving the estimation error problem under low signal-to-noise ratio or co-frequency interference. Patent document (CN113992599B) discloses a method and apparatus for training a delay prediction model and a congestion control method and apparatus. This document uses an encoding neural network and a decoding neural network for delay prediction, and uses the difference between the delay prediction value and the true delay value as a loss function. The solution disclosed in this patent document can improve the accuracy of delay prediction to a certain extent. However, another prediction solution that can accurately perform delay prediction can be provided. Summary of the Invention

[0003] For the above technical problems, the technical solution adopted by the present invention is as follows:

[0004] According to a first aspect of the present invention, there is provided a method for training a network delay prediction model, the method comprising the following steps:

[0005] S100, construct an initial network delay prediction model, the network delay prediction model including a convolutional module, a first dropout layer, a first feature extraction network, a second dropout layer, a second feature extraction network, and an output prediction layer connected in sequence, wherein the first feature extraction network includes a first LSTM module and a first multi-head attention module, and the second feature extraction network includes a second LSTM module and a second multi-head attention module.

[0006] S200, obtain a historical time series data set, the historical time series data set being a data set obtained by sorting multiple network traffic data of a target network link collected within a set time period according to time series, and the data entities of the network traffic data including network transmission parameters and network delay;

[0007] S300, divide the historical time series data set into multiple sample time series data by using a sliding time window with a sliding stride of 1 and a window size of w, each sample time series data being a matrix of w×m, where m is the number of data entities of the network traffic data.

[0008] The S400 uses the values of the network transmission parameters in the sample time-series data as input information, and the values of the network delay in the sample time-series data as label information, and trains the initial network delay prediction model using the multiple sample time-series data to obtain a trained network delay prediction model. Wherein, during the training process, the total loss of the network delay prediction model includes a data fitting loss and a network characteristic constraint loss, and the network characteristic constraint loss includes a load constraint loss determined based on network load, an attenuation constraint loss determined based on network burst traffic, and a window constraint loss determined based on the receive window of the network protocol.

[0009] According to a second aspect of the present invention, there is provided a training device for a network delay prediction model, including:

[0010] A model construction module, configured to construct an initial network delay prediction model, where the network delay prediction model includes a convolutional module, a first dropout layer, a first feature extraction network, a second dropout layer, a second feature extraction network, and an output prediction layer connected in sequence. Among them, the first feature extraction network includes a first LSTM module and a first multi-head attention module, and the second feature extraction network includes a second LSTM module and a second multi-head attention module.

[0011] A first data acquisition module, configured to acquire a historical time-series data set, where the historical time-series data set is a data set obtained by sorting multiple network traffic data of a target network link collected within a set time period in time sequence, and the data entities of the network traffic data include network transmission parameters and network delay.

[0012] A second data acquisition module, configured to divide the historical time-series data set into multiple sample time-series data by using a sliding time window with a sliding step of 1 and a window size of w. Each time-series data is a matrix of w×m, and m is the number of data entities of the network traffic data.

[0013] A model training module, configured to use the values of the network transmission parameters in the sample time-series data as input information, and the values of the network delay in the sample time-series data as label information, and train the initial network delay prediction model using the multiple sample time-series data to obtain a trained network delay prediction model. Wherein, during the training process, the total loss of the network delay prediction model includes a data fitting loss and a network characteristic constraint loss, and the network characteristic constraint loss includes a load constraint loss determined based on network load, an attenuation constraint loss determined based on network burst traffic, and a window constraint loss determined based on the receive window of the network protocol.

[0014] According to a third aspect of the present invention, there is provided a network status detection method, and the method includes the following steps:

[0015] S10. Obtain the network traffic data of the target network link collected at the current monitoring moment as the current data to be processed.

[0016] S20. Input the numerical values of the network transmission parameters in the current data to be processed into the trained network delay prediction model obtained in the first aspect of the present invention to obtain the network delay prediction value corresponding to the current data to be processed.

[0017] S30. Compare the network delay prediction value corresponding to the current data to be processed with the corresponding actual network delay value, and determine the network state of the target network link at the current monitoring moment based on the comparison result.

[0018] The present invention has at least the following beneficial effects:

[0019] In the training method of the network delay prediction model provided by the embodiment of the present invention, since the LSTM module integrated with the multi-head attention module is used as the network delay prediction model, and during the model training process, the load constraint loss determined based on the network load, the attenuation constraint loss determined based on the network burst traffic, and the window constraint loss determined for the receive window based on the network protocol are added as the losses of the model, it can alleviate the influence of complex network activities on the prediction accuracy and make the predicted delay accuracy more accurate.

[0020] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present invention, nor is it used to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0022] Figure 1 It is a flowchart of the training method of the network delay prediction model provided by the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0023] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0024] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the technical field to which this invention belongs. The terms used in the description of the present invention herein are for the purpose of describing specific embodiments only and are not intended to limit the present invention. The term "and / or" used herein includes any and all combinations of one or more of the associated listed items.

[0025] It should be noted that some exemplary embodiments are described as processes or methods depicted as flowcharts. Although the flowcharts depict the steps as sequential processes, many of the steps can be implemented in parallel, concurrently, or simultaneously. In addition, the order of the steps can be rearranged. The process can be terminated when its operations are completed, but it can also have additional steps not included in the figures. The process can correspond to a method, function, procedure, subroutine, subprogram, etc.

[0026] An embodiment of the present invention provides a method for training a network delay prediction model, as Figure 1 shown, the method may include the following steps:

[0027] S100, construct an initial network delay prediction model.

[0028] In an embodiment of the present invention, the network delay prediction model may include a convolution module, a first dropout layer, a first feature extraction network, a second dropout layer, a second feature extraction network, and an output prediction layer connected in sequence. Among them, the first feature extraction network includes a first LSTM module and a first multi-head attention module, and the second feature extraction network includes a second LSTM module and a second multi-head attention module.

[0029] Among them, the convolution module is used to perform data size normalization operations on the input sequence data so that the sizes of all data are the same. The convolution module may include multiple convolutional layers, and the specific number may be determined based on actual situations. In a schematic embodiment, it may include 64 convolutions and a convolutional layer with a size of 32.

[0030] The first dropout layer is used to apply random inactivation to the features obtained by the convolution module, while maintaining the consistency of the time series feature dimension, suppressing the redundant features extracted by the convolutional layer. The dropout probability of the first dropout layer can be set to 0.3. The second dropout layer is used to apply random inactivation to the output of the attention module to prevent overfitting of the model. The dropout probability of the second dropout layer can be set to 0.3.

[0031] Each LSTM module may include multiple LSTM units connected in series, and the specific number of units can be determined based on experiments, for example, by experimental means. Each multi-head attention module may include multiple attention heads, and the specific number of attention heads can be determined by experimental means.

[0032] In an embodiment of the present invention, the output prediction layer may be a fully connected layer.

[0033] S200. Obtain a historical time series data set.

[0034] In an embodiment of the present invention, the historical time series data set is a data set obtained by sorting multiple network traffic data of a target network link collected within a set time period in time sequence.

[0035] In an embodiment of the present invention, the target network link may be a network link specified by a user. In a schematic embodiment, the target network link is a civil aviation data network link, specifically, it may be a communication link between an airport and a data center.

[0036] In an embodiment of the present invention, the data entities of the network traffic data may include network transmission parameters and network delay. The network delay refers to the total time required for data to be transmitted from one end of the network to the other end. In an embodiment of the present invention, the network transmission parameters may include the number of source IPs, the number of destination IPs, the packet size, and the source address window size. That is, the data entities of the network traffic data in the embodiment of the present invention include the number of source IPs, the number of destination IPs, the packet size, the source address window size, and the network delay. Among them, the source IP is the IP that initiates a network session, and the destination IP is the IP that receives the network session. The source address window size is the amount of data that the destination IP can receive.

[0037] In an embodiment of the present invention, the historical time series data set may be a one-way communication data set, which can be obtained based on the traffic file of the target network link. Since the number of traffic files is huge, it is necessary to group the files by date according to the start and end values of the internal timestamp field of the traffic, including using the Python language and the Scapy function library to read the content of the pcap file, and taking the date as the classification criterion to extract the features of the traffic data occurring on the same day. The obtained historical time series data set D = {D1, D2,..., D i , ……, D n}, D i is the network traffic data of the target network link obtained at the collection moment i, D i ={d i1 , d i2 , ……, d ij , ……, d im}, d ij is the value of the jth data entity in D i . The value range of i is from 1 to n, where n is the length of the historical time series data set, and the value range of j is from 1 to m, where m is the number of data entities of the network traffic data.

[0038] In an embodiment of the present invention, network traffic data can be collected at each collection moment. The interval between two adjacent collection moments can be set according to actual needs. In a schematic embodiment, the time interval between two adjacent collection moments can be 1 minute.

[0039] In an embodiment of the present invention, when processing each data packet, if the data packet is a data packet transmitted through the TCP protocol, a session key will be generated according to the source IP, destination IP, source port, and destination port. The timestamp of the data packet will be added to the timestamp list of the session, and the end time of the session will be updated. That is, for each session, if there are h timestamps (h>1) in the timestamp list, the value of the network delay is calculated as the difference between the last timestamp and the first timestamp in the timestamp list divided by the number of timestamps minus one. That is, the network delay RTT at each moment = (t h -t1) / (h - 1), where t h is the last timestamp in the timestamp list corresponding to each moment, and t1 is the first timestamp in the timestamp list corresponding to each moment.

[0040] In an embodiment of the present invention, the historical time series data set is a data set obtained after preprocessing. The preprocessing includes clearing abnormal data with a network delay of 0 and filling in missing data. For a specific communication direction, there is no communication in every time slice within 24 hours. Therefore, there will be vacancies and interruptions in the obtained one-way communication baseline, and filling processing is required. In an embodiment of the present invention, linear interpolation is used for data filling.

[0041] In an embodiment of the present invention, the historical time series data set can be a data set of a set historical time period before the time period to be predicted. For example, to predict the network delay on date a, the historical time series data set can be a data set formed by multiple network traffic data within a set time period before date a. The duration of the set historical time period can be set according to actual needs, for example, it can be 6 days.

[0042] S300, use a sliding time window with a sliding stride of 1 and a window size of w to divide the historical time series data set into multiple sample time series data, and each sample time series data is a matrix of w×m.

[0043] In an embodiment of the present invention, a sliding stride of 1 means that each time 1 time interval of data is slid. For example, the data within 1 minute is slid. w can be set according to actual needs. In a schematic embodiment, w = 60, that is, the time series data within 60 minutes is used as a sample time series data. Each row of data in each sample time series data is the network traffic data collected at the corresponding moment.

[0044] S400 uses the numerical value of the network transmission parameter in the sample time-series data as the input information, and the numerical value of the network delay in the sample time-series data as the label information, and trains the initial network delay prediction model using the multiple sample time-series data to obtain a trained network delay prediction model.

[0045] Among them, during the training process, the total loss of the network delay prediction model includes a data fitting loss and a network characteristic constraint loss. The network characteristic constraint loss includes a load constraint loss determined based on network load, an attenuation constraint loss determined based on network burst traffic, and a window constraint loss determined based on the receive window of the network protocol.

[0046] Furthermore, S400 specifically includes:

[0047] S410 inputs the current batch of sample time-series data into the convolutional module of the current network delay prediction model for convolutional processing to obtain corresponding convolutional processing features, and inputs them to the first dropout layer; the initial value of the current network delay prediction model is the initial network delay prediction model.

[0048] S420 uses the first dropout layer to perform random dropout processing on the received convolutional processing features to obtain first dropout processing features, and inputs them to the first LSTM module.

[0049] S430 uses the first LSTM module to extract features from the received first dropout processing features to obtain first extraction features, divides the first extraction features into p1 first sub-features, and inputs the p1 first sub-features to the p1 attention heads of the first multi-head attention module respectively to perform feature enhancement processing on the first sub-features to obtain enhanced features of each first sub-feature.

[0050] S440 splices the enhanced features of the p1 first sub-features and performs linear transformation on the splicing result to obtain the first enhanced feature corresponding to the first extraction feature and inputs it to the second dropout layer.

[0051] S450 uses the second dropout layer to perform random dropout processing on the received first enhanced feature to obtain second dropout processing features, and inputs them to the second LSTM module.

[0052] S460 uses the second LSTM module to extract features from the received second dropout processing features to obtain second extraction features, divides the second extraction features into p2 second sub-features, and inputs the p2 second sub-features to the p2 attention heads of the second multi-head attention module respectively to perform feature enhancement processing on the second sub-features to obtain enhanced features of each second sub-feature.

[0053] In the embodiments of the present invention, the dimension of each first sub - feature is equal to d1 / p1, where d1 is the dimension of the first extracted feature. In a schematic embodiment, d1 = 32 and p1 = 8. The dimension of each second sub - feature is equal to d2 / p2, where d2 is the dimension of the second extracted feature. In a schematic embodiment, d2 = 16 and p2 = 2.

[0054] Those skilled in the art know that any method of using an LSTM module to extract features from input features to obtain corresponding feature extraction results belongs to the protection scope of the present invention. To avoid redundancy, the present invention omits the detailed description thereof.

[0055] In the embodiments of the present invention, the first sub - feature is the hidden state of the first LSTM module, and the second sub - feature is the hidden state of the second LSTM module. That is, the query matrix, key matrix, and value matrix in the first multi - head attention module are respectively obtained by multiplying the hidden state of the first LSTM module with the corresponding query weight matrix, key weight matrix, and value weight matrix. The query matrix, key matrix, and value matrix in the second multi - head attention module are respectively obtained by multiplying the hidden state of the second LSTM module with the corresponding query weight matrix, key weight matrix, and value weight matrix. Those skilled in the art know that the working principle of each attention head can be the prior art, that is, the method of each attention head obtaining enhanced features belongs to the prior art. To avoid redundancy, the present invention omits the detailed description thereof.

[0056] S470, splice the enhanced features of p2 second sub - features and perform a linear transformation on the splicing result to obtain the second enhanced feature corresponding to the second extracted feature and input it to the output prediction layer to obtain the delay prediction result corresponding to the sample time - series data of the current batch.

[0057] In the embodiments of the present invention, the splicing result can be multiplied by the multi - head output projection matrix for linear transformation. In the embodiments of the present invention, by mapping the features output by the LSTM module to multiple sub - spaces for parallel calculation of attention weights, each attention head can focus on different aspects of the sequence, can capture the features of data at different stages from data in different dimensions respectively, and can effectively capture the multi - dimensional and multi - level dependency relationships in the sequence.

[0058] S480, obtain the total loss of the current network delay prediction model based on the delay prediction result corresponding to the sample time - series data of the current batch and the true delay result as the current total loss. If the current model training round reaches the preset training rounds or the current total loss meets the preset conditions, use the current network delay prediction model as the trained network delay prediction model. Otherwise, update the parameters of the current network delay prediction model based on the current total loss, and use the next batch of sample time - series data as the sample time - series data of the current batch, and execute S410.

[0059] In an embodiment of the present invention, the preset condition may be that the total loss does not change for a continuous number of rounds, for example, the total loss does not change for 10 consecutive rounds. The preset number of training rounds may be an empirical value.

[0060] Furthermore, in an embodiment of the present invention, the total loss of the current network delay prediction model satisfies the following conditions:

[0061] Loss = L d + L c , where Loss is the total loss of the current network delay prediction model, and L d is the data fitting loss of the current network delay prediction model, and L d = (1 / T) ∑ T t=1 ∑ w s=1 (y ts p - y ts true ) 2 , where y ts p is the network delay prediction value corresponding to the s-th row data in the t-th sample time series data in the current batch of sample time series data. The value of t ranges from 1 to T, where T is the number of the current batch of sample time series data, and the value of s ranges from 1 to w. y ts true is the true network delay value corresponding to y ts p . L c is the network characteristic constraint loss of the current network delay prediction model, and L c = λ1 × L1 + λ2 × L2 + λ3 × L3, where L1 is the load constraint loss, L2 is the attenuation constraint loss, L3 is the window constraint loss, and λ1, λ2, and λ3 are weight hyperparameters respectively.

[0062] In an embodiment of the present invention, λ1, λ2, and λ3 can be obtained through experiments. In a schematic embodiment, λ1 = 0.3, λ2 = 0.1, and λ3 = 0.9, which can ensure the best model accuracy.

[0063] In network communication, the network delay is usually positively correlated with the network load, that is, the traffic intensity. Based on this, L1 in an embodiment of the present invention satisfies the following conditions:

[0064] L1 = (1 / T) ∑ T t=1 ∑ w s=1 (y ts p - (k load×N ts src ×S ts p ) / C), k load is the load balancing ratio coefficient, N ts src is the number of source IPs corresponding to the s-th row data in the t-th sample time series data of the current batch of sample time series data, S ts p is the data packet size corresponding to the s-th row data in the t-th sample time series data of the current batch of sample time series data, and C is the capacity of the target network link.

[0065] The inventors of the present invention found that the range of the number of source IPs per unit time granularity is between 1 and 60, and the data packet size fluctuates around 50 - 60 bytes, that is, the maximum instantaneous traffic intensity is 28.8 kbps. In the traditional VDL2 link mode (C = 31.5 kbps), the traffic intensity accounts for about 91%, which may cause an overloaded link. Assuming that the traffic intensity remains below 50% continuously to avoid network congestion, C = 57.6 kbps (28.8 kbps / 0.5) is selected. To balance the quantitative relationship between the loss calculation and the delay prediction value, k load can be set to 0.5.

[0066] In network communication, when encountering burst traffic, the RTT will increase instantaneously, and then gradually recover due to network self-adaptive functions (such as congestion control, etc.). Therefore, based on this, L2 satisfies the following conditions:

[0067] L2 = (1 / T) ∑ T t=1 ∑ w s=1 (R ts + k decay × (S ts p - S t(s-1) p )) 2 , where R ts is the delay change rate corresponding to the s-th row data in the t-th sample time series data of the current batch of sample time series data, k decay is the attenuation balance ratio coefficient, which can be an empirical value. In a schematic embodiment, k decay is set to 0.1. Where R ts = (y ts p - y t(s-1) p ) / △t, where y ts pis the predicted network delay value corresponding to the s-th row data in the t-th sample time series data of the current batch of sample time series data, y t(s-1) p is the predicted network delay value corresponding to the (s - 1)-th row data in the t-th sample time series data of the current batch of sample time series data, and △t is the time interval between two adjacent data in the sample time series data, that is, the time interval between two adjacent acquisition times.

[0068] In the embodiment of the present invention, L3 can be obtained according to the TCP throughput formula. Specifically, L3 = (1 / T)∑ T t=1 ∑ w s=1 (y ts p ×S ts p / k window ), k window is the window balance constant and can be set to 100.

[0069] The trained network delay prediction model obtained by the training method of the network delay prediction model provided by the embodiment of the present invention can detect the network state of the target network link during actual application.

[0070] Furthermore, in the embodiment of the present invention, a network state detection method is also provided. The method may include the following steps:

[0071] S10, Obtain the network traffic data of the target network link collected at the current monitoring moment as the current data to be processed.

[0072] S20, Input the numerical values of the network transmission parameters in the current data to be processed into the trained network delay prediction model to obtain the predicted network delay value corresponding to the current data to be processed.

[0073] In the embodiment of the present invention, the trained network delay prediction model is the trained network delay prediction model obtained in the foregoing embodiment.

[0074] S30, Compare the predicted network delay value corresponding to the current data to be processed with the corresponding actual network delay value, and determine the network state of the target network link at the current monitoring moment based on the comparison result.

[0075] Furthermore, S30 specifically includes:

[0076] If the predicted network delay value corresponding to the currently to-be-processed data is less than the corresponding actual network delay value, or if the difference between the predicted network delay value and the corresponding actual network delay value of the currently to-be-processed data is less than or equal to a set difference threshold, it is determined that the target network link is in a normal state. If the difference between the predicted network delay value and the corresponding actual network delay value of the currently to-be-processed data is greater than the set difference threshold, it is determined that the target network link is in an abnormal state.

[0077] In the embodiment of the present invention, the set difference threshold can be set according to actual needs. For example, it can be a value infinitely close to 0. If the predicted network delay value corresponding to the currently to-be-processed data is less than the corresponding actual network delay value, or if the difference between the predicted network delay value and the corresponding actual network delay value of the currently to-be-processed data is less than or equal to the set difference threshold, that is, if the predicted network delay value corresponding to the currently to-be-processed data is greater than the corresponding actual network delay value, but the difference between them is very small, it indicates that the transmission of the target network link is normal. If the difference between the predicted network delay value and the corresponding actual network delay value of the currently to-be-processed data is greater than the set difference threshold, that is, if the predicted network delay value corresponding to the currently to-be-processed data is greater than the corresponding actual network delay value and the difference between them is greater than the set difference threshold, it indicates that the transmission of the target network link is abnormal.

[0078] Furthermore, it further includes:

[0079] S40, visually display the network state of the target network link at the current monitoring moment.

[0080] Those skilled in the art know that any method for visually displaying the network state of the target network link at the current monitoring moment belongs to the protection scope of the present invention.

[0081] Another embodiment of the present invention further provides a training device for a network delay prediction model, including:

[0082] A model construction module, configured to construct an initial network delay prediction model. The network delay prediction model includes a convolutional module, a first dropout layer, a first feature extraction network, a second dropout layer, a second feature extraction network, and an output prediction layer connected in sequence. Among them, the first feature extraction network includes a first LSTM module and a first multi-head attention module, and the second feature extraction network includes a second LSTM module and a second multi-head attention module.

[0083] A first data acquisition module, configured to acquire a historical time series data set. The historical time series data set is a data set obtained by sorting multiple network traffic data of the target network link collected within a set time period according to time series. The data entities of the network traffic data include network transmission parameters and network delays.

[0084] A second data acquisition module, configured to divide the historical time series data set into multiple sample time series data by using a sliding time window with a sliding step of 1 and a window size of w. Each time series data is a matrix of w×m, where m is the number of data entities of the network traffic data.

[0085] A model training module, configured to use the numerical values of the network transmission parameters in the sample time series data as input information, use the numerical values of the network delay in the sample time series data as label information, and train an initial network delay prediction model by using the multiple sample time series data to obtain a trained network delay prediction model. During the training process, the total loss of the network delay prediction model includes a data fitting loss and a network characteristic constraint loss. The network characteristic constraint loss includes a load constraint loss determined based on network load, an attenuation constraint loss determined based on network burst traffic, and a window constraint loss determined based on the receive window of the network protocol.

[0086] This device can be used to execute Figure 1 the method shown in the embodiments shown, and therefore, for the functions that can be realized by each functional module of this device, reference can be made to Figure 1 the description of the embodiments shown, which will not be elaborated here.

[0087] An embodiment of the present invention further provides an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are configured to execute the method of the embodiment of the present invention.

[0088] An embodiment of the present invention further provides a computer-readable storage medium, storing computer-executable instructions, and the computer instructions are used to execute the method of the embodiment of the present invention.

[0089] It should be understood that various forms of the processes shown above can be used, steps can be reordered, added or deleted. For example, the steps described in the present invention can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in the present invention can be achieved. No limitation is imposed herein.

[0090] The above specific embodiments do not constitute a limitation to the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.

Claims

1. A training method for a network delay prediction model, characterized in that: The method comprises the following steps: S100, constructing an initial network delay prediction model, wherein the network delay prediction model includes a convolution module, a first drop layer, a first feature extraction network, a second drop layer, a second feature extraction network, and an output prediction layer connected in sequence, wherein the first feature extraction network includes a first LSTM module and a first multi-head attention module, and the second feature extraction network includes a second LSTM module and a second multi-head attention module; S200, obtaining a historical time series data set, wherein the historical time series data set is a data set obtained by sorting a plurality of network flow data of a target network link collected within a set time period in time series, wherein the data entity of the network flow data includes a network transmission parameter and a network delay; S300, using a sliding time window with a sliding step of 1 and a window size of w to divide the historical time series data set into a plurality of sample time series data, each sample time series data being a w×m matrix, where m is the number of data entities of the network traffic data; S400, using the numerical values ​​of network transmission parameters in the sample time series data as input information, using the numerical values ​​of network delay in the sample time series data as label information, and using the multiple sample time series data to train an initial network delay prediction model to obtain a trained network delay prediction model, wherein during the training process, the total loss of the network delay prediction model includes data fitting loss and network characteristic constraint loss, and the network characteristic constraint loss includes a load constraint loss determined based on the network load, an attenuation constraint loss determined based on the network burst traffic, and a window constraint loss determined for a receiving window based on the network protocol.

2. The method according to claim 1, characterized in that S400 specifically includes: S410, inputting the sample time series data of the current batch into the convolution module of the current network delay prediction model for convolution processing, obtaining corresponding convolution processing features, and inputting them into the first discard layer; the initial value of the current network delay prediction model is the initial network delay prediction model; S420, using the first drop layer to perform random drop processing on the received convolution processing features to obtain first drop processing features, and input them to the first LSTM module; S430, using the first LSTM module to perform feature extraction on the received first discarded processing feature to obtain a first extracted feature and divide the first extracted feature into p1 first sub-features, and respectively input the p1 first sub-features to p1 attention heads of the first multi-head attention module to perform feature enhancement processing on the first sub-features to obtain enhanced features of each first sub-feature; S440, concatenating the enhanced features of the p1 first sub-features and performing a linear transformation on the concatenated result to obtain a first enhanced feature corresponding to the first extracted feature and input the first enhanced feature to the second discard layer; S450, using the second discard layer to perform random discard processing on the received first enhanced feature to obtain a second discard processing feature, and input it to the second LSTM module; S460, using a second LSTM module to perform feature extraction on the received second discarded processing feature to obtain a second extracted feature and divide the second extracted feature into p2 second sub-features, and respectively input the p2 second sub-features to p2 attention heads of a second multi-head attention module to perform feature enhancement processing on the second sub-features to obtain enhanced features of each second sub-feature; S470, concatenating the enhanced features of the p2 second sub-features and performing linear transformation on the concatenated results to obtain the second enhanced features corresponding to the second extracted features and inputting them to the output prediction layer to obtain the delay prediction results corresponding to the sample time series data of the current batch; S480, based on the delay prediction results corresponding to the current batch of sample timing data and the actual delay results, obtain the total loss of the current network delay prediction model as the current total loss. If the current model training round number reaches the preset training round number or the current total loss meets the preset conditions, the current network delay prediction model is used as the trained network delay prediction model. Otherwise, the parameters of the current network delay prediction model are updated based on the current total loss, and the next batch of sample timing data is used as the sample timing data of the current batch, and S410 is executed.

3. The method according to claim 2, characterized in that The total loss of the current network delay prediction model meets the following conditions: Loss=L d +L c , Loss is the total loss of the current network delay prediction model, L d is the data fitting loss of the current network delay prediction model, L d =(1 / T)∑ T t=1 ∑ w s=1 (y ts p -y ts true ) 2 ,y ts p is the network delay prediction value corresponding to the sth row of data in the tth sample time series data in the current batch of sample time series data. The value of t ranges from 1 to T, T is the number of sample time series data in the current batch, and the value of s ranges from 1 to w. ts true for y ts p The corresponding real value of network delay, L c is the network characteristic constraint loss of the current network delay prediction model, L c =λ1×L1+λ2×L2+λ3×L3, where L1 is the load constraint loss, L2 is the decay constraint loss, L3 is the window constraint loss, and λ1, λ2 and λ3 are weight hyperparameters respectively.

4. The method according to claim 3, characterized in that L1=(1 / T)∑ T t=1 ∑ w s=1 (y ts p -(k load ×N ts src ×S ts p ) / C), k load is the load balancing coefficient, N ts src is the number of source IP addresses corresponding to the s-th row of data in the t-th sample time series data in the current batch of sample time series data, S ts p is the data packet size corresponding to the s-th row of data in the t-th sample time series data in the current batch of sample time series data, and C is the capacity of the target network link; L2=(1 / T)∑ T t=1 ∑ w s=1 (R ts +k decay ×(S ts p -S t(s-1) p ) 2 , where R ts is the delay change rate corresponding to the s-th row of data in the t-th sample time series data in the current batch of sample time series data, k decay is the attenuation balance coefficient, L3=(1 / T)∑ T t=1 ∑ w s=1 (y ts p ×S ts p / k window ), k window is the window equilibrium constant.

5. The method according to claim 1, characterized in that: The network transmission parameters include the number of source IPs, the number of destination IPs, the data packet size, and the source address window size.

6. The method according to claim 1, characterized in that The target network link is a civil aviation data network link.

7. A training device for a network delay prediction model, characterized in that: include: A model building module, used to build an initial network delay prediction model, wherein the network delay prediction model includes a convolution module, a first drop layer, a first feature extraction network, a second drop layer, a second feature extraction network and an output prediction layer connected in sequence, wherein the first feature extraction network includes a first LSTM module and a first multi-head attention module, and the second feature extraction network includes a second LSTM module and a second multi-head attention module; A first data acquisition module is used to acquire a historical time series data set, wherein the historical time series data set is a data set obtained by sorting a plurality of network flow data of a target network link collected within a set time period in time series, and the data entity of the network flow data includes a network transmission parameter and a network delay; A second data acquisition module is used to divide the historical time series data set into a plurality of sample time series data using a sliding time window with a sliding step of 1 and a window size of w, each time series data is a w×m matrix, where m is the number of data entities of the network traffic data; A model training module is used to use the numerical values ​​of network transmission parameters in sample time series data as input information, and the numerical values ​​of network delay in the sample time series data as label information, and use the multiple sample time series data to train an initial network delay prediction model to obtain a trained network delay prediction model, wherein during the training process, the total loss of the network delay prediction model includes data fitting loss and network characteristic constraint loss, and the network characteristic constraint loss includes a load constraint loss determined based on the network load, an attenuation constraint loss determined based on the network burst traffic, and a window constraint loss determined for a receiving window based on the network protocol.

8. A network status detection method, characterized in that: The method comprises the following steps: S10, obtaining the network traffic data of the target network link collected at the current monitoring time as the current data to be processed; S20, inputting the numerical value of the network transmission parameter in the current data to be processed into the trained network delay prediction model obtained by the method described in any one of claims 1 to 6, to obtain the network delay prediction value corresponding to the current data to be processed; S30, comparing the predicted network delay value corresponding to the current data to be processed with the corresponding actual network delay value, and determining the network state of the target network link at the current monitoring time based on the comparison result.

9. The network status detection method according to claim 8, characterized in that: S30 specifically includes: If the network delay prediction value corresponding to the current data to be processed is less than the corresponding network delay actual value, or if the difference between the network delay prediction value corresponding to the current data to be processed and the corresponding network delay actual value is less than or equal to the set difference threshold, it is determined that the target network link is in a normal state; if the difference between the network delay prediction value corresponding to the current data to be processed and the corresponding network delay actual value is greater than the set difference threshold, it is determined that the target network link is in an abnormal state.

Citation Information

Patent Citations

  • Training method and device of delay prediction model and congestion control method and device

    CN113992599B

  • Optimization method of efficient throughput capacity of multipath parallel transmission system

    CN105873096A

  • Time delay prediction model training method and device, and congestion control method and device

    CN113992599A