Network time delay prediction model training method and device and network state detection method
By using the LSTM module of the multi-head attention module in the network delay prediction model and adding network characteristic constraint loss, the problem of insufficient delay prediction accuracy in the prior art is solved, and a higher delay prediction accuracy is achieved.
Patent Information
- Application Number
- CN202510519635.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-24
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2045-04-24
AI Technical Summary
The accuracy of delay prediction in the prior art under low signal-to-noise ratio or homofrequency interference is insufficient, making it difficult to effectively solve the impact of complex network activities on prediction accuracy.
The LSTM module with a multi-head attention module is used as the network delay prediction model, and the constraint loss based on network load, burst traffic and protocol reception window is added during the model training process to improve prediction accuracy.
By alleviating the impact of complex network activities on prediction accuracy, the accuracy of delay prediction is significantly improved, making the prediction delay accuracy more accurate.
Smart Images

Figure CN120034449A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology applications, and in particular to a training method and device for a network delay prediction model and a network status detection method. Background Art
[0002] Delay prediction refers to the technology of estimating the time delay in the data transmission process through algorithms or models. In communication networks, delay is affected by factors such as noise, interference, and bandwidth fluctuations. Accurately predicting delay can optimize data scheduling and improve service quality. For example, in civil aviation data communication networks, it is necessary to ensure smooth communication between data airports and data centers and avoid unnecessary losses. The core of delay prediction is to analyze historical data and real-time signal characteristics, and to achieve dynamic modeling and prediction of delay through mathematical modeling or intelligent learning methods to solve the problem of estimation error under low signal-to-noise ratio or co-channel interference. Patent document (CN113992599B) discloses a training method and device for delay prediction model and a congestion control method and device. The document uses an encoding neural network and a decoding neural network for delay prediction, and uses the difference between the delay prediction value and the actual delay value as the loss function. The solution disclosed in the patent document can improve the accuracy of delay prediction to a certain extent. However, another prediction scheme that can accurately predict delay can be provided. Summary of the invention
[0003] In view of the above technical problems, the technical solution adopted by the present invention is: According to a first aspect of the present invention, a method for training a network delay prediction model is provided, the method comprising the following steps: S100, constructing an initial network delay prediction model, wherein the network delay prediction model includes a convolution module, a first drop layer, a first feature extraction network, a second drop layer, a second feature extraction network and an output prediction layer connected in sequence, wherein the first feature extraction network includes a first LSTM module and a first multi-head attention module, and the second feature extraction network includes a second LSTM module and a second multi-head attention module.
[0004] S200, obtaining a historical time series data set, wherein the historical time series data set is a data set obtained by sorting a plurality of network flow data of a target network link collected within a set time period in time series, wherein the data entity of the network flow data includes a network transmission parameter and a network delay; S300, using a sliding time window with a sliding step of 1 and a window size of w to divide the historical time series data set into multiple sample time series data, each sample time series data is a w×m matrix, and m is the number of data entities of the network traffic data.
[0005] S400, using the numerical values of network transmission parameters in the sample time series data as input information, using the numerical values of network delay in the sample time series data as label information, and using the multiple sample time series data to train an initial network delay prediction model to obtain a trained network delay prediction model, wherein during the training process, the total loss of the network delay prediction model includes data fitting loss and network characteristic constraint loss, and the network characteristic constraint loss includes a load constraint loss determined based on the network load, an attenuation constraint loss determined based on the network burst traffic, and a window constraint loss determined for a receiving window based on the network protocol.
[0006] According to a second aspect of the present invention, there is provided a training device for a network delay prediction model, comprising: A model building module is used to construct an initial network delay prediction model, which includes a convolution module, a first drop layer, a first feature extraction network, a second drop layer, a second feature extraction network and an output prediction layer connected in sequence, wherein the first feature extraction network includes a first LSTM module and a first multi-head attention module, and the second feature extraction network includes a second LSTM module and a second multi-head attention module.
[0007] The first data acquisition module is used to acquire a historical time series data set, where the historical time series data set is a data set obtained by sorting multiple network flow data of a target network link collected within a set time period in time series, and the data entity of the network flow data includes network transmission parameters and network delay.
[0008] The second data acquisition module is used to divide the historical time series data set into multiple sample time series data using a sliding time window with a sliding step of 1 and a window size of w, each time series data is a w×m matrix, and m is the number of data entities of the network traffic data.
[0009] A model training module is used to use the numerical values of network transmission parameters in sample time series data as input information, and the numerical values of network delay in the sample time series data as label information, and use the multiple sample time series data to train an initial network delay prediction model to obtain a trained network delay prediction model, wherein during the training process, the total loss of the network delay prediction model includes data fitting loss and network characteristic constraint loss, and the network characteristic constraint loss includes a load constraint loss determined based on the network load, an attenuation constraint loss determined based on the network burst traffic, and a window constraint loss determined for a receiving window based on the network protocol.
[0010] According to a third aspect of the present invention, a network status detection method is provided, the method comprising the following steps: S10, obtaining the network traffic data of the target network link collected at the current monitoring time as the current data to be processed.
[0011] S20, inputting the numerical value of the network transmission parameter in the current data to be processed into the trained network delay prediction model obtained in the first aspect of the present invention, to obtain the network delay prediction value corresponding to the current data to be processed.
[0012] S30, comparing the predicted network delay value corresponding to the current data to be processed with the corresponding actual network delay value, and determining the network state of the target network link at the current monitoring time based on the comparison result.
[0013] The present invention has at least the following beneficial effects: The training method of the network delay prediction model provided by the embodiment of the present invention uses an LSTM module integrated with a multi-head attention module as the network delay prediction model, and in the model training process, adds a load constraint loss determined based on the network load, an attenuation constraint loss determined based on the network burst traffic, and a window constraint loss determined for a receiving window based on the network protocol as the loss of the model. This can alleviate the impact of complex network activities on the prediction accuracy and make the predicted delay accuracy more accurate.
[0014] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present invention, nor are they intended to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0016] Figure 1 A flowchart of a method for training a network delay prediction model provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0017] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present invention.
[0018] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art of the present invention. The terms used herein in the specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. The term "and / or" used herein includes any and all combinations of one or more related listed items.
[0019] It should be noted that some exemplary embodiments are described as processes or methods depicted as flow charts. Although the flow charts describe the steps as sequential processes, many of the steps can be implemented in parallel, concurrently or simultaneously. In addition, the order of the steps can be rearranged. The process can be terminated when its operation is completed, but it can also have additional steps not included in the accompanying drawings. The process can correspond to a method, function, procedure, subroutine, subprogram, etc.
[0020] The embodiment of the present invention provides a method for training a network delay prediction model, such as Figure 1 As shown, the method may include the following steps: S100, constructing an initial network delay prediction model.
[0021] In an embodiment of the present invention, the network delay prediction model may include a convolution module, a first discard layer, a first feature extraction network, a second discard layer, a second feature extraction network and an output prediction layer connected in sequence, wherein the first feature extraction network includes a first LSTM module and a first multi-head attention module, and the second feature extraction network includes a second LSTM module and a second multi-head attention module.
[0022] The convolution module is used to perform a data size standardization operation on the input sequence data so that the size of each data is the same. The convolution module may include multiple convolution layers, and the specific number may be determined based on actual conditions. In an illustrative embodiment, 64 convolutions and a convolution layer of size 32 may be included.
[0023] The first dropout layer is used to apply random inactivation to the features obtained by the convolution module, suppressing the redundant features extracted by the convolution layer while maintaining the consistency of the temporal feature dimension. The dropout probability of the first dropout layer can be set to 0.3. The second dropout layer is used to apply random inactivation to the output of the attention module to prevent the model from overfitting. The dropout probability of the second dropout layer can be set to 0.3.
[0024] Each LSTM module may include multiple LSTM units connected in series, and the specific number of units may be determined based on experiments, for example, by experimental means. Each multi-head attention module may include multiple attention heads, and the specific number of attention heads may be determined by experimental means.
[0025] In an embodiment of the present invention, the output prediction layer may be a fully connected layer.
[0026] S200, obtaining a historical time series data set.
[0027] In the embodiment of the present invention, the historical time series data set is a data set obtained by sorting a plurality of network flow data of a target network link collected within a set time period in time series.
[0028] In an embodiment of the present invention, the target network link may be a network link specified by a user. In an exemplary embodiment, the target network link is a civil aviation data network link, specifically a communication link between an airport and a data center.
[0029] In an embodiment of the present invention, the data entity of the network traffic data may include network transmission parameters and network delay. Network delay refers to the total time required for data to be transmitted from one end of the network to the other end. In an embodiment of the present invention, the network transmission parameters may include the number of source IPs, the number of destination IPs, the data packet size, and the source address window size, that is, the data entity of the network traffic data in an embodiment of the present invention includes the number of source IPs, the number of destination IPs, the data packet size, the source address window size, and the network delay. Among them, the source IP is the IP that initiates the network session, and the destination IP is the IP that receives the network session. The source address window size is the amount of data that the destination IP can receive.
[0030] In the embodiment of the present invention, the historical time series data set can be a one-way communication data set, which can be obtained based on the traffic file of the target network link. Due to the large number of traffic files, it is necessary to group the files with respect to the date according to the start and end values of the timestamp field inside the traffic, including using the Python language and the scapy function library to read the pcap file content, taking the date as the classification standard, and extracting the features of the traffic data occurring on the same day. The obtained historical time series data set D={D 1 , D 2 , ..., D i , ..., D n}, D i is the network traffic data of the target network link obtained at the collection time i, D i ={d i1 , d i2 , ..., d ij , ..., d im}, d ij D i The value of the j-th data entity in , i ranges from 1 to n, n is the length of the historical time series data set, j ranges from 1 to m, m is the number of data entities of the network traffic data.
[0031] In the embodiment of the present invention, network flow data may be collected at each collection moment. The interval between two adjacent collection moments may be set based on actual needs. In an exemplary embodiment, the time interval between two adjacent collection moments may be 1 minute.
[0032] In the embodiment of the present invention, when processing each data packet, if the data packet is a data packet transmitted via the TCP protocol, a session key is generated according to the source IP, destination IP, source port and destination port, the timestamp of the data packet is added to the timestamp list of the session, and the end time of the session is updated. That is, for each session, if there are h timestamps in the timestamp list (h>1), the value of the network delay is calculated as the difference between the last timestamp and the first timestamp in the timestamp list divided by the number of timestamps minus one, that is, the network delay RTT at each moment = (t h -t 1 ) / (h-1), t h is the last timestamp in the timestamp list corresponding to each moment, t 1 The first timestamp in the timestamp list corresponding to each moment.
[0033] In the embodiment of the present invention, the historical time series data set is a data set obtained after preprocessing. The preprocessing includes clearing abnormal data with a network delay of 0 and filling in missing data. For a specific communication direction, there is not communication in every time slice in 24 hours, so the obtained unidirectional communication baseline will have vacancies and interruptions, which need to be filled. In the embodiment of the present invention, linear interpolation is used for data filling.
[0034] In an embodiment of the present invention, the historical time series data set may be a data set of a set historical time period before the time period to be predicted. For example, if the network delay of date a needs to be predicted, the historical time series data set may be a data set formed by a plurality of network traffic data within a set time period before date a. The length of the set historical time period may be set based on actual needs, for example, it may be 6 days.
[0035] S300, using a sliding time window with a sliding step of 1 and a window size of w to divide the historical time series data set into a plurality of sample time series data, each sample time series data being a w×m matrix.
[0036] In the embodiment of the present invention, a sliding step of 1 means sliding data of 1 time interval each time, for example, sliding data within 1 minute. w can be set based on actual needs. In an exemplary embodiment, w=60, that is, the time series data within 60 minutes is taken as a sample time series data. Each row of data in each sample time series data is the network traffic data collected at the corresponding time.
[0037] S400, using the numerical values of network transmission parameters in the sample time series data as input information, using the numerical values of network delays in the sample time series data as label information, and using the multiple sample time series data to train an initial network delay prediction model to obtain a trained network delay prediction model.
[0038] Among them, during the training process, the total loss of the network delay prediction model includes data fitting loss and network characteristic constraint loss, and the network characteristic constraint loss includes load constraint loss determined based on network load, attenuation constraint loss determined based on network burst traffic, and window constraint loss determined for a receiving window based on a network protocol.
[0039] Furthermore, S400 specifically includes: S410, input the sample time series data of the current batch into the convolution module of the current network delay prediction model for convolution processing, obtain the corresponding convolution processing features, and input them to the first discard layer; the initial value of the current network delay prediction model is the initial network delay prediction model.
[0040] S420, using the first discard layer to perform random discard processing on the received convolution processing features, obtain a first discard processing feature, and input it to the first LSTM module.
[0041] S430, using the first LSTM module to perform feature extraction on the received first discarded processing feature, obtain a first extracted feature and divide the first extracted feature into p1 first sub-features, and input the p1 first sub-features into p1 attention heads of the first multi-head attention module respectively to perform feature enhancement processing on the first sub-features to obtain enhanced features of each first sub-feature.
[0042] S440 , concatenating the enhanced features of the p1 first sub-features and performing a linear transformation on the concatenated result to obtain a first enhanced feature corresponding to the first extracted feature and input the first enhanced feature to the second discard layer.
[0043] S450, using the second discard layer to perform random discard processing on the received first enhanced feature, obtain a second discard processing feature, and input it to the second LSTM module.
[0044] S460, using the second LSTM module to perform feature extraction on the received second discarded processing feature, obtain the second extracted feature and divide the second extracted feature into p2 second sub-features, and input the p2 second sub-features into the p2 attention heads of the second multi-head attention module respectively to perform feature enhancement processing on the second sub-features to obtain enhanced features of each second sub-feature.
[0045] In an embodiment of the present invention, the dimension of each first sub-feature is equal to d1 / p1, where d1 is the dimension of the first extracted feature. In an exemplary embodiment, d1=32 and p1=8. The dimension of each second sub-feature is equal to d2 / p2, where d2 is the dimension of the second extracted feature. In an exemplary embodiment, d2=16 and p2=2.
[0046] Those skilled in the art know that any method of using an LSTM module to extract input features to obtain corresponding feature extraction results falls within the protection scope of the present invention. To avoid redundancy, the present invention omits a detailed description of this.
[0047] In an embodiment of the present invention, the first sub-feature is the hidden state of the first LSTM module, and the second sub-feature is the hidden state of the second LSTM module, that is, the query matrix, key matrix and value matrix in the first multi-head attention module are obtained by multiplying the hidden state of the first LSTM module and the corresponding query weight matrix, key weight matrix and value weight matrix, respectively, and the query matrix, key matrix and value matrix in the second multi-head attention module are obtained by multiplying the hidden state of the second LSTM module and the corresponding query weight matrix, key weight matrix and value weight matrix, respectively. It is known to those skilled in the art that the working principle of each attention head can be the prior art, that is, the method for obtaining enhanced features by each attention head belongs to the existing method. To avoid redundancy, the present invention omits a detailed description of this.
[0048] S470, concatenate the enhanced features of the p2 second sub-features and perform linear transformation on the concatenated result to obtain the second enhanced feature corresponding to the second extracted feature and input it to the output prediction layer to obtain the delay prediction result corresponding to the current batch of sample time series data.
[0049] In an embodiment of the present invention, the splicing result can be multiplied by the multi-head output projection matrix for linear transformation. In an embodiment of the present invention, by mapping the features output by the LSTM module to multiple subspaces and calculating the attention weights in parallel, each attention head can focus on different aspects of the sequence, and can capture the features of the data at different stages from data of different dimensions, and can effectively capture the multi-dimensional and multi-level dependencies in the sequence.
[0050] S480, based on the delay prediction results corresponding to the current batch of sample timing data and the actual delay results, obtain the total loss of the current network delay prediction model as the current total loss. If the current model training round number reaches the preset training round number or the current total loss meets the preset conditions, the current network delay prediction model is used as the trained network delay prediction model. Otherwise, the parameters of the current network delay prediction model are updated based on the current total loss, and the next batch of sample timing data is used as the sample timing data of the current batch, and S410 is executed.
[0051] In the embodiment of the present invention, the preset condition may be that the total loss does not change for a plurality of consecutive rounds, for example, the total loss does not change for 10 consecutive rounds. The preset training round may be an experience value.
[0052] Furthermore, in an embodiment of the present invention, the total loss of the current network delay prediction model satisfies the following conditions: Loss=L d +L c , Loss is the total loss of the current network delay prediction model, L d is the data fitting loss of the current network delay prediction model, L d =(1 / T)∑ T t=1 ∑ w s=1 (y ts p -y ts true ) 2 ,y ts p is the network delay prediction value corresponding to the sth row of data in the tth sample time series data in the current batch of sample time series data. The value of t ranges from 1 to T, T is the number of sample time series data in the current batch, and the value of s ranges from 1 to w. ts true for y ts p The corresponding real value of network delay, L c is the network characteristic constraint loss of the current network delay prediction model, L c =λ1×L1+λ2×L2+λ3×L3, where L1 is the load constraint loss, L2 is the decay constraint loss, L3 is the window constraint loss, and λ1, λ2 and λ3 are weight hyperparameters respectively.
[0053] In the embodiment of the present invention, λ1, λ2 and λ3 can be obtained by experiments. In an illustrative embodiment, λ1=0.3, λ2=0.1, and λ3=0.9, which can ensure the best model accuracy.
[0054] In network communication, network delay is usually positively correlated with network load, that is, traffic intensity. Based on this, L1 in the embodiment of the present invention satisfies the following conditions: L1=(1 / T)∑ T t=1 ∑ w s=1 (y ts p -(k load ×N ts src×S ts p ) / C), k load is the load balancing coefficient, N ts src is the number of source IP addresses corresponding to the s-th row of data in the t-th sample time series data in the current batch of sample time series data, S ts p is the data packet size corresponding to the s-th row of data in the t-th sample time series data in the current batch of sample time series data, and C is the capacity of the target network link.
[0055] The inventors of the present invention found that the number of source IPs per unit time granularity ranges from 1 to 60, and the packet size fluctuates around 50-60 bytes, that is, the maximum flow intensity is 28.8 kbps. In the traditional VDL2 link mode (C = 31.5 kbps), the flow intensity accounts for about 91%, which may cause excessive link burden. Assuming that the flow intensity continues to remain below 50% to avoid network congestion, C = 57.6 kbps (28.8 kbps / 0.5) is selected. In order to balance the quantitative relationship between the loss calculation and the delay prediction value, k load Can be set to 0.5.
[0056] In network communications, burst traffic will cause the RTT to increase instantaneously, and then gradually recover due to network adaptive functions (such as congestion control, etc.). Therefore, based on this, L2 meets the following conditions: L2=(1 / T)∑ T t=1 ∑ w s=1 (R ts +k decay ×(S ts p -S t(s-1) p ) 2 , where R ts is the delay change rate corresponding to the s-th row of data in the t-th sample time series data in the current batch of sample time series data, k decay is the attenuation balance proportional coefficient, which can be an empirical value. In an exemplary embodiment, k decay Set to 0.1. Among them, R ts =(y ts p -y t(s-1) p ) / △t, where y ts p is the network delay prediction value corresponding to the s-th row of data in the t-th sample time series data in the current batch of sample time series data, y t(s-1) pis the network delay prediction value corresponding to the s-1th row of data in the tth sample time series data in the current batch of sample time series data, and △t is the time interval between two adjacent data in the sample time series data, that is, the time interval between two adjacent collection moments.
[0057] In the embodiment of the present invention, L3 can be obtained according to the TCP throughput formula. Specifically, L3=(1 / T)∑ T t=1 ∑ w s=1 (y ts p ×S ts p / k window ), k window is the window equilibrium constant and can be set to 100.
[0058] The trained network delay prediction model obtained by the network delay prediction model training method provided in the embodiment of the present invention can detect the network status of the target network link during actual application.
[0059] Furthermore, in an embodiment of the present invention, a network status detection method is also provided, and the method may include the following steps: S10, obtaining the network traffic data of the target network link collected at the current monitoring time as the current data to be processed.
[0060] S20, inputting the numerical value of the network transmission parameter in the current data to be processed into the trained network delay prediction model to obtain the network delay prediction value corresponding to the current data to be processed.
[0061] In the embodiment of the present invention, the trained network delay prediction model is the trained network delay prediction model obtained in the above-mentioned embodiment.
[0062] S30, comparing the predicted network delay value corresponding to the current data to be processed with the corresponding actual network delay value, and determining the network state of the target network link at the current monitoring time based on the comparison result.
[0063] Furthermore, S30 specifically includes: If the network delay prediction value corresponding to the current data to be processed is less than the corresponding network delay actual value, or if the difference between the network delay prediction value corresponding to the current data to be processed and the corresponding network delay actual value is less than or equal to the set difference threshold, it is determined that the target network link is in a normal state; if the difference between the network delay prediction value corresponding to the current data to be processed and the corresponding network delay actual value is greater than the set difference threshold, it is determined that the target network link is in an abnormal state.
[0064] In an embodiment of the present invention, the difference threshold value may be set based on actual needs, for example, it may be a value that is wirelessly close to 0. If the network delay prediction value corresponding to the current data to be processed is less than the corresponding network delay actual value, or if the difference between the network delay prediction value corresponding to the current data to be processed and the corresponding network delay actual value is less than or equal to the set difference threshold value, that is, if the network delay prediction value corresponding to the current data to be processed is greater than the corresponding network delay actual value, but the difference between the two is very small, it indicates that the transmission of the target network link is normal, if the difference between the network delay prediction value corresponding to the current data to be processed and the corresponding network delay actual value is greater than the set difference threshold value, that is, if the network delay prediction value corresponding to the current data to be processed is greater than the corresponding network delay actual value and the difference between the two is greater than the set difference threshold value, it indicates that the transmission of the target network link is abnormal.
[0065] Furthermore, it also includes: S40, visually displaying the network status of the target network link at the current monitoring time.
[0066] Those skilled in the art know that any method for visually displaying the network status of a target network link at a current monitoring moment falls within the protection scope of the present invention.
[0067] Another embodiment of the present invention further provides a training device for a network delay prediction model, comprising: A model building module is used to construct an initial network delay prediction model, which includes a convolution module, a first drop layer, a first feature extraction network, a second drop layer, a second feature extraction network and an output prediction layer connected in sequence, wherein the first feature extraction network includes a first LSTM module and a first multi-head attention module, and the second feature extraction network includes a second LSTM module and a second multi-head attention module.
[0068] The first data acquisition module is used to acquire a historical time series data set, where the historical time series data set is a data set obtained by sorting multiple network flow data of a target network link collected within a set time period in time series, and the data entity of the network flow data includes network transmission parameters and network delay.
[0069] The second data acquisition module is used to divide the historical time series data set into multiple sample time series data using a sliding time window with a sliding step of 1 and a window size of w, each time series data is a w×m matrix, and m is the number of data entities of the network traffic data.
[0070] A model training module is used to use the numerical values of network transmission parameters in sample time series data as input information, and the numerical values of network delay in the sample time series data as label information, and use the multiple sample time series data to train an initial network delay prediction model to obtain a trained network delay prediction model, wherein during the training process, the total loss of the network delay prediction model includes data fitting loss and network characteristic constraint loss, and the network characteristic constraint loss includes a load constraint loss determined based on the network load, an attenuation constraint loss determined based on the network burst traffic, and a window constraint loss determined for a receiving window based on the network protocol.
[0071] The device can be used to perform Figure 1 Therefore, for the functions that can be realized by each functional module of the device, please refer to Figure 1 The description of the illustrated embodiment is not repeated in detail.
[0072] An embodiment of the present invention also provides an electronic device, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are configured to execute the method described in the embodiment of the present invention.
[0073] The embodiment of the present invention further provides a computer-readable storage medium storing computer-executable instructions, wherein the computer instructions are used to execute the method described in the embodiment of the present invention.
[0074] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps described in the present invention can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solution disclosed in the present invention can be achieved, and this document does not limit this.
[0075] The above specific implementations do not constitute a limitation on the protection scope of the present invention. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modification, equivalent substitution and improvement made within the spirit and principle of the present invention should be included in the protection scope of the present invention.
Claims
1. A training method for a network delay prediction model, characterized in that: The method comprises the following steps: S100, constructing an initial network delay prediction model, wherein the network delay prediction model includes a convolution module, a first drop layer, a first feature extraction network, a second drop layer, a second feature extraction network, and an output prediction layer connected in sequence, wherein the first feature extraction network includes a first LSTM module and a first multi-head attention module, and the second feature extraction network includes a second LSTM module and a second multi-head attention module; S200, obtaining a historical time series data set, wherein the historical time series data set is a data set obtained by sorting a plurality of network flow data of a target network link collected within a set time period in time series, wherein the data entity of the network flow data includes a network transmission parameter and a network delay; S300, using a sliding time window with a sliding step of 1 and a window size of w to divide the historical time series data set into a plurality of sample time series data, each sample time series data being a w×m matrix, where m is the number of data entities of the network traffic data; S400, using the numerical values of network transmission parameters in the sample time series data as input information, using the numerical values of network delay in the sample time series data as label information, and using the multiple sample time series data to train an initial network delay prediction model to obtain a trained network delay prediction model, wherein during the training process, the total loss of the network delay prediction model includes data fitting loss and network characteristic constraint loss, and the network characteristic constraint loss includes a load constraint loss determined based on the network load, an attenuation constraint loss determined based on the network burst traffic, and a window constraint loss determined for a receiving window based on the network protocol.
2. The method according to claim 1, characterized in that S400 specifically includes: S410, inputting the sample time series data of the current batch into the convolution module of the current network delay prediction model for convolution processing, obtaining corresponding convolution processing features, and inputting them into the first discard layer; the initial value of the current network delay prediction model is the initial network delay prediction model; S420, using the first drop layer to perform random drop processing on the received convolution processing features to obtain first drop processing features, and input them to the first LSTM module; S430, using the first LSTM module to perform feature extraction on the received first discarded processing feature to obtain a first extracted feature and divide the first extracted feature into p1 first sub-features, and respectively input the p1 first sub-features to p1 attention heads of the first multi-head attention module to perform feature enhancement processing on the first sub-features to obtain enhanced features of each first sub-feature; S440, concatenating the enhanced features of the p1 first sub-features and performing a linear transformation on the concatenated result to obtain a first enhanced feature corresponding to the first extracted feature and input the first enhanced feature to the second discard layer; S450, using the second discard layer to perform random discard processing on the received first enhanced feature to obtain a second discard processing feature, and input it to the second LSTM module; S460, using a second LSTM module to perform feature extraction on the received second discarded processing feature to obtain a second extracted feature and divide the second extracted feature into p2 second sub-features, and respectively input the p2 second sub-features to p2 attention heads of a second multi-head attention module to perform feature enhancement processing on the second sub-features to obtain enhanced features of each second sub-feature; S470, concatenating the enhanced features of the p2 second sub-features and performing linear transformation on the concatenated results to obtain the second enhanced features corresponding to the second extracted features and inputting them to the output prediction layer to obtain the delay prediction results corresponding to the sample time series data of the current batch; S480, based on the delay prediction results corresponding to the current batch of sample timing data and the actual delay results, obtain the total loss of the current network delay prediction model as the current total loss. If the current model training round number reaches the preset training round number or the current total loss meets the preset conditions, the current network delay prediction model is used as the trained network delay prediction model. Otherwise, the parameters of the current network delay prediction model are updated based on the current total loss, and the next batch of sample timing data is used as the sample timing data of the current batch, and S410 is executed.
3. The method according to claim 2, characterized in that The total loss of the current network delay prediction model meets the following conditions: Loss=L d +L c , Loss is the total loss of the current network delay prediction model, L d is the data fitting loss of the current network delay prediction model, L d =(1 / T)∑ T t=1 ∑ w s=1 (y ts p -y ts true ) 2 ,y ts p is the network delay prediction value corresponding to the sth row of data in the tth sample time series data in the current batch of sample time series data. The value of t ranges from 1 to T, T is the number of sample time series data in the current batch, and the value of s ranges from 1 to w. ts true for y ts p The corresponding real value of network delay, L c is the network characteristic constraint loss of the current network delay prediction model, L c =λ1×L1+λ2×L2+λ3×L3, where L1 is the load constraint loss, L2 is the decay constraint loss, L3 is the window constraint loss, and λ1, λ2 and λ3 are weight hyperparameters respectively.
4. The method according to claim 3, characterized in that L1=(1 / T)∑ T t=1 ∑ w s=1 (y ts p -(k load ×N ts src ×S ts p ) / C), k load is the load balancing coefficient, N ts src is the number of source IP addresses corresponding to the s-th row of data in the t-th sample time series data in the current batch of sample time series data, S ts p is the data packet size corresponding to the s-th row of data in the t-th sample time series data in the current batch of sample time series data, and C is the capacity of the target network link; L2=(1 / T)∑ T t=1 ∑ w s=1 (R ts +k decay ×(S ts p -S t(s-1) p ) 2 , where R ts is the delay change rate corresponding to the s-th row of data in the t-th sample time series data in the current batch of sample time series data, k decay is the attenuation balance coefficient, L3=(1 / T)∑ T t=1 ∑ w s=1 (y ts p ×S ts p / k window ), k window is the window equilibrium constant.
5. The method according to claim 1, characterized in that The network transmission parameters include the number of source IPs, the number of destination IPs, the data packet size, and the source address window size.
6. The method according to claim 1, characterized in that The target network link is a civil aviation data network link.
7. A training device for a network delay prediction model, characterized in that: include: A model building module, used to build an initial network delay prediction model, wherein the network delay prediction model includes a convolution module, a first drop layer, a first feature extraction network, a second drop layer, a second feature extraction network and an output prediction layer connected in sequence, wherein the first feature extraction network includes a first LSTM module and a first multi-head attention module, and the second feature extraction network includes a second LSTM module and a second multi-head attention module; A first data acquisition module is used to acquire a historical time series data set, wherein the historical time series data set is a data set obtained by sorting a plurality of network flow data of a target network link collected within a set time period in time series, and the data entity of the network flow data includes a network transmission parameter and a network delay; A second data acquisition module is used to divide the historical time series data set into a plurality of sample time series data using a sliding time window with a sliding step of 1 and a window size of w, each time series data is a w×m matrix, where m is the number of data entities of the network traffic data; A model training module is used to use the numerical values of network transmission parameters in sample time series data as input information, and the numerical values of network delay in the sample time series data as label information, and use the multiple sample time series data to train an initial network delay prediction model to obtain a trained network delay prediction model, wherein during the training process, the total loss of the network delay prediction model includes data fitting loss and network characteristic constraint loss, and the network characteristic constraint loss includes a load constraint loss determined based on the network load, an attenuation constraint loss determined based on the network burst traffic, and a window constraint loss determined for a receiving window based on the network protocol.
8. A network status detection method, characterized in that: The method comprises the following steps: S10, obtaining the network traffic data of the target network link collected at the current monitoring time as the current data to be processed; S20, inputting the numerical value of the network transmission parameter in the current data to be processed into the trained network delay prediction model obtained by the method described in any one of claims 1 to 6, to obtain the network delay prediction value corresponding to the current data to be processed; S30, comparing the predicted network delay value corresponding to the current data to be processed with the corresponding actual network delay value, and determining the network state of the target network link at the current monitoring time based on the comparison result.
9. The network status detection method according to claim 8, characterized in that: S30 specifically includes: If the network delay prediction value corresponding to the current data to be processed is less than the corresponding network delay actual value, or if the difference between the network delay prediction value corresponding to the current data to be processed and the corresponding network delay actual value is less than or equal to the set difference threshold, it is determined that the target network link is in a normal state; if the difference between the network delay prediction value corresponding to the current data to be processed and the corresponding network delay actual value is greater than the set difference threshold, it is determined that the target network link is in an abnormal state.
Citation Information
Patent Citations
Training method and device of delay prediction model and congestion control method and device
CN113992599B
Optimization method of efficient throughput capacity of multipath parallel transmission system
CN105873096A
Time delay prediction model training method and device, and congestion control method and device
CN113992599A
Network capacity planning
US20070067296A1