Network packet loss model classification identification and multi-parameter calculation method
Through the combination of a fully connected neural network and a deep neural network framework, the problem of being difficult to accurately identify packet loss model categories and calculate internal parameters in complex network environments is solved, and high-precision packet loss model classification and parameter calculation are realized.
Patent Information
- Application Number
- CN202510110677.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-23
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-01-23
AI Technical Summary
In a complex and changeable actual network environment, it is difficult to accurately identify the categories of packet loss models and their internal parameter values. The classification accuracy of the prior art is low and the parameter solution is inaccurate.
A fully connected neural network is used to train multi-dimensional features to improve the accuracy of packet loss model type recognition, and a deep neural network framework combining multi-path convolution and bidirectional long and short-term memory network is designed to accurately calculate the internal parameter values of packet loss model.
The accuracy of packet loss model type recognition and the accuracy of internal parameter calculation are significantly improved, and the problems of low classification accuracy and inaccurate parameter solution in the prior art are overcome.
Smart Images

Figure CN119945945A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of network communications, and in particular relates to a network packet loss model classification identification and multi-parameter calculation method. Background Art
[0002] Among the many indicators of network performance evaluation, packet loss, as one of the key factors to measure data transmission reliability and network quality of service (QoS), has long been an important focus in the field of network engineering research. The packet loss phenomenon is not only directly related to the integrity and reliability of data transmission, but also deeply reflects the degree of network congestion, quality of service (QoS) and potential failure points, and plays an irreplaceable role in the evaluation and diagnosis of network health status. As an important tool to explain the correlation between discrete packet loss events, the packet loss model has become increasingly important. In recent years, with the continuous expansion of the classic packet loss model and the emergence of a series of new models, the existing packet loss model system has been greatly enriched, and now it can more comprehensively and accurately describe various packet loss characteristics and their distribution laws in actual network environments. These models not only provide a powerful means to accurately outline network packet loss behavior, but also lay a solid theoretical foundation for the precise measurement of network packet loss events, the forward-looking prediction of network behavior, and the rapid and accurate location of faults.
[0003] However, in actual network scenarios, different types of packet loss models vary greatly, and their internal parameters are also different. Taking the Gilbert-Elliot model as an example, it has 2 states and involves 4 parameters; the Adaptation of Extended Gilbert model has 8 states and 8 parameters; the Four-state Hidden Markov model has 4 states and 20 parameters. Therefore, how to accurately identify the category of packet loss models in complex and changeable actual network environments and how to accurately calculate the internal parameter values of packet loss models have become urgent problems to be solved. Unfortunately, the current research on network packet loss model type identification faces many challenges. Although there are many feature extraction methods such as positive and negative coding, state transition, and RGB images, these methods do not perform well in complex packet loss model classification tasks, and the classification accuracy can only reach about 50%, which is difficult to meet the high-precision requirements of practical applications. In addition, there are obvious problems in solving the internal parameters of the packet loss model. Although the expectation-maximization (EM) algorithm has been widely used, its sensitivity to initial parameter settings and the limitation of being easily trapped in local optimal solutions affect the accuracy and reliability of the solution results. As an advanced machine learning method, reinforcement learning has shown certain potential in the field of parameter solution. However, for Markov models with complex multi-implicit states, due to their large state and action spaces, using reinforcement learning for parameter solution will lead to a sharp increase in computational complexity, making it difficult to be widely used in actual network operation and maintenance. Summary of the invention
[0004] In order to solve the problem of how to accurately identify the category of packet loss model in a complex and changeable actual network environment and how to accurately calculate the internal parameter value of the packet loss model, the present invention provides a network packet loss model classification identification and multi-parameter calculation method. Therefore, the present invention aims to overcome the limitations of the above-mentioned prior art and proposes the following innovative solutions:
[0005] (1) Extraction of multi-dimensional features (time domain, frequency domain, and statistical domain features) for multiple types of complex packet loss models: By generating packet loss sequences that match multiple complex packet loss models, we extract time domain, frequency domain, and statistical features from them and fuse these features together. We then use a fully connected neural network (FCNN) to train the fused features, thereby significantly improving the accuracy of packet loss model type recognition.
[0006] (2) A deep neural network framework for calculating parameters of multiple types of complex packet loss models: In order to solve the problems existing in the existing packet loss model parameter solution methods, the present invention designs a deep neural network framework that combines multi-path convolution (ResNeXt) with a bidirectional long short-term memory network (BiLSTM). The framework is mainly composed of multi-path convolution ResNeXt, BatchNormalization layer, ReLU activation function layer and BiLSTM. The multi-branch structure of ResNeXt enhances the feature representation capability, and BiLSTM effectively captures the time dependency of time series data, thereby realizing the accurate calculation of the internal parameters of the packet loss model.
[0007] In order to achieve the above object, the present invention adopts the following technical solutions:
[0008] A network packet loss model classification identification and multi-parameter calculation method, the method comprising the following steps:
[0009] Step 1: For various complex packet loss models, various parameters are set to generate data sets, and preprocessing operations are performed on them to extract time domain, frequency domain and statistical domain features, and then fuse them to form a multi-dimensional feature vector; the specific steps are:
[0010] Step 1.1: Use random numbers as judgment conditions, compare the random numbers generated by the system with the state transition probability and output probability of each packet loss model, realize the transition between states, and generate the bit value sequence under the corresponding state;
[0011] Step 1.2: Extract multi-dimensional features of the bit value sequence, where the multi-dimensional features include three domains: time domain, frequency domain and statistical domain;
[0012] Step 1.2.1: Time domain feature extraction, the main goal is to capture the dynamic change characteristics of the packet loss bit value data in the time dimension. To this end, the original binary bit value sequence is converted into an integer time series, and the following key features are extracted from it:
[0013] Zero Crossing Rate (ZCR): It is the frequency of the quantized bit value sequence crossing the zero point, reflecting the conversion frequency from "0" to "1" or "1" to "0". The formula is as follows:
[0014]
[0015] Where N is the length of the data sequence, s i is the ith data point, (·) is the indicator function, which takes the value of 1 when the condition is met and 0 otherwise.
[0016] Number of Max Peaks: This feature focuses on the local maximum points in the data sequence, which can reveal sudden events in the packet loss pattern. This feature is obtained by detecting the local maximum points (i.e. peaks) in the bit value sequence and counting the total number of these peaks. The calculation formula is as follows:
[0017] Peaks = |find_peaks(data)|
[0018] Among them, find_peaks is a function for finding peaks, and data is a data sequence.
[0019] Number of Min Valleys: Corresponding to the number of maximum peaks, this feature focuses on the local minimum points (valleys) in the bit value sequence, reflecting the number of significant drops or low activity levels in the sequence. The number of valleys is obtained by taking the negative value of the bit value sequence and using a similar peak detection method to find the local minimum points. The calculation formula is as follows:
[0020] Valleys=|find_peaks(-data)|
[0021] Among them, find_peaks is a function for finding peaks, and data is a data sequence.
[0022] Step 1.2.2: Frequency domain feature extraction aims to reveal the periodic components behind the bit value sequence. The fast Fourier transform (FFT) algorithm is used to convert the bit value sequence into a frequency domain representation, and the amplitude spectrum is further calculated and the first half of the spectrum is selected to highlight the importance of the positive frequency components. In particular, the first ten frequency points with the largest amplitude are focused on to characterize the most significant periodic features in the bit value sequence. The specific implementation steps are as follows:
[0023] Step 1.2.1.1: Perform a fast Fourier transform (FFT) on the bit value sequence data to obtain a frequency domain representation frequency_data in complex form:
[0024] frequency_data=FFT(data)
[0025] Step 1.2.1.2: Calculate the amplitude spectrum of the frequency domain data, that is, the absolute value of the frequency domain data:
[0026] amplitude_spectrum=∣frequency_data∣
[0027] Step 1.2.1.3: Take the first half of the spectrum and keep the first half of the spectrum (i.e. the positive frequency part):
[0028] half_spectrum=amplitude_spectrum[:len(amplitude_spectrum) / / 2]
[0029] Step 1.2.1.4: Identify the top 10 frequency points with the largest amplitude. In order to identify the most significant periodic components in the signal, use the sorting method to find the indexes of the top 10 frequency points with the largest amplitude, and extract the corresponding amplitude values from half_spectrum:
[0030] max_magnitude_indices=argsort(half_spectrum)[::-1][:10]
[0031] max_magnitude_values=half_spectrum[max_magnitude_indices]
[0032] Step 1.2.3: Extract statistical domain features, which reflects the packet loss phenomenon and its characteristics in the network according to the frequency of occurrence of specific patterns in the bit value sequence; the specific patterns are: "01", "10", "00";
[0033] "01": represents the beginning of packet loss, that is, the number of transitions from no packet loss to packet loss. It can reflect the frequency of sudden packet loss in the network.
[0034] "10": represents the end of packet loss, that is, the number of transitions from packet loss to normal state. This helps to understand the network's ability to recover from abnormal state.
[0035] "00": indicates that there is no packet loss and the frequency of packet loss can reflect the stability and reliability of the network.
[0036] Step 1.3: Concatenate the feature vectors from different analysis domains to construct a new high-dimensional feature vector as the data set for subsequent classification tasks.
[0037] Step 1.3.1: Initialize the feature vector: Represent the features extracted from each analysis domain (time domain, frequency domain, statistical domain) as three different vectors: time domain feature vector F time , frequency domain feature vector F freq , Statistical domain eigenvector F stat According to the previous steps, the specific eigenvalues contained in these eigenvectors are as follows:
[0038] F time =[ZCR,Peaks,Valleys]
[0039] F freq =[A1,A2,...,A 10 ]
[0040] F stat =[N 01 ,N 10 ,N 00 ]
[0041] Among them, ZCR is the zero crossing rate, Peaks and Valleys are the maximum peak number and the minimum peak number respectively; A i Indicates the amplitude values of the first ten frequency points with the largest amplitude. 01 ,N 10 ,N 00 These are the number of occurrences of the patterns “01”, “10”, and “00” respectively.
[0042] Step 1.3.2: To construct the final multidimensional feature vector F, concatenate the above three feature vectors in sequence:
[0043] F=[F time ; F freq ; F stat ]
[0044] Specifically, if the dimensions of each feature vector are d time , d freq , d stat , then the dimension of the final feature vector F is:
[0045] d=d time +d freq +d stat
[0046] Step 2: A multi-model classification method for network packet loss that integrates the time-frequency domain and the statistical domain is proposed. A fully connected neural network model is constructed. The multi-dimensional feature vector obtained in step 1 is input into the fully connected neural network model. The deep characteristics of the data are gradually mined with the help of the fully connected layer to achieve the classification of the network packet loss model. The specific steps are as follows:
[0047] Step 2.1: Data preprocessing, ensure that all features are standardized so that they are comparable on the same scale, which helps improve training efficiency and classification accuracy. For each feature x i , normalized by the following formula:
[0048]
[0049] Among them, μ is the mean of the feature and σ is the standard deviation.
[0050] Step 2.2: Build fully connected layers and use fully connected layers (FC layers) to gradually explore the deep characteristics of the data. Each neuron receives all the outputs from the previous layer and introduces nonlinear transformations through activation functions (such as ReLU). The fully connected layer can learn the complex relationship between features, thereby enhancing the expressive power of the model. It is expressed as follows:
[0051] a [l+1] =ReLU(W [l+1] a [l] +b [l+1] )
[0052] Among them, W [l+1] and b [l+1] are the weight matrix and bias vector respectively; a [l] represents the input vector;
[0053] Step 2.3: Classifier design, add an output layer with the number of neurons equal to the number of packet loss models to be classified. Use the softmax function as the activation function, which converts the score of each category into a probability distribution so that the sum of the output values is 1; the specific calculation is as follows:
[0054] z [L] =W [L] a [L-1] +b [L]
[0055]
[0056] Among them, z [L] represents the linear combination output of the Lth layer (the last layer), W [L] is the weight matrix connecting the L-1th layer and the Lth layer, a [L-1] is the activation output of layer L-1, b [L] is the bias vector of the Lth layer, represents the probability of the i-th category;
[0057] Step 2.4: Model training: Using the samples in the generated dataset and their corresponding labels (i.e. the actual packet loss model type), adjust the weight parameters through the back-propagation algorithm to minimize the loss function:
[0058]
[0059] Among them, K is the number of categories of the packet loss model, y is the true label, Represents the predicted probability distribution. The optimizer uses the Adam optimizer, whose adaptive learning rate adjustment mechanism can achieve efficient convergence in a shorter time while significantly reducing computing resource consumption;
[0060] Step 2.5: Model validation and tuning. Evaluate the performance of the model through cross-validation and other means, and adjust the hyperparameters (such as learning rate, batch size, etc.) according to the evaluation results until a satisfactory classification effect is obtained. The update formula of the Adam optimizer is:
[0061] m t =β1m t-1 +(1-β1)▽ θ L
[0062] v t =β2v t-1 +(1-β2)(▽ θ L) 2
[0063]
[0064] Among them, m t and v t Represent the first-order moment estimate of the gradient (i.e., momentum) and the second-order moment estimate (i.e., the exponentially weighted average of the square of the gradient), which are used to calculate the adaptive learning rate in the parameter update step; β1 and β2 are hyperparameters of the past gradient decay rate; ▽ θ L is the gradient of the loss function L with respect to the parameter θ; ò is a small constant to prevent division by zero; η is the learning rate, which controls the step size of each update; and is the bias-corrected moment estimate; θ t represents the parameter value at time step t;
[0065] Step 3: For various complex packet loss models, generate bit value sequence data suitable for packet loss model parameter calculation by setting various parameters, and attach the corresponding model parameter values after it; the data set generated in this way can meet the multi-parameter calculation requirements, and its packet loss data set format covers the packet loss time series and the matching model parameter values. The specific steps are as follows:
[0066] Step 3.1: Data generation: Based on the selected parameters, simulate the corresponding packet loss events and generate the packet loss time series;
[0067] Step 3.2: Label the true value. For each generated packet loss sequence, record the specific parameter configuration used to generate it as the true label of the sample. This not only helps supervised learning, but also provides a basis for direct comparison for subsequent parameter estimation.
[0068] Step 3.3: Formatting: Organize all packet loss time series and their matching model parameter values in a unified format to form a structured data set for subsequent reading, processing and analysis;
[0069] Step 4: Construct a ResNeXt-BiLSTM deep neural network framework for multi-parameter calculation of complex packet loss models. The framework is mainly composed of multi-path convolution ResNeXt, BatchNorm layer, ReLU layer and bidirectional long short-term memory network BiLSTM. Input the test set into the trained ResNeXt-BiLSTM model, and the model outputs the internal multi-parameter values of the target packet loss model. Use multiple evaluation indicators to comprehensively evaluate the model performance. The construction of the framework includes the following steps:
[0070] Step 4.1: ResNeXt Block uses a multi-path convolutional structure, which can reduce computational complexity while retaining powerful feature extraction capabilities, thereby effectively capturing local and global features in the packet loss sequence. For example, it can identify key features such as the pattern of continuous packet loss and periodic packet loss. The specific convolution process includes 1×1 convolution for dimensionality reduction, 3×3 grouped convolution to enhance feature diversity, BatchNorm layer for normalization, and ReLU activation function to introduce nonlinearity. The input feature map is Among them C in is the number of input channels, T is the length of the feature map;
[0071] Step 4.1.1: 1×1 convolution to reduce dimension, and pass through BatchNorm layer and ReLU layer, the formula is as follows:
[0072] Y1=W1*X+b1
[0073]
[0074] Among them, Y1 represents the output after 1x1 convolution operation; Represents the weight matrix of the convolution kernel; b1 represents the bias term in the 1x1 convolution operation; Y2 represents the output after batch normalization (BatchNorm) and ReLU activation function processing; γ1 is the scaling parameter of BatchNormalization; μ1 is the mean after convolution; is the variance after convolution; ò is a constant to prevent division by zero; β1 is the translation parameter of BatchNormalization; D is the intermediate dimension; cardinality is the number of paths; C out is the number of output channels; widen_factor is the width expansion factor;
[0075] Step 4.1.2: Perform group convolution operation on Y1, the formula is as follows:
[0076] Y3=GroupConv(W2,Y2)+b2
[0077]
[0078] Among them, the grouped convolution divides the intermediate dimension D into G groups, each group size is G is the number of groups; W2 is the weight matrix of the convolution kernel; Y3 is the output after the group convolution operation; b2 is the bias term of the group convolution operation; Y4 is the output after batch normalization (BatchNorm) and ReLU activation function processing; μ2 is the mean after convolution; is the variance after convolution; γ2 is the scaling parameter of BatchNormalization; β2 is the translation parameter of BatchNormalization; ò is a constant to prevent division by zero;
[0079] Step 4.1.3: 1×1 convolution restores the number of channels and restores the feature map after group convolution to the original number of channels so that it can be connected with subsequent layers. The formula is as follows:
[0080] Y5=W3*Y4+b3
[0081]
[0082] in, Represents the weight matrix of the convolution kernel; Y5 represents the output of the 1x1 convolution operation; b3 represents the bias term of the 1x1 convolution; μ3 is the mean after convolution; is the variance after convolution; γ3 is the scaling parameter of BatchNormalization; β3 is the translation parameter of BatchNormalization; Y6 represents the output of the batch normalization operation; ò is a constant to prevent division by zero;
[0083] Step 4.1.4: Perform residual link and ReLU activation on Y6 to retain input information and prevent information loss; ReLU activation function increases nonlinearity and improves model expression ability:
[0084] Y out =RELU(Y6+X residual )
[0085] Among them, X residual is the residual branch adjusted by downsampling; Y out Represents the output after residual connection and ReLU activation;
[0086] When the input and output dimensions are different, downsample the input:
[0087] X residual =Conv1d(X)+BN(X)
[0088] Step 4.2: Input the feature map obtained by two ResNeXtBlocks into the global pooling layer. The pooling operation compresses the feature map to a fixed size to accommodate subsequent LSTM processing. The formula is as follows:
[0089]
[0090] Among them, L is the length of the input feature map; Y pool Represents the feature vector after global pooling;
[0091] Step 4.3: Treat the globally pooled feature vector as the input of a sequence and process it using the BiLSTM layer. By combining the results of the forward and backward LSTMs, the BiLSTM is able to capture long-term dependencies in the packet loss sequence. This helps understand the temporal dependencies of packet loss events, such as identifying situations where packet loss is more frequent in certain time periods. The calculation formula is as follows:
[0092]
[0093] In the formula, represents the hidden state of the forward LSTM at time step t, represents the hidden state of the backward LSTM at time step t; σ represents the activation function; W f , W b Represents the weight matrices of the forward and reverse LSTM layers respectively; U f , U b Represents the weight matrices of the forward and reverse LSTM layers for the hidden state of the previous time step; b f 、b b denote the bias terms of the forward and reverse LSTM layers respectively; h t represents the output of bidirectional LSTM; X pool,t represents the input sequence;
[0094] Step 4.4: After the BiLSTM layer processes the sequence features, take the hidden state h of the last time step t As input, it passes through the fully connected layer to generate the final output Y final , used to predict the multi-parameter values of the packet loss model, the calculation formula is as follows:
[0095] Y final =W fc ·h t +b fc
[0096] In the formula, is the weight matrix of the fully connected layer; b fc is bias;
[0097] Step 4.5: Use mean square error MSE as the loss function, and its formula is:
[0098]
[0099] Where n is the number of samples; y i is the true value of the i-th sample; y i ′ is the predicted value of the i-th sample;
[0100] The optimization algorithm uses the Adam optimizer, and its update formula is:
[0101]
[0102] Among them, m t and v t are the first and second moments of the gradient respectively; η is the learning rate; ò is a constant to prevent division by zero; θ t Indicates the updated parameter value;
[0103] Step 4.6: Use multiple metrics to evaluate model performance:
[0104] 1) Mean square error (MSE) measures the average value of the square error between the predicted value and the true value. The smaller the value, the closer the prediction result of the model is to the true value. The formula is as follows:
[0105]
[0106] Where n is the number of samples; y i is the true value of the i-th sample; y i ′ is the predicted value of the i-th sample;
[0107] 2) Mean absolute error (MAE), the average value of the absolute error between the predicted value and the true value, is used to measure the average error of the model prediction results. The formula is as follows:
[0108]
[0109] Where n is the number of samples; y i is the true value of the i-th sample; y i ′ is the predicted value of the i-th sample;
[0110] 3) Coefficient of determination R-squared, which is used to measure the explanatory power of the model, that is, the proportion of the model explaining the variation of the independent variable to the dependent variable; R 2 The value range is [0, 1], where 1 means the model predicts perfectly and 0 means the model has no explanatory power. The formula is as follows:
[0111]
[0112] Where n is the number of samples; y i is the true value of the i-th sample; y i ′ is the predicted value of the i-th sample; is the average of the true values.
[0113] Step 4.6: The test set is input into the trained ResNeXt-BiLSTM model, and the model outputs the internal multi-parameter values of the target packet loss model, and the model performance is comprehensively evaluated using multiple evaluation indicators.
[0114] Compared with the prior art, the present invention has the following advantages:
[0115] 1. Combining Fourier transform, time series feature extraction and statistical analysis, the present invention realizes efficient classification of multiple network packet loss models. This method can fully capture the characteristics of data in the time domain, frequency domain and statistical domain, accurately identify different types of packet loss models, significantly improve the classification accuracy, and overcome the problem of low classification accuracy of traditional methods;
[0116] 2. A deep neural network framework combining ResNeXt and BiLSTM was constructed. This framework fully utilizes the advantages of multi-path convolution (ResNeXt) in feature extraction and the powerful ability of bidirectional long short-term memory network (BiLSTM) in processing long-term dependencies of sequence data. This innovation fills the gap in deep learning in the field of multi-parameter calculation of network packet loss models and provides a new technical means for network performance optimization and fault diagnosis. BRIEF DESCRIPTION OF THE DRAWINGS
[0117] Figure 1 It is a flow chart of a network packet loss model classification identification and multi-parameter calculation method;
[0118] Figure 2 This is the ResNeXt-BiLSTM model structure diagram;
[0119] Figure 3 It is a density scatter plot between the results calculated based on the CNN model and the true value;
[0120] Figure 4 It is a density scatter plot between the results calculated based on the CNN-LSTM model and the true value;
[0121] Figure 5 It is a density scatter plot between the results calculated based on the CNN-BiLSTM model and the true value;
[0122] Figure 6 It is a density scatter plot between the results calculated based on the ResNeXt model and the true value;
[0123] Figure 7 It is a density scatter plot between the results calculated based on the ResNeXt-LSTM model and the true value;
[0124] Figure 8 It is a density scatter plot between the results calculated based on the ResNeXt-BiLSTM model and the true value;
[0125] Fig. 9 Error distribution diagram of multi-parameter calculation of network packet loss model for each neural network model. DETAILED DESCRIPTION
[0126] In order to gain a deeper understanding of the present invention, we will provide a comprehensive and detailed description of the present invention. However, the present invention has multiple implementations and is not limited to the specific examples listed herein. The presentation of these examples is intended to deepen the comprehensive understanding of the disclosure of the present invention.
[0127] A network packet loss model classification identification and multi-parameter calculation method, the flow chart is as follows Figure 1 As shown, the following steps are included:
[0128] Step 1: For various complex packet loss models, various parameters are reasonably set to generate a data set, and preprocessing operations are performed on it to extract time domain, frequency domain and statistical domain features, and then fuse them to form a multi-dimensional feature vector. The specific steps are as follows:
[0129] Step 1.1: Use random numbers as judgment conditions and compare the random numbers generated by the system with the state transition probability and output probability of each packet loss model to achieve the transition between states and generate the bit value sequence under the corresponding state;
[0130] Step 1.2: Extract multi-dimensional features of the bit value sequence, where the multi-dimensional features include three domains: time domain, frequency domain and statistical domain;
[0131] Step 1.2.1: Time domain feature extraction, the main goal is to capture the dynamic change characteristics of the packet loss bit value data in the time dimension. To this end, the original binary bit value sequence is converted into an integer time series, and the following key features are extracted from it:
[0132] Zero Crossing Rate (ZCR): It is the frequency of the quantized bit value sequence crossing the zero point, reflecting the conversion frequency from "0" to "1" or "1" to "0". The formula is as follows:
[0133]
[0134] Where N is the length of the data sequence, s iis the ith data point, (·) is the indicator function, which takes the value of 1 when the condition is met and 0 otherwise.
[0135] Number of Max Peaks: This feature focuses on the local maximum points in the data sequence, which can reveal sudden events in the packet loss pattern. This feature is obtained by detecting the local maximum points (i.e. peaks) in the bit value sequence and counting the total number of these peaks. The calculation formula is as follows:
[0136] Peaks = |find_peaks(data)|
[0137] Among them, find_peaks is a function for finding peaks, and data is a data sequence.
[0138] Number of Min Valleys: Corresponding to the number of maximum peaks, this feature focuses on the local minimum points (valleys) in the bit value sequence, reflecting the number of significant drops or low activity levels in the sequence. The number of valleys is obtained by taking the negative value of the bit value sequence and using a similar peak detection method to find the local minimum points. The calculation formula is as follows:
[0139] Valleys=|find_peaks(-data)|
[0140] Among them, find_peaks is a function for finding peaks, and data is a data sequence.
[0141] Step 1.2.2: Frequency domain feature extraction aims to reveal the periodic components behind the packet loss sequence. The Fast Fourier Transform (FFT) algorithm is used to convert the bit value sequence into a frequency domain representation. The amplitude spectrum is further calculated and the first half of the spectrum is selected to highlight the importance of the positive frequency components. In particular, the first ten frequency points with the largest amplitude are focused on to characterize the most significant periodic features in the packet loss sequence. The specific implementation steps are as follows:
[0142] Step 1.2.1.1: Perform a fast Fourier transform (FFT) on the bit value sequence data to obtain the frequency domain representation frequency_data in complex form:
[0143] frequency_data=FFT(data)
[0144] Step 1.2.1.2: Calculate the amplitude spectrum of the frequency domain data, that is, the absolute value of the frequency domain data:
[0145] amplitude_spectrum=∣frequency_data∣
[0146] Step 1.2.1.3: Take the first half of the spectrum and keep the first half of the spectrum (i.e. the positive frequency part):
[0147] half_spectrum=amplitude_spectrum[:len(amplitude_spectrum) / / 2]
[0148] Step 1.2.1.4: Identify the top 10 frequency points with the largest amplitude. In order to identify the most significant periodic components in the signal, use the sorting method to find the indexes of the top 10 frequency points with the largest amplitude, and extract the corresponding amplitude values from half_spectrum:
[0149] max_magnitude_indices=argsort(half_spectrum)[::-1][:10]
[0150] max_magnitude_values=half_spectrum[max_magnitude_indices]
[0151] Step 1.2.3: Extract statistical domain features, which reflects the packet loss phenomenon and its characteristics in the network according to the frequency of occurrence of specific patterns in the bit value sequence; the specific patterns are: "01", "10", "00";
[0152] "01": represents the beginning of packet loss, that is, the number of transitions from no packet loss to packet loss. Reflects the frequency of sudden packet loss in the network.
[0153] "10": represents the end of packet loss, that is, the number of transitions from packet loss to normal state. This helps to understand the network's ability to recover from abnormal state.
[0154] "00": indicates that there is no packet loss and the frequency of packet loss can reflect the stability and reliability of the network.
[0155] Step 1.3: Concatenate the feature vectors from different analysis domains to construct a new high-dimensional feature vector as the data set for subsequent classification tasks.
[0156] Step 1.3.1: Initialize the feature vector: Represent the features extracted from each analysis domain (time domain, frequency domain, statistical domain) as three different vectors: time domain feature vector F time , frequency domain feature vector F freq , Statistical domain eigenvector F stat According to the previous steps, the specific eigenvalues contained in these eigenvectors are as follows:
[0157] F time =[ZCR,Peaks,Valleys]
[0158] F freq =[A1,A2,...,A 10 ]
[0159] F stat =[N 01 ,N 10 ,N 00 ]
[0160] Among them, ZCR is the zero crossing rate, Peaks and Valleys are the maximum peak number and the minimum peak number respectively; A i Indicates the amplitude values of the first ten frequency points with the largest amplitude. 01 ,N 10 ,N 00 These are the number of occurrences of the patterns “01”, “10”, and “00” respectively.
[0161] Step 1.3.2: To construct the final multidimensional feature vector F, concatenate the above three feature vectors in sequence:
[0162] F=[F time ; F freq ; F stat ]
[0163] Specifically, if the dimensions of each feature vector are d time , d freq , d stat , then the dimension of the final feature vector F is:
[0164] d=d time +d freq +d stat
[0165] Step 2: A multi-model classification method for network packet loss that integrates the time-frequency domain and the statistical domain is proposed. A fully connected neural network model is constructed. The multi-dimensional feature vector obtained in step 1 is input into the fully connected neural network model. The deep characteristics of the data are gradually mined with the help of the fully connected layer to achieve the classification of the network packet loss model. The specific steps are as follows:
[0166] Step 2.1: Data preprocessing, ensure that all features are standardized so that they are comparable on the same scale, which helps improve training efficiency and classification accuracy. For each feature x i , normalized by the following formula:
[0167]
[0168] Among them, μ is the mean of the feature and σ is the standard deviation.
[0169] Step 2.2: Build fully connected layers and use fully connected layers (FC layers) to gradually explore the deep characteristics of the data. Each neuron receives all the outputs from the previous layer and introduces nonlinear transformations through activation functions (such as ReLU). The fully connected layer can learn the complex relationship between features, thereby enhancing the expressive power of the model. It is expressed as follows:
[0170] a [l+1] =ReLU(W [l+1] a [l] +b [l+1] )
[0171] Among them, W [l+1] and b [l+1] are the weight matrix and bias vector, respectively, a [l] represents the input vector;
[0172] Step 2.3: Classifier design, add an output layer with the number of neurons equal to the number of packet loss models to be classified. Use the softmax function as the activation function, which converts the score of each category into a probability distribution so that the sum of the output values is 1; the specific calculation is as follows:
[0173] z [L] =W [L] a [L-1] +b [L]
[0174]
[0175] Among them, z [L] represents the linear combination output of the Lth layer (the last layer), W [L] is the weight matrix connecting the L-1th layer and the Lth layer, a [L-1] is the activation output of layer L-1, b [L] is the bias vector of the Lth layer, represents the probability of the i-th category;
[0176] Step 2.4: Model training: Using the samples in the generated dataset and their corresponding labels (i.e. the real packet loss model type), adjust the weight parameters through the back propagation algorithm to minimize the loss function:
[0177]
[0178] Among them, K is the number of categories of the packet loss model, y is the true label, represents the probability distribution of the prediction.
[0179] Step 2.5: Model verification and tuning. Evaluate the performance of the model through cross-validation and other means, and adjust the hyperparameters (such as learning rate, batch size, etc.) according to the evaluation results until a satisfactory classification effect is obtained. The optimizer uses the Adam optimizer, whose adaptive learning rate adjustment mechanism can achieve efficient convergence in a shorter time and significantly reduce computing resource consumption. The update formula of the Adam optimizer is:
[0180] m t =β1m t-1 +(1-β1)▽ θ L
[0181] v t =β2v t-1 +(1-β2)(▽ θ L) 2
[0182]
[0183] Among them, m t and v t Represent the first-order moment estimate of the gradient (i.e., momentum) and the second-order moment estimate (i.e., the exponentially weighted average of the square of the gradient), which are used to calculate the adaptive learning rate in the parameter update step; β1 and β2 are hyperparameters of the past gradient decay rate; ▽ θ L is the gradient of the loss function L with respect to the parameter θ; ò is a small constant to prevent division by zero; η is the learning rate, which controls the step size of each update; and is the bias-corrected moment estimate; θ t represents the parameter value at time step t;
[0184] In order to verify the effectiveness of this method, we conducted extensive experiments. These experiments generated a total of 12 data sets for three packet loss rate scenarios (30%, 15% and 5%), and examined their classification results under different feature extraction strategies. The performance of the present invention is evaluated by comparing the classification accuracy of different feature extraction methods. Table 1 clearly shows the accuracy comparison of classification using different feature extraction strategies under different packet loss rates, which intuitively reflects the significant advantages of the present invention in improving classification accuracy.
[0185] Table 1: Comparison of classification accuracy of different methods for multi-network packet loss models
[0186]
[0187]
[0188] The experimental results show that under different packet loss rates (i.e., 0.3, 0.15, and 0.05), the multi-domain combination method that integrates statistical domain, time domain, and frequency domain features shows significantly higher classification accuracy than the method that relies solely on a single feature domain. Specifically, the "statistics + time-frequency domain" method achieved the highest classification accuracy under all tested packet loss rates, verifying the effectiveness of the multi-feature fusion strategy. Further analysis shows that as the packet loss rate decreases, the classification performance of the single-domain method, especially the time domain analysis, has significantly degraded. In contrast, the multi-domain combination method can maintain a high level of classification accuracy. The above findings not only demonstrate the superior performance of the multi-domain feature fusion method proposed in this study in improving the classification accuracy of the network packet loss model, but also emphasize the importance of combining multiple features to enhance the classification effect.
[0189] Step 3: For various complex packet loss models, generate bit value sequence data suitable for packet loss model parameter calculation by setting diversified parameters, and attach the corresponding model parameter values after it; the specific steps are as follows:
[0190] Step 3.1: Data generation: Based on the selected parameters, simulate the corresponding packet loss events and generate the packet loss time series;
[0191] Step 3.2: Label the true value. For each generated packet loss sequence, record the specific parameter configuration used to generate it as the true label of the sample. This not only helps supervised learning, but also provides a basis for direct comparison for subsequent parameter estimation.
[0192] Step 3.3: Formatting: Organize all packet loss time series and their matching model parameter values in a unified format to form a structured data set for subsequent reading, processing and analysis;
[0193] Step 4: Construct a ResNeXt-BiLSTM deep neural network framework for multi-parameter calculation of complex packet loss models. The framework is mainly composed of multi-path convolution ResNeXt, BatchNorm layer, ReLU layer and bidirectional long short-term memory network BiLSTM; it aims to achieve efficient calculation of multi-parameters of network packet loss models. The construction of the framework includes the following steps:
[0194] Step 4.1: ResNeXt Block uses a multi-path convolutional structure, which can reduce computational complexity while retaining powerful feature extraction capabilities, thereby effectively capturing local and global features in the packet loss sequence. For example, it can identify key features such as the pattern of continuous packet loss and periodic packet loss. The specific convolution process includes 1×1 convolution for dimensionality reduction, 3×3 grouped convolution to enhance feature diversity, BatchNorm layer for normalization, and ReLU activation function to introduce nonlinearity. The input feature map is Among them C in is the number of input channels, L is the length of the feature map;
[0195] Step 4.1.1: 1×1 convolution to reduce dimension, and pass through BatchNorm layer and ReLU layer, the formula is as follows:
[0196] Y1=W1*X+b1
[0197]
[0198] Among them, Y1 represents the output after 1x1 convolution operation; Represents the weight matrix of the convolution kernel; b1 represents the bias term in the 1x1 convolution operation; Y2 represents the output after batch normalization (BatchNorm) and ReLU activation function processing; γ1 is the scaling parameter of BatchNormalization; μ1 is the mean after convolution; is the variance after convolution; ò is a constant to prevent division by zero; β1 is the translation parameter of BatchNormalization; D is the intermediate dimension; cardinality is the number of paths; C out is the number of output channels; widen_factor is the width expansion factor;
[0199] Step 4.1.2: Perform group convolution operation on Y1, the formula is as follows:
[0200] Y3=GroupConv(W2,Y2)+b2
[0201]
[0202] Among them, the grouped convolution divides the intermediate dimension D into G groups, each group size is G is the number of groups; W2 is the weight matrix of the convolution kernel; Y3 is the output after the group convolution operation; b2 is the bias term of the group convolution operation; Y4 is the output after batch normalization (BatchNorm) and ReLU activation function processing; μ2 is the mean after convolution; is the variance after convolution; γ2 is the scaling parameter of BatchNormalization; β2 is the translation parameter of BatchNormalization; ò is a constant to prevent division by zero;
[0203] Step 4.1.3: 1×1 convolution restores the number of channels and restores the feature map after group convolution to the original number of channels so that it can be connected with subsequent layers. The formula is as follows:
[0204] Y5=W3*Y4+b3
[0205]
[0206] in, Represents the weight matrix of the convolution kernel; Y5 represents the output of the 1x1 convolution operation; b3 represents the bias term of the 1x1 convolution; μ3 is the mean after convolution; is the variance after convolution; γ3 is the scaling parameter of BatchNormalization; β3 is the translation parameter of BatchNormalization; Y6 represents the output of the batch normalization operation; ò is a constant to prevent division by zero;
[0207] Step 4.1.4: Perform residual link and ReLU activation on Y6 to retain input information and prevent information loss; ReLU activation function increases nonlinearity and improves model expression ability:
[0208] Y out =RELU(Y6+X residual )
[0209] Among them, X residual is the residual branch adjusted by downsampling; Y out Represents the output after residual connection and ReLU activation;
[0210] When the input and output dimensions are different, downsample the input:
[0211] X residual =Conv1d(X)+BN(X)
[0212] Step 4.2: Input the feature map obtained by two ResNeXtBlocks into the global pooling layer. The pooling operation compresses the feature map to a fixed size to accommodate subsequent LSTM processing. The formula is as follows:
[0213]
[0214] Among them, L is the length of the input feature map; Y pool Represents the feature vector after global pooling;
[0215] Step 4.3: Treat the globally pooled feature vector as the input of a sequence and process it using the BiLSTM layer. By combining the results of the forward and backward LSTMs, the BiLSTM is able to capture long-term dependencies in the packet loss sequence. This helps understand the temporal dependencies of packet loss events, such as identifying situations where packet loss is more frequent in certain time periods. The calculation formula is as follows:
[0216]
[0217] In the formula, represents the hidden state of the forward LSTM at time step t, represents the hidden state of the backward LSTM at time step t; σ represents the activation function; W f , W b Represents the weight matrices of the forward and reverse LSTM layers respectively; U f , U b Represents the weight matrices of the forward and reverse LSTM layers for the hidden state of the previous time step; b f 、b b denote the bias terms of the forward and reverse LSTM layers respectively; h t represents the output of bidirectional LSTM; X pool,t represents the input sequence;
[0218] Step 4.4: After the BiLSTM layer processes the sequence features, take the hidden state h of the last time step t As input, it passes through the fully connected layer to generate the final output Y final , used to predict the multi-parameter values of the packet loss model, the calculation formula is as follows:
[0219] Y final =W fc ·h t +b fc
[0220] In the formula, is the weight matrix of the fully connected layer; b fc is bias;
[0221] Step 4.5: Use mean square error MSE as the loss function, and its formula is:
[0222]
[0223] Where n is the number of samples; y i is the true value of the i-th sample; y i ′ is the predicted value of the i-th sample;
[0224] The optimization algorithm uses the Adam optimizer, and its update formula is:
[0225]
[0226] Among them, m t and v t are the first and second moments of the gradient respectively; η is the learning rate; ò is a constant to prevent division by zero; θ t Indicates the updated parameter value;
[0227] Step 4.6: Use multiple metrics to evaluate model performance:
[0228] 1) Mean square error (MSE) measures the average value of the square error between the predicted value and the true value. The smaller the value, the closer the prediction result of the model is to the true value. The formula is as follows:
[0229]
[0230] Where n is the number of samples; y i is the true value of the i-th sample; y i ′ is the predicted value of the i-th sample;
[0231] 2) Mean absolute error (MAE), the average value of the absolute error between the predicted value and the true value, is used to measure the average error of the model prediction results. The formula is as follows:
[0232]
[0233] Where n is the number of samples; y i is the true value of the i-th sample; y i ′ is the predicted value of the i-th sample;
[0234] 3) Coefficient of determination R-squared, which is used to measure the explanatory power of the model, that is, the proportion of the model explaining the variation of the independent variable to the dependent variable; R 2 The value range is [0, 1], where 1 means the model predicts perfectly and 0 means the model has no explanatory power. The formula is as follows:
[0235]
[0236] Where n is the number of samples; y i is the true value of the i-th sample; y i ′ is the predicted value of the i-th sample; is the average of the true values.
[0237] Step 4.6: The test set is input into the trained ResNeXt-BiLSTM model, and the model outputs the internal multi-parameter values of the target packet loss model, and the model performance is comprehensively evaluated using multiple evaluation indicators.
[0238] The mean square error (MSE) and mean absolute error (MAE) are important criteria for measuring model prediction errors. The smaller their values are, the smaller the deviation between the results calculated by the model and the actual values is, which means that the model is more effective. 2 ) is a statistic for evaluating the goodness of fit of the model. The closer its value is to 1, the stronger the model's ability to explain the data, that is, the better the model effect.
[0239] The core of the multi-parameter solution task of the Gilbert-Elliot model is to accurately calculate the four key parameters (p, q, h, k) within it. In order to comprehensively compare the performance of different deep learning models on this task, Table 2 designs various performance indicators to intuitively understand the advantages and disadvantages of each model in solving parameters.
[0240] In order to better understand the accuracy of neural network in solving multi-parameter packet loss model, Figures 3 to 8 The correspondence between the calculated values and actual values of each deep learning model is vividly demonstrated in the form of density scatter plots. The density distribution of these scatter plots intuitively reveals the degree of agreement between the model calculation results and the actual values, and is a powerful auxiliary tool for evaluating model performance. When the scatter points are dense and close to the diagonal, it means that the model calculation is extremely accurate; conversely, if the scatter points are scattered or deviate from the diagonal, it indicates that there is a deviation in the model calculation.
[0241] also, Fig. 9 The error distribution diagram further refines the error performance of different deep learning models on each output parameter. This diagram intuitively shows the distribution form and range of the error, providing an important reference for evaluating model accuracy. If the error distribution is compact and biased to the left, it means that the model error is small and the performance is excellent; on the contrary, if the error distribution is broad and biased to the right, it implies that the model error is large and needs further optimization.
[0242] It is worth noting that for other types of network packet loss models, although the specific task parameters may be different, the expected effect evaluation method of their multi-parameter computing tasks is similar to that of the Gilbert-Elliot model, and similar performance indicators and visualization methods can be used to comprehensively measure the performance of deep learning models.
[0243] Table 2: Comparison of evaluation indicators of different models
[0244]
[0245]
[0246] Table 3: Comparison of calculated results and true values of different model parameters
[0247]
[0248]
[0249] According to the calculation results, the ResNeXt-BiLSTM model has good performance in terms of mean square error (MSE), mean absolute error (MAE) and determination coefficient (R 2 ) showed the best performance in all three evaluation indicators. This shows that it has the highest accuracy in solving the multi-parameter task of the Gilbert-Elliot model. From the specific parameter calculation results in Table 3, the gap between the calculated value and the true value of the ResNeXt-BiLSTM model is also relatively small, which further verifies its superiority. Its advantages can be attributed to the following aspects:
[0250] 1. Stronger feature extraction capability: The ResNeXt model is designed with grouped convolution (Cardinality), which can capture more diverse features than traditional convolutional layers. This feature enables ResNeXt to show stronger feature extraction capabilities when processing complex data. When combined with BiLSTM, this advantage can be fully utilized in time series data processing.
[0251] 2. Bidirectional sequence modeling: Compared with ordinary LSTM, BiLSTM can capture both forward and reverse time dependencies in the data. Therefore, the ResNeXt-BiLSTM model performs better when processing sequence data with bidirectional time dependencies, providing more comprehensive integration of time information.
[0252] 3. Higher generalization ability: ResNeXt's multi-path feature extraction method enables it to better adapt and generalize when facing diverse data distributions. Compared with pure CNN, LSTM or their simple combination models, ResNeXt-BiLSTM can more effectively avoid overfitting and maintain good performance on different test sets.
[0253] 4. Better training stability: ResNeXt uses a residual connection mechanism, which can alleviate the gradient vanishing problem in deep networks and ensure that the model converges more stably and quickly during training. This makes it easier for ResNeXt-BiLSTM to achieve higher accuracy in a shorter training time than other models.
[0254] 5. Strong ability to integrate spatiotemporal features: ResNeXt-BiLSTM combines the advantages of ResNeXt in spatial feature extraction and the ability of BiLSTM in temporal dependency modeling, and can better integrate spatiotemporal features, thereby improving the overall performance of the model.
[0255] 6. Improvement of comprehensive capabilities: Compared with ResNeXt-LSTM, ResNeXt-BiLSTM further enhances the modeling capabilities of bidirectional time series. Compared with using only a single model (such as CNN, LSTM) or a simple combination model (such as CNN-LSTM, ResNeXt-LSTM), ResNeXt-BiLSTM has significant improvements in feature extraction, time series modeling, generalization capabilities, etc., which makes its calculation results better than other models.
[0256] In summary, the excellent performance of the ResNeXt-BiLSTM model in various evaluation indicators is mainly due to its powerful feature extraction ability, bidirectional sequence modeling ability, good generalization ability and stable training process. Therefore, in the multi-parameter calculation task of the packet loss model, the ResNeXt-BiLSTM model can achieve better results.
[0257] The contents not described in detail in the specification of the present invention belong to the prior art known to the professional and technical personnel in the field. Although the illustrative specific embodiments of the present invention are described above to facilitate the understanding of the present invention by the technical personnel in the field, it should be clear that the present invention is not limited to the scope of the specific embodiments. For the ordinary technical personnel in the field, as long as various changes are within the spirit and scope of the present invention defined and determined by the attached claims, these changes are obvious, and all inventions and creations using the concept of the present invention are protected.
Claims
1. A network packet loss model classification identification and multi-parameter calculation method, characterized in that: The method comprises the following steps: Step 1: For various complex packet loss models, various parameters are set to generate data sets, and preprocessing operations are performed on them to extract time domain, frequency domain and statistical domain features, and then fuse them to form a multi-dimensional feature vector; Step 2: A multi-model classification method for network packet loss that integrates the time-frequency domain and the statistical domain is proposed. A fully connected neural network model is constructed. The multi-dimensional feature vector obtained in step 1 is input into the fully connected neural network model. The deep characteristics of the data are gradually mined with the help of the fully connected layer to achieve the classification of the network packet loss model. Step 3: For various complex packet loss models, generate bit value sequence data suitable for packet loss model parameter calculation by setting diversified parameters, and attach the corresponding model parameter values after it; Step 4: Construct a ResNeXt-BiLSTM deep neural network framework for multi-parameter calculation of complex packet loss models. The framework is mainly composed of multi-path convolution ResNeXt, BatchNorm layer, ReLU layer and bidirectional long short-term memory network BiLSTM. Input the test set into the trained ResNeXt-BiLSTM model, and the model outputs the internal multi-parameter values of the target packet loss model. Use multiple evaluation indicators to comprehensively evaluate the model performance.
2. According to claim 1, a network packet loss model classification identification and multi-parameter calculation method is characterized in that: In step 1, various parameters are reasonably set for various complex packet loss models to generate a data set, and preprocessing operations are performed on it to extract time domain, frequency domain and statistical domain features from it, and then fuse them to form a multi-dimensional feature vector; the specific steps are: Step 1.1: Use the random number generated by the system as the judgment condition, compare the random number with the state transition probability and output probability of each packet loss model, realize the transition between states, and generate the bit value sequence under the corresponding state; Step 1.2: Extract multi-dimensional features of the bit value sequence, where the multi-dimensional features include three domains: time domain, frequency domain and statistical domain; Step 1.2.1: Time domain feature extraction, convert the original binary bit value sequence into a time series in integer form, and extract the following features from it: Zero Crossing Rate ZCR: The frequency of the quantized bit value sequence crossing the zero point, reflecting the conversion frequency from "0" to "1" or "1" to "0". The formula is as follows: Where N is the length of the bit value sequence, s i is the i-th bit value, (·) is the indicator function, which takes the value 1 when the condition is met and 0 otherwise; Maximum number of peaks: This feature is obtained by detecting the peaks in the bit value sequence and counting the total number of peaks. The calculation formula is as follows: Peaks = |find_peaks(data)| Among them, find_peaks is a function for finding peaks, and data is a sequence of bit values; Minimum number of peaks: The number of valley values is obtained by taking the negative value of the bit value sequence and using the peak detection method to find the valley value. The calculation formula is as follows: Valleys=|find_peaks(-data)| Among them, find_peaks is a function for finding peaks, and data is a data sequence; Step 1.2.2: Frequency domain feature extraction; use the Fast Fourier Transform (FFT) algorithm to convert the bit value sequence into a frequency domain representation, calculate the amplitude spectrum and select the first half of the spectrum to highlight the importance of the positive frequency components; the specific implementation steps are as follows: Step 1.2.1.1: Perform a fast Fourier transform on the bit value sequence data to obtain the frequency domain representation frequency_data in complex form: frequency_data=FFT(data) Step 1.2.1.2: Calculate the amplitude spectrum of the frequency domain data, that is, the absolute value of the frequency domain data: amplitude_spectrum=∣frequency_data∣ Step 1.2.1.3: Take the first half of the spectrum and keep the first half of the spectrum, that is, the positive frequency part: half_spectrum=amplitude_spectrum[:len(amplitude_spectrum) / / 2] Step 1.2.1.4: Use the sorting method to find the indices of the top 10 frequency points with the largest amplitudes, and extract the corresponding amplitude values from half_spectrum: max_magnitude_indices=argsort(half_spectrum)[::-1][:10] max_magnitude_values=half_spectrum[max_magnitude_indices] Step 1.2.3: Extract statistical domain features, which reflects the packet loss phenomenon and its characteristics in the network according to the frequency of occurrence of specific patterns in the bit value sequence; the specific patterns are: "01", "10", "00"; "01": represents the beginning of packet loss, that is, the number of transitions from no packet loss to packet loss; reflects the frequency of sudden packet loss in the network; "10": represents the end of packet loss, that is, the number of transitions from packet loss to normal state; reflects the ability of the network to recover from abnormal state; "00": indicates that there is no packet loss phenomenon. The frequency of this phenomenon reflects the stability and reliability of the network. Step 1.3: Concatenate the feature vectors from different analysis domains to construct a new high-dimensional feature vector as the data set for subsequent classification tasks; Step 1.3.1: Initialize feature vectors: Represent the features extracted from each analysis domain as three different vectors: time domain feature vector F time , frequency domain feature vector F freq , Statistical domain eigenvector F stat ; According to the previous steps, the specific eigenvalues contained in these eigenvectors are as follows: F time =[ZCR,Peaks,Valleys] <h2 style=";text-align:left;direction:ltr">F<h2 style=";text-align:left;direction:ltr"> freq <h2 style=";text-align:left;direction:ltr"> (A1,A2,...,A)<h2 style=";text-align:left;direction:ltr"> 10 <h2 style=";text-align:left;direction:ltr"> ] F stat =[N 01 ,N 10 ,N 00 ] Among them, ZCR is the zero crossing rate, Peaks and Valleys are the maximum peak number and the minimum peak number respectively; A i Indicates the amplitude values of the first ten frequencies with the largest amplitude; N 01 ,N 10 ,N 00 are the number of occurrences of the patterns "01", "10", and "00" respectively; Step 1.3.2: To construct the final multidimensional feature vector F, concatenate the above three feature vectors in sequence: F=[F time ;F freq ;F stat ] Specifically, when the dimensions of each feature vector are d time , d freq , d stat , then the dimension d of the final feature vector F is: d=d time +d freq +d stat 。 3. According to claim 2, a network packet loss model classification identification and multi-parameter calculation method is characterized in that: The step 2: proposes a network packet loss multi-model classification method that integrates the time-frequency domain and the statistical domain, constructs a fully connected neural network model, inputs the multi-dimensional feature vector obtained in step 1 into the fully connected neural network model, and gradually mines the deep characteristics of the data with the help of the fully connected layer. The specific steps for classifying the network packet loss model are as follows: Step 2.1: Data preprocessing; for each feature x i , normalized by the following formula: Among them, μ is the mean of the feature and σ is the standard deviation; Step 2.2: Build a fully connected layer and use it to gradually explore the deep characteristics of the data; it is expressed as follows: a [l+1] =ReLU(W [l+1] a [l] +b [l+1] ) Among them, W [l+1] and b [l+1] are the weight matrix and bias vector respectively; a [l] represents the input vector; Step 2.3: Design the classifier. Add an output layer whose number of neurons is equal to the number of packet loss models to be classified. Use the softmax function as the activation function to convert the score of each category into a probability distribution so that the sum of the output values is 1. The specific calculation is as follows: z [L] =W [L] a [L-1] +b [L] Among them, z [L] represents the linear combination output of the Lth layer, W [L] is the weight matrix connecting the L-1th layer and the Lth layer, a [L-1] is the activation output of layer L-1, b [L] is the bias vector of the Lth layer, represents the probability of the i-th category; Step 2.4: Model training: Using the samples in the generated dataset and their corresponding labels, i.e. the real packet loss model type, the weight parameters are adjusted through the back propagation algorithm to minimize the loss function: Among them, K is the number of categories of the packet loss model, y is the true label, represents the predicted probability distribution; Step 2.5: Model verification and tuning. Evaluate the performance of the model through cross-validation. Adjust the hyperparameters according to the evaluation results until a satisfactory classification effect is obtained. The optimizer uses the Adam optimizer, whose adaptive learning rate adjustment mechanism can achieve efficient convergence in a shorter time and significantly reduce computing resource consumption. The update formula of the Adam optimizer is: Among them, m t and v t Represent the first-order moment estimate and second-order moment estimate of the gradient, which are used to calculate the adaptive learning rate in the parameter update step; β1 and β2 are hyperparameters of the past gradient decay rate; is the gradient of the loss function L with respect to the parameter θ; is a small constant to prevent division by zero; η is the learning rate, which controls the step size of each update; and is the bias-corrected moment estimate; θ t represents the parameter value at time step t.
4. The network packet loss model classification identification and multi-parameter calculation method according to claim 3 is characterized in that: The specific steps of step 3: generating bit value sequence data suitable for packet loss model parameter calculation by setting diversified parameters for various complex packet loss models, and appending corresponding model parameter values thereto are as follows: Step 3.1: Data generation: Based on the selected parameters, simulate the corresponding packet loss events and generate the packet loss time series; Step 3.2: Label the true value. For each generated packet loss sequence, record the specific parameter configuration used to generate it as the true label of the sample. Step 3.3: Formatting: Organize all packet loss timing sequences and their matching model parameter values in a unified format to form a structured data set for subsequent reading, processing, and analysis.
5. According to claim 4, a network packet loss model classification identification and multi-parameter calculation method is characterized in that: In step 4, a ResNeXt-BiLSTM deep neural network framework for multi-parameter calculation of a network packet loss model is constructed, and the construction of the framework includes the following steps: Step 4.1: ResNeXtBlock adopts a multi-path convolution structure. The specific convolution process includes 1×1 convolution for dimensionality reduction, 3×3 group convolution to enhance feature diversity, BatchNorm layer for standardization, and ReLU activation function to introduce nonlinearity; the input feature map is Among them C in is the number of input channels, T is the length of the feature map; Step 4.1.1: 1×1 convolution to reduce dimension, and pass through BatchNorm layer and ReLU layer, the formula is as follows: Y1=W1*X+b1 Among them, Y1 represents the output after 1×1 convolution operation; Represents the weight matrix of the convolution kernel; b1 represents the bias term in the 1x1 convolution operation; Y2 represents the output after batch normalization and ReLU activation function processing; γ1 is the scaling parameter of BatchNormalization; μ1 is the mean after convolution; is the variance after convolution; ò is a constant to prevent division by zero; β1 is the translation parameter of BatchNormalization; D is the intermediate dimension; cardinality is the number of paths; C out is the number of output channels; widen_factor is the width expansion factor; Step 4.1.2: Perform group convolution operation on Y1, the formula is as follows: Y3=GroupConv(W2,Y2)+b2 Among them, the grouped convolution divides the intermediate dimension D into G groups, each group size is G is the number of groups; W2 is the weight matrix of the convolution kernel; Y3 is the output after the group convolution operation; b2 is the bias term of the group convolution operation; Y4 is the output after batch normalization and ReLU activation function processing; μ2 is the mean after convolution; is the variance after convolution; γ2 is the scaling parameter of BatchNormalization; β2 is the translation parameter of Batch Normalization; ò is a constant to prevent division by zero; Step 4.1.3: 1×1 convolution restores the number of channels and restores the feature map after group convolution to the original number of channels so that it can be connected with subsequent layers; the formula is as follows: Y5=W3*Y4+b3 in, Represents the weight matrix of the convolution kernel; Y5 represents the output of the 1×1 convolution operation; b3 represents the bias term of the 1×1 convolution; μ3 is the mean after convolution; is the variance after convolution; γ3 is the scaling parameter of Batch Normalization; β3 is the translation parameter of BatchNormalization; Y6 represents the output of the batch normalization operation; ò is a constant to prevent division by zero; Step 4.1.4: Perform residual link and ReLU activation on Y6 to retain input information and prevent information loss; ReLU activation function increases nonlinearity and improves model expression ability: AND out =RELU(Y6+X residual ) Among them, X residual is the residual branch adjusted by downsampling; Y out Represents the output after residual connection and ReLU activation; When the input and output dimensions are different, downsample the input: X residual =Conv1d(X)+BN(X) Step 4.2: Input the feature map obtained by two ResNeXtBlocks into the global pooling layer. The pooling operation compresses the feature map to a fixed size to accommodate subsequent LSTM processing; the formula is as follows: Among them, L is the length of the input feature map; Y pool Represents the feature vector after global pooling; Step 4.3: Treat the globally pooled feature vector as the input of a sequence and process it using the BiLSTM layer. By combining the results of the forward and backward LSTMs, the BiLSTM is able to capture long-term dependencies in the packet loss sequence. This helps understand the temporal dependencies of packet loss events, such as identifying situations where packet loss is more frequent in certain time periods. The calculation formula is as follows: In the formula, represents the hidden state of the forward LSTM at time step t, represents the hidden state of the backward LSTM at time step t; σ represents the activation function; W f , W b Represents the weight matrices of the forward and reverse LSTM layers respectively; U f , U b Represents the weight matrices of the forward and reverse LSTM layers for the hidden state of the previous time step; b f , b b denote the bias terms of the forward and reverse LSTM layers respectively; h t represents the output of bidirectional LSTM; X pool,t represents the input sequence; Step 4.4: After the BiLSTM layer processes the sequence features, take the hidden state h of the last time step t As input, it passes through the fully connected layer to generate the final output Y final , used to predict the multi-parameter values of the packet loss model, the calculation formula is as follows: Y final =W fc ·h t +b fc In the formula, is the weight matrix of the fully connected layer; b fc is bias; Step 4.5: Use mean square error MSE as the loss function, and its formula is: Where n is the number of samples; y i is the true value of the i-th sample; y′ i is the predicted value of the i-th sample; The optimization algorithm uses the Adam optimizer, and its update formula is: Among them, m t and v t are the first and second moments of the gradient respectively; η is the learning rate; ò is a constant to prevent division by zero; θ t Indicates the updated parameter value; Step 4.6: Use multiple metrics to evaluate model performance: Mean square error (MSE) measures the average value of the square error between the predicted value and the true value. The smaller the value, the closer the prediction result of the model is to the true value. The formula is as follows: Where n is the number of samples; y i is the true value of the i-th sample; y′ i is the predicted value of the i-th sample; Mean absolute error MAE, the average value of the absolute error between the predicted value and the true value, is used to measure the average error of the model prediction results. The formula is as follows: Where n is the number of samples; y i is the true value of the i-th sample; y′ i is the predicted value of the i-th sample; The coefficient of determination R-squared is used to measure the explanatory power of the model, that is, the proportion of the model explaining the variation of the independent variable to the dependent variable; R 2 The value range is [0, 1], where 1 means the model predicts perfectly and 0 means the model has no explanatory power. The formula is as follows: Where n is the number of samples; y i is the true value of the i-th sample; y′ i is the predicted value of the ith sample, is the average of the true values.
6. A network packet loss model classification identification and multi-parameter calculation method according to claim 5, characterized in that: In step 4.6, the test set is input into the trained ResNeXt-BiLSTM model, the model outputs the internal multi-parameter values of the target packet loss model, and multiple evaluation indicators are used to comprehensively evaluate the model performance.
Citation Information
Patent Citations
Multi-sound music human sound main melody extraction method based on deep learning
CN114627892A
Method for realizing a multi-channel convolutional recurrent neural network EEG emotion recognition model using transfer learning
US20230039900A1
Cited By
Photovoltaic cell parameter identification method and system based on intelligent hierarchical model
CN120687997A