Network Monitoring Data Compression Method, Terminal Device, and Storage Medium

The non-zero parameter vector index is filtered through the deterministic CUR algorithm to avoid zero parameter updates, and combined with column and row filtering, the storage resource occupation problem of large-scale network monitoring data is solved, achieving an efficient, stable and interpretable compression effect.

CN116405404BActive Publication Date: 2025-07-29HUNAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202310462476.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-26
Publication Date
2025-07-29
Estimated Expiration
2043-04-26

AI Technical Summary

Technical Problem

The existing network monitoring system occupies too much storage resources in large-scale networks, and the existing matrix compression algorithm lacks interpretability, stability and speed, especially under noise interference.

Method used

The deterministic CUR algorithm is used to build a column selection model. By filtering non-zero parameter vector indexes, the zero parameter vector update is avoided, and the dual filtering of columns and rows is combined to achieve compression of the monitoring matrix.

Benefits of technology

The compression results of stability and interpretability are achieved, which reduces computing costs and storage resource usage, and improves compression speed and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116405404B_ABST
    Figure CN116405404B_ABST
Patent Text Reader

Abstract

The present invention discloses a network monitoring data compression method, a terminal device and a storage medium, which fully utilize the correlation of end-to-end performance monitoring data between network nodes. For a network composed of n nodes, the present invention models the end-to-end performance monitoring data of the entire network into a monitoring matrix of size n×n, and compresses the monitoring matrix by utilizing the correlation of the monitoring data; the present invention designs a column selection model based on the deterministic CUR algorithm, and directly uses the selected partial rows and columns as the compression result, so that the compression result has interpretability and stability; the present invention effectively reduces the calculation cost and improves the compression speed of the monitoring matrix by avoiding the update of zero parameter vectors; the present invention further screens the selected rows and columns according to the accuracy requirement, further reduces the data compression ratio, and makes the storage resources required for network monitoring less.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to network monitoring technology, and in particular to a network monitoring data compression method, a terminal device, and a storage medium. Background Art

[0002] Monitoring the network to track network status (such as packet loss, latency peaks, and traffic bursts, etc.) is crucial for resolving network events and ensuring the expected performance of the network. Taking a real-time video conference as an example, latency changes in milliseconds may affect the user experience. Therefore, in order to ensure the transmission quality of the network, it is necessary to closely monitor the end-to-end performance changes of the network.

[0003] Some currently known network monitoring systems, such as Ningmesh and NetNORAD, need to detect the end-to-end performance between each pair of nodes in the network. Specifically, for a network with n nodes, each node needs to send probe packets to the other n - 1 nodes for measuring performance metrics such as round-trip latency, available bandwidth, and packet loss rate. Therefore, each time a full-network measurement is performed, it will bring an O(n 2 ) storage overhead. Modern data centers usually contain hundreds of thousands of servers, and the scale is growing rapidly at an exponential rate [1] . As the network scale continues to increase, the storage resources occupied by network monitoring data are also continuously growing. Therefore, how to compress network monitoring data has become a current research hotspot.

[0004] It is observed that the network paths starting from a certain terminal node usually have overlapping path segments or pass through some common network nodes. Therefore, the end-to-end performance data of different node pairs usually has strong correlation [2] . For example, congestion on a certain link will cause higher latency for all paths passing through that link. If we model the network monitoring data as a monitoring matrix P ∈ R n×n , where n is the number of nodes in the network, the rows represent the source nodes, the columns represent the destination nodes, and the element represents the performance monitoring data obtained by sending probe packets with node i as the source node and node j as the destination node. The correlation of the monitoring data indicates that there is redundant information in the monitoring matrix. In other words, the monitoring matrix has low rankness and can be compressed to reduce the occupation of storage resources.

[0005] Truncated SVD is a most common matrix compression method. For an n×n matrix X, truncated SVD decomposes X into the form of the product of three matrices, that is where is a diagonal matrix composed of the first k singular values, U k ∈ R n×k and Vk ∈R n×k is a singular matrix composed of the left and right singular vectors corresponding to the first k singular values, and k << n. After the truncated SVD algorithm, the three factor matrices (U k , V k , ∑ k ) are stored as the compression result; when the complete data matrix is needed, is used as the optimal rank-k approximation of X. When the sum of the data volumes of U k , V k and ∑ k is much smaller than the data volume of X, that is, 2nk + k 2 << n 2 , the compression of X is achieved. However, the singular vectors obtained by truncated SVD lack interpretability, that is, each value in the singular vector lacks the physical meaning corresponding to the row and column attributes of the matrix, which brings difficulties to further analyze and interpret the data [3] .

[0006] To address this issue, experts in the relevant field attempt to utilize the self-expression ability of data and approximately represent the low-rank matrix under investigation with some of the original data. It is known that when a matrix is low-rank, some of its columns (rows) can be represented by linear combinations of other columns (rows). Based on this conclusion, experts have proposed CUR decomposition, that is, representing the low-rank approximation of the matrix with a small number of actual columns and rows of the original data matrix. Figure 1 is a schematic diagram of CUR decomposition for a matrix X of size n×n. CUR decomposes X into the form of the product of three matrices, that is, X≈CUR; where, C ∈ R n×c is a column submatrix composed of c columns of X, R ∈ R r×n is a row submatrix composed of r rows of X, and U ∈ R c×r is a construction matrix that makes the product approximate to X, and c, r << m. After CUR decomposition, the three factor matrices C, U, and R will be stored to replace the matrix X. When the sum of the data volumes of C, U, and R is much smaller than the data volume of X, that is, n(c + r) + cr << n 2 , the compression of X is achieved. Compared with truncated SVD, CUR decomposition constructs factor matrices using some rows and columns of the data matrix, retains the characteristics of the original data, and makes the decomposition result interpretable.

[0007] Currently, the research on matrix CUR decomposition algorithms is mainly divided into two categories, one is the random algorithm and the other is the deterministic algorithm. The following briefly introduces the related work and respective limitations of these two types of algorithms.

[0008] The random algorithm assigns sampling probabilities to each column and each row of the matrix. Through random sampling, the decomposition result can be obtained relatively quickly. Frieze et al. [4] assign sampling probabilities to each column (and row) of the matrix according to the 2-norms of the columns (and rows) of matrix X. This method does not require the entire matrix to be stored in memory, solving the problem that when the matrix scale is too large to be fully input into memory for calculation. However, since using the 2-norms of the matrix columns (and rows) to measure the column (and row) space information is not accurate enough, the error between the columns and rows selected by this method and the original matrix is relatively large. Subsequently, Mahoney et al. [3] proposed to design sampling probabilities using the leverage scores of columns (and rows), which can better measure the column (and row) space information and effectively reduce the approximation error of the random algorithm. However, due to the uncertainty of the random algorithm, its approximation result is not stable. Especially when the matrix scale is small, due to the large variance of the approximation result, a poor approximation result may be obtained [5] .

[0009] The deterministic algorithm formulates the CUR decomposition problem as a convex optimization problem with a regularization term. It takes the parameter vectors corresponding to the columns and rows of the data matrix as the optimization objectives, so it can uniquely determine the columns and rows used to construct the factor matrix [6] . The deterministic algorithm usually uses the coordinate descent method to solve [7] , iteratively updating each parameter vector corresponding to each column and each row of the data matrix until convergence. However, the computational cost of updating the parameter vectors is the cube of the number of columns or rows of the data matrix, so it becomes the efficiency bottleneck of the entire algorithm. For large-scale matrices, this method is difficult to apply due to the excessive computational cost. To accelerate the speed of the deterministic algorithm, Ida et al. [5] defined the upper and lower bounds of the optimal condition score with the parameter vectors being zero. First, they concentrated on updating the non-zero parameter vectors identified by the lower bound, and then skipped the update of the zero parameter vectors identified by the upper bound during the iterative update process, thus reducing the total computational cost of the deterministic CUR decomposition. However, the method of Ida et al. still has some deficiencies: (1) The upper and lower bounds of the optimal condition are relatively far from the exact values, resulting in a limited number of identified zero parameter vectors and non-zero parameter vectors; (2) The computational cost of the lower bound of the optimal condition score is relatively large, resulting in a large cost of identifying zero parameter vectors. In addition, since the real network monitoring data usually contains some noise, many parameter vectors corresponding to the columns (rows) of the data matrix cannot converge to zero vectors, causing the deterministic CUR algorithm to select too many columns and rows, so the compression effect is poor.

[0010] In summary, the existing research has the following deficiencies:

[0011] (1) The decomposition result of truncated SVD lacks physical meaning. Although truncated SVD can compress the monitoring matrix, the singular vectors obtained by truncating the monitoring matrix with SVD do not provide physical meanings corresponding to the row and column attributes of the monitoring matrix, lacking interpretability. Therefore, the decomposition result is difficult to guide network control and management work.

[0012] (2) CUR decomposition based on random algorithms is not stable enough. Although CUR decomposition based on random algorithms makes the decomposition result interpretable, due to the random sampling of rows and columns by random algorithms, which has uncertainty, its approximate result is not stable. Especially when the matrix scale is small, due to the large variance of the approximate result, a poor approximate result may be obtained.

[0013] (3) CUR decomposition based on deterministic algorithms is slow. Although the results of deterministic CUR algorithms can satisfy both interpretability and stability at the same time, since deterministic algorithms usually use coordinate descent method to solve, the computational cost of updating the parameter vector by coordinate descent method is high, and each parameter vector corresponding to each column and each row of the data matrix needs to be iteratively updated until convergence. Therefore, the speed of such algorithms is very slow.

[0014] (4) CUR decomposition based on deterministic algorithms has a poor compression effect on network monitoring data affected by noise. Since real network monitoring data usually contains some noise, many parameter vectors corresponding to the columns (rows) of the data matrix cannot converge to the zero vector, resulting in too many columns and rows selected by the deterministic CUR algorithm. Therefore, the compression effect is poor. Summary of the Invention

[0015] The technical problem to be solved by the present invention is to provide a network monitoring data compression method, a terminal device and a storage medium in view of the deficiencies of the prior art, making full use of the correlation of end-to-end performance data of different node pairs in the network, realizing the compression of end-to-end monitoring data of all network nodes in the network, and improving the network data compression efficiency.

[0016] To solve the above technical problem, the technical solution adopted by the present invention is: A network monitoring data compression method, including the following steps:

[0017] S1. Construct a monitoring matrix P using the collected network monitoring data; let X = P or X = P T ; generate a sequence of regularization constants N + 1 is the length of the sequence of regularization constants, λ q > 0;

[0018] S2. Construct a column selection model of X: where W is a parameter matrix, W iThe parameter vector of the i-th row of W, where i = 1, …, n and n is the number of nodes in the network;

[0019] S3. Establish a set I for recording the indexes of non-zero parameter vectors. Update the parameter vectors according to the upper bounds of the optimal condition scores of each parameter vector corresponding to the indexes recorded in the set I, where the set I is initially an empty set;

[0020] S4. Increment the value of q by 1, and return to step S3 until q = N. Screen the columns indicated by the indexes in the set I to obtain the matrix C or the matrix R, where the matrix C corresponds to the output matrix of X = P, and the matrix R corresponds to the output matrix of X = P T ;

[0021] S5. Let the matrix U = C + PR + , to obtain the compressed network monitoring data C, U, and R, where C + , R + are the pseudo-inverses of the matrices C and R, respectively.

[0022] The network monitoring data compression method proposed by the present invention not only makes the compression result stable and interpretable, but also has a faster compression speed and a lower compression rate. The present invention constructs a column selection model based on the deterministic CUR algorithm, and uses part of the original monitoring data as the compression result, making the compression result interpretable and stable. The present invention improves the compression speed by avoiding unnecessary updates of zero parameter vectors. The present invention further screens the selected rows and columns according to the accuracy requirements, further reducing the compression rate. The present invention makes full use of the correlation of the end-to-end performance monitoring data of different node pairs in the network to achieve the compression of the whole network performance monitoring data.

[0023] In step S1, the implementation process of constructing the monitoring matrix P using the collected network monitoring data includes:

[0024] 1) Construct an n-dimensional vector from the end-to-end monitoring data of all network paths with i as the source node where represents the monitoring data obtained by sending a probe packet with node i as the source node and node j as the destination node;

[0025] 2) Model the end-to-end performance monitoring data of n nodes as an n×n monitoring matrix P = (P1, P2, …, P n ) T ; where the i-th row of P represents the end-to-end performance monitoring data from the source node i to all network nodes, and the j-th column of P represents the end-to-end performance monitoring data at the destination node j from all network nodes.

[0026] ​In step S1, the regularization constant sequence is generated by the formula λ q = λ max 10 -γqN , where γ = 4, N = 99, and λ max = max i ||X iT X||2, X iT is the transpose of the i-th column of matrix X.

[0027] In the present invention, λ0 < λ1 < … < λ N .

[0028] In step S3, the update formula for the parameter vector W i is as follows:

[0029]

[0030] where

[0031] In step S3, the process of establishing the set I includes: in ascending order of the indices, using the update formula of the parameter vector W i to calculate the updated parameter vector. If the parameter vector W i ≠ 0, then the index i is included in the set I.

[0032] In step S3, the specific implementation process of updating the parameter vector according to the upper bound of the optimal condition score of each parameter vector corresponding to the indices recorded in the set I includes:

[0033] A. In ascending order of the index values, calculate the upper bound of the optimal condition score of the non-zero parameter vectors indicated by the indices recorded in the set I Update the parameter vector according to . For each updated parameter vector, perform a synchronous update on δ; where the update method of the parameter vector is: calculate If calculate the updated parameter vector according to the update formula of the parameter vector W i ; otherwise set W i = 0; where is ‖z i ‖2 used in the previous update of W i , and ‖G i ‖2 represents the 2-norm of G i . G i represents the i-th row of matrix G, G = X T X; the initial value of δ is zero, and its update formula is where ΔW i ' represents the difference between the updated parameter vector and the difference, represents the matrix of the i-th row, the initial value of the matrix is a zero matrix of size n×n;

[0034] B. Delete the indices corresponding to the zero vectors indicated by the set I;

[0035] C. If δ is less than the threshold, all parameter vectors converge; otherwise, let and repeat steps A to C.

[0036] After the above operations, a parameter matrix W with all-zero partial behaviors is obtained. The product XW in the column selection model is equivalent to making a linear combination of the columns of X with the column values of W as coefficients to approximate the columns of X; if W i = 0, then X i contributes zero to the linear combination.

[0037] In step S4, the specific implementation process of screening the columns indicated by the indices in the set I to obtain the matrix C or the matrix R includes:

[0038] a) If is not the maximum value in the i-th column of the parameter matrix W, then delete the index i from the set I; represents the element in the i-th row and i-th column of the parameter matrix W;

[0039] b) Establish a 2-row m-column array M. The first row of M stores the indices recorded in the set I, and the second row of M stores the 2-norms of the rows of W corresponding to the indices stored in the first row; m represents the number of indices in the set I;

[0040] c) Sort M column-wise in descending order according to the values in the second row of M, and delete the indices with lower rankings from the set I

[0041] d) Select the columns of the matrix X indicated by the updated set I after step c) to construct the matrix C or the matrix R.

[0042] The present invention also provides a network monitoring data compression system, including a memory, a processor, and a computer program stored on the memory; the processor executes the computer program to implement the steps of the method of the present invention.

[0043] The present invention also provides a computer program product, including computer programs / instructions; when the computer programs / instructions are executed by a processor, the steps of the method of the present invention are implemented.

[0044] The present invention also provides a computer-readable storage medium, on which a computer program / instructions are stored; when the computer program / instructions are executed by a processor, the steps of the above method of the present invention are implemented.

[0045] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0046] 1. The present invention uses some rows and columns of the monitoring matrix as the compression result, which not only achieves the purpose of compressing data and saving storage resources, but also enables the compression result to guide network management. The present invention models the detection data of a network containing n nodes into a monitoring matrix, uses CUR decomposition to compress the monitoring matrix, and selects some columns and rows of the monitoring matrix as the compression result, making the decomposition result interpretable. These information is of great guiding significance for advanced network management services such as path setting, network resource allocation, anomaly detection, and fault recovery.

[0047] 2. The present invention improves the compression speed of the deterministic CUR algorithm for the monitoring matrix, which not only makes the compression result stable, but also reduces the computational cost of network monitoring. The deterministic CUR algorithm uses the coordinate descent method to solve. The coordinate descent method iteratively updates all parameter vectors until convergence, and the update cost each time is O(n 3 ), so the compression speed is very slow. The present invention believes that not updating the zero parameter vectors will not affect the approximation accuracy; therefore, the present invention establishes a set to record the indices of the non-zero parameter vectors, and iteratively updates the non-zero parameter vectors corresponding to the indices in the set, thereby reducing the number of updates of the zero parameter vectors. During the iterative update of the non-zero parameter vectors, some non-zero parameter vectors will become zero vectors; therefore, the present invention calculates the upper bound of the optimal condition score, and identifies and skips the update of the zero parameter vectors before updating the parameter vectors. Therefore, the present invention reduces the number of updates of the zero parameter vectors, improves the compression speed, and reduces the computational cost of network monitoring.

[0048] 3. While ensuring the approximation accuracy, the present invention further screens the columns and rows selected by the deterministic CUR algorithm, reduces the compression ratio of the monitoring data, and reduces the storage resources required for network monitoring. Since the real network monitoring data usually contains some noise, the number of columns and rows selected by the deterministic CUR algorithm is too large, so the compression effect is poor. The present invention uses the values in the parameter matrix as coefficients, compares the contribution sizes of the linear combinations of the columns or rows in the monitoring matrix, and screens the columns and rows selected by the deterministic CUR algorithm twice. In the first screening, the column with the largest contribution is selected when it is approximately itself by linear combination with other columns; in the second screening, the column with a relatively large overall contribution is selected when it participates in the linear combination to approximate itself and other columns. Description of the Drawings

[0049] Figure 1Schematic diagram of matrix CUR decomposition;

[0050] Figure 2 Schematic diagram of the multiplication of matrix X and parameter matrix W;

[0051] Figure 3 Indicates the processing time of three data compression methods on the first 6 groups of network monitoring data of PlanetLab;

[0052] Figure 4 Indicates the relative error corresponding to the number of columns (or rows) selected in the order of contribution size on PL01;

[0053] Figure 5 Indicates the relative error corresponding to the number of columns (or rows) selected in the order of contribution size on PL02;

[0054] Figure 6 Indicates the relative error corresponding to the number of columns (or rows) selected in the order of contribution size on PL03;

[0055] Figure 7 Indicates the relative error corresponding to the number of columns (or rows) selected in the order of contribution size on PL04;

[0056] Figure 8 Indicates the relative error corresponding to the number of columns (or rows) selected in the order of contribution size on PL05;

[0057] Figure 9 Indicates the relative error corresponding to the number of columns (or rows) selected in the order of contribution size on PL06. Detailed implementation manners

[0058] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0059] In this article, the terms "comprise", "include" and other similar words are intended to represent logical relationships and should not be regarded as representing spatial structural relationships. For example, "A includes B" is intended to mean that logically B belongs to A, rather than indicating that B is located inside A in terms of space. Additionally, the meanings of the terms "comprise", "include" and other similar words should be regarded as open-ended rather than closed. For example, "A includes B" is intended to mean that B belongs to A, but B does not necessarily constitute all of A, and A may also include other elements such as C, D, E, etc.

[0060] Example 1

[0061] This embodiment provides a network monitoring data compression method, including the following steps:

[0062] S1. Model the collected network monitoring data (such as end-to-end round-trip delay, available bandwidth, and packet loss rate) to form a monitoring matrix P and assign it to the matrix X.

[0063] S2. Generate regularization constant sequence according to the sequence rule Precalculate constants and initialize variables; set the initial value of q to 0;

[0064] S3. Based on the deterministic CUR algorithm, build a column selection model for X;

[0065] S4, for the regularization constant λ q , according to the column selection model, determine all parameter vectors W i (i=0,1,...,n) update method;

[0066] S5. Update parameter vector W i , establish a set I that records the indexes of non-zero parameter vectors; iteratively update the non-zero parameter vectors recorded in set I, and predict their changing trends based on the upper bound of the optimal condition score of the parameter vector, avoiding updating the zero parameter vector until convergence; set q = q + 1, and repeat steps S3, S4 and S5 until q = N;

[0067] S6. Further filter the columns indicated by the non-zero parameter vectors obtained by the above process to obtain a column submatrix Y of X;

[0068] S7, let C = Y, so far the column selection process is completed; for the row selection of the monitoring matrix P, let X = P T , repeat steps S2, S3, S4, S5 and S6, let matrix R = Y T , so far the row selection process of P has been completed;

[0069] S8. Let matrix U = C + PR + , that is, the compressed network monitoring data (C, U, R) is obtained; among them, C + , R + are the pseudo-inverses of matrices C and R respectively.

[0070] In the above implementation steps, the column selection process and the row selection process can be performed in different orders or simultaneously. The embodiment of the present invention does not limit the execution order.

[0071] The network monitoring data compression method proposed in this embodiment not only makes the compression result stable and interpretable, but also has a faster compression speed and a lower compression rate: (1) In this embodiment, a column selection model is constructed based on the deterministic CUR algorithm, and part of the original monitoring data is used as the compression result, making the compression result interpretable and stable. (2) In this embodiment, the compression speed is improved by avoiding unnecessary updates of zero parameter vectors. (3) In this embodiment, the compression rate is further reduced by further screening the selected rows and columns. This embodiment makes full use of the correlation of end-to-end performance monitoring data of different node pairs in the network to achieve the compression of the whole network performance monitoring data.

[0072] The specific implementation process of step S1 includes:

[0073] 1) Construct all the end-to-end monitoring data of the network paths with node i as the source node into an n-dimensional vector where, represents the monitoring data obtained by sending a probe packet with node i as the source node and node j as the destination node;

[0074] 2) Model the end-to-end performance monitoring data of n nodes into an n×n monitoring matrix P = (P1, P2, …, P n ) T ; where, the i-th row of P represents the end-to-end performance monitoring data from the source node i to all network nodes; the j-th column of P represents the end-to-end performance monitoring data at the destination node j from all network nodes.

[0075] 3) Assign the monitoring matrix P to the matrix X, that is, X = P.

[0076] In step S2, the regularization constant sequence is generated by the formula λ q = λ max 10 -γqN , where γ = 4, N = 99, and λ max = max i ||X iT X||2, X iT is the transpose of the i-th column of matrix X.

[0077] In step S2, the specific implementation process of the pre-computation is: Let the matrix G = X T X, calculate and save the values of ‖G i ‖2 (i = 1, …, n); where, ‖G i ‖2 represents the 2-norm of G i , and G i represents the i-th row of matrix G.

[0078] Since the value of ‖Gi For all \(i\in\{1,\ldots,n\}\), to improve the computational efficiency, pre - calculations are performed on them.

[0079] In this embodiment, the specific implementation process of initialization is as follows:

[0080] 1) Create an \(n\) - dimensional array Scores, initialized as an all - zero array;

[0081] 2) Create a parameter matrix \(W\) of size \(n\times n\), initialized as an all - zero matrix;

[0082] 3) Create a matrix of size \(n\times n\) Initialized as an all - zero matrix;

[0083] 4) Create a variable \(\delta\), initialized as zero.

[0084] The column - selection model in step S3 is:

[0085]

[0086] where \(W\) is the parameter matrix, and its \(i\) - th row \(W\) i is the parameter vector; \(\lambda\) q \(>0\) is the regularization constant, \(q = 0,\ldots,N\).

[0087] During the training process of \(W\), the regularization term \(\lambda\) q \(\|W\|\) i \((i = 1,\ldots,n)\) will induce \(W\) i to become a zero vector, prompting \(W\) to become a sparse matrix with all - zero partial behaviors, where \(\lambda\) q is used to control its sparsity. The product \(XW\) in the column - selection model is equivalent to making a linear combination of the columns of \(X\) with the column values of \(W\) as coefficients to approximate the columns of \(X\); if \(W\) i \(=0\), then the contribution of \(X\) i to the linear combination is zero, as Figure 2 shown. If the set \(I\) is used to record the indices of the non - zero parameter vectors in \(W\), then the columns of \(X\) indicated by the indices in \(I\) are the columns to be selected by the column - selection model.

[0088] In step S4, the update method of the parameter vector \(W\) i obtained by the coordinate - descent method is:

[0089]

[0090] where,

[0091]

[0092]

[0093] In step S5, the process of establishing set I is as follows:

[0094] 1) Initialize the index set I as an empty set;

[0095] 2) Update each parameter vector W in ascending order of the index value i (i = 0, 1,..., n). If

[0096] W i ≠ 0, then store K i into Scores[i], and include the index i in the index set I.

[0097] In step S5, the specific implementation process of predicting the change trend according to the upper bound of the optimal condition score and avoiding the update of the zero parameter vector includes:

[0098] A. For the non-zero parameter vectors indicated by set I in ascending order of the index value, calculate the upper bound of its optimal condition score K i :

[0099]

[0100] where is ‖z i ‖2 used in the previous update of W i ; update the parameter vector according to (if

[0101] indicates that W i will become a zero vector, let W i = 0; otherwise, update W i ), and synchronously update δ:

[0102]

[0103] where, ΔW i ′ represents the difference between the updated W i and ; the initial value of the matrix is an all-zero matrix of size n×n, and the initial value of δ is zero;

[0104] B. If, after step A, the non-zero parameter vectors indicated by set I become zero vectors, then delete their corresponding indices from set I;

[0105] C. If δ is less than the threshold, it means that all the parameter vectors in W have converged, and stop the iterative update; otherwise, let and repeat steps A to C.

[0106] In step S5, non-iterative update of zero parameter vectors outside set I will not affect the approximation accuracy of the column selection model for the data matrix. The theoretical basis is that after each update of each parameter vector W i (i = 1, 2, …, n), the zero parameter vectors therein remain zero after the next round of update.

[0107] In step S5, predicting its change trend based on the upper bound of the optimal condition score of the parameter vector and avoiding the update of zero parameter vectors is based on the following theory:

[0108] A. As can be seen from formula (1), when K i ≤λ q , W i = 0; As the upper bound of K i , if then K i ≤λ q must hold, and it can be predicted that W i = 0. Therefore, it can be predicted whether the parameter vector is a zero vector;

[0109] B. Given that K i = ‖z i ‖2, where the computational cost of z i is O(n 3 ); if the computational cost of 3 can be reduced to O(n), which is much lower than O(n ), then using

[0110] for prediction has a lower cost. In this embodiment, formula (4) is used to calculate

[0111] [[ID=�3]]1) Formula (3) can be split into the following form:

[0112]

[0113] Using the triangle inequality and the Cauchy-Schwarz inequality, the following inequality is obtained:

[0114]

[0115] Given that K i = ‖z i ‖2, an upper bound of K i can be obtained from the above formula:

[0116]

[0117] 2) As can be seen from formula (6), The computational cost is mainly generated by three parts, namely ‖ΔW‖ F , ‖ΔW i ‖² and their computational costs are O(n 2 ), O(n), O(n 3 );

[0118] 3) In this embodiment, δ is used to store the value of ‖ΔW‖ F , and the synchronous update method is adopted to reduce the computational cost of ‖ΔW‖ F to O(n); in this embodiment, is used to store the result of the previous round of update. When W i has not been updated yet, therefore the value of ‖ΔW i ‖² is always zero, and its computational cost is O(0); in this embodiment, the K i obtained in the previous round is used to replace to reduce the computational cost to O(1); the calculation formula (4) of is obtained, and its computational cost is O(n).

[0119] In step S5, after each round of update, check whether the non-zero parameter vector indicated in set I becomes a zero vector, delete the index corresponding to the zero parameter vector in time, and gradually narrow the update range of the parameter vector, effectively reducing the computational cost.

[0120] In step S5, changing the regularization constant according to the sequence rule [8] can accelerate the convergence of W. The theoretical basis is that the regularization constant sequence is arranged in descending order, that is, λ0 > λ1 > … > λ N , and the larger values among them can make W converge quickly; when W has reached convergence based on a larger regularization constant, the convergence speed is much faster when training based on a smaller regularization constant than directly using a small regularization constant; therefore, the sequence rule can generally improve the convergence speed of W.

[0121] Since the real network monitoring data usually contains some noise, many parameter vectors corresponding to the columns (rows) of the data matrix cannot converge to zero vectors, resulting in too many columns selected by the column selection model and a poor compression effect. Therefore, step S6 further screens the columns indicated by the non-zero parameter vectors obtained from the above process.

[0122] The specific implementation process of step S6 is as follows:

[0123] 1) If is not the maximum value in the i-th column of W, then delete the index i from set I;

[0124] 2) Calculate the number of indices in set I, denoted by m;

[0125] 3) Create an array M with 2 rows and m columns. The first row of M stores the index of the record in set I. The second row of M stores the 2-norm of the row of W corresponding to the index stored in the first row.

[0126] 4) Sort M by column according to the values in the second row of M from largest to smallest, and delete the indexes other than the first 100 from set I;

[0127] 5) Select the columns of matrix X indicated by set I and construct matrix Y.

[0128] Step S6 performs two screenings on the columns indicated by the non-zero parameter vectors. The first screening selects the columns that contribute the most when linearly combining with other columns to approximate themselves; the second screening selects the columns that contribute the most when participating in the linear combination to approximate themselves and other columns.

[0129] like Figure 2 As shown in the example, if Mainly by X 1 Expressed as, then the coefficient It should be The maximum value in ; if Mainly by X 2 Expressed as, then the coefficient It should be The maximum value in ; if Mainly by X 3 Expressed as, then the coefficient It should be From the above conclusion, it can be inferred that if the coefficient W i i no The maximum value in Mainly by the addition of X i The other columns besides X i can be approximated by a linear combination of other columns. Therefore, if If it is not the maximum value in the i-th column of W, its corresponding index i is deleted from the set I during the first screening.

[0130] like Figure 2 As shown in the example, Each column of is obtained by linear combination of all columns of X. i right The contribution of each column is determined by the parameter vector Control. Use W i The 2-norm of X i right The contribution size of each column is calculated, and the indexes in the set I are sorted according to the contribution size, with the indexes with the highest sorting priority being selected.

[0131] To verify the network monitoring data compression method proposed in this embodiment, the proposed network monitoring data compression method (abbreviated as Ours) in this embodiment is compared with two other data compression methods, namely the deterministic CUR decomposition method proposed by Bien et al. [6] (abbreviated as Origin) and the fast deterministic CUR decomposition method proposed by Ida et al. [5] (abbreviated as Ida). The metrics used include: the processing time of the data compression method, the number of updates of the parameter vector during the data compression process, the number of rows, columns, and compression ratio selected by the data compression method, and the relative error generated by the compressed result approximating the monitoring matrix.

[0132] This embodiment conducts experiments on the first 6 groups of monitoring data in the PlanetLab dataset. PlanetLab is a set of monitoring data containing the end-to-end round-trip delay of 490 nodes in the network. It includes 18 groups of network delay data measured in 18 time periods. The first 6 groups of monitoring data used in the experiment are represented by PL01, PL02, PL03, PL04, PL05, and PL06 respectively.

[0133] Table 1 shows the number of columns and rows indicated by the non-zero parameter vectors of the three data compression methods on the first 6 groups of monitoring data in PlanetLab, as well as their compression ratios and relative errors. The formula for the compression ratio is: (n*c + n*r + c*r) / (n×n), where the numerator represents the sum of the data volumes of the compressed matrices C, U, and R, and the denominator represents the data volume of the monitoring data; c represents the number of columns of matrix C, r represents the number of rows of matrix R, and n represents the number of network nodes. The formula for the relative error is: ‖P - CUR‖ F / ‖P‖ F . It can be seen from the table that the number of columns, rows, compression ratio, and relative error indicated by the non-zero parameter vectors obtained by the three methods are the same. Among them, the number of rows obtained by each method is equal to the number of columns because PlanetLab is symmetric round-trip delay data.

[0134] Figure 3 Shows the processing times of the three data compression methods on the first 6 groups of network monitoring data in PlanetLab. It can be seen from the figure that the method proposed in this embodiment has the least processing time. This shows that the method proposed in this embodiment has a faster compression speed while ensuring the compression ratio and relative error.

[0135] Since real network monitoring data usually contains some noise, many parameter vectors corresponding to the columns (rows) of the data matrix cannot converge to the zero vector, resulting in too many columns and rows selected by the deterministic CUR algorithm and poor compression effect. Therefore, this embodiment proposes a method for screening the columns and rows indicated by non-zero parameter vectors to reduce the compression ratio of the monitoring matrix. The screening method proposed in this embodiment can also be used in other data compression methods based on the deterministic CUR algorithm.

[0136] In step S6 of the compression method proposed in this embodiment, the columns and rows indicated by non-zero parameter vectors are sorted and screened according to their contribution to the approximation result, and the columns and rows with higher rankings are preferentially selected. Figures 4 to 9 Shows the relative errors generated by the number of columns (or rows) selected in the order of contribution on the monitoring data PL01, PL02, PL03, PL04, PL05, and PL06. It can be seen from the figure that when the number of selected columns (or rows) reaches about 100, the relative error basically converges to zero, and the compression ratio at this time is 0.45. Usually, network administrators can set the number of columns and rows selected by the compression method according to the requirements for the compression ratio and relative error in the actual situation.

[0137] Table 1 The number of columns and rows indicated by non-zero parameter vectors, their compression ratios, and relative errors obtained by the network monitoring data compression method proposed in this embodiment and two other data compression methods on the first 6 groups of monitoring data of PlanetLab

[0138]

[0139] Embodiment 2

[0140] Embodiment 2 of the present invention provides a terminal device corresponding to Embodiment 1 above. The terminal device can be a processing device for a client, such as a mobile phone, a laptop computer, a tablet computer, a desktop computer, etc., to execute the method of the above embodiment.

[0141] The terminal device of this embodiment includes a memory, a processor, and a computer program stored on the memory; the processor executes the computer program on the memory to implement the steps of the method of Embodiment 1 above.

[0142] In some implementations, the memory can be a high-speed random access memory (RAM: Random Access Memory), and may also include a non-volatile memory, such as at least one disk memory.

[0143] In other implementations, the processor can be a central processing unit (CPU), a digital signal processor (DSP), or various other types of general-purpose processors, which are not limited here.

[0144] Example 3

[0145] Example 3 of the present invention provides a computer-readable storage medium corresponding to Example 1 above, on which computer programs / instructions are stored. When the computer programs / instructions are executed by a processor, the steps of the method in Example 1 above are implemented.

[0146] A computer-readable storage medium can be a tangible device that holds and stores instructions for use by an instruction execution device. A computer-readable storage medium can be, for example, but is not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any combination of the above.

[0147] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code. The solutions in the embodiments of the present application can be implemented in various computer languages. For example, object-oriented programming languages such as Java and interpreted scripting languages such as JavaScript.

[0148] The present application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, as well as the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one or more of the flows Figure 1 or blocks or multiple blocks.

[0149] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are performed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one or more of the flows Figure 1 or blocks or multiple blocks.

[0150] Although the preferred embodiments of the present application have been described, those skilled in the art can make additional changes and modifications once they learn the basic creative concepts. Therefore, the appended claims are intended to be construed to include the preferred embodiments as well as all changes and modifications falling within the scope of the present application.

[0151] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalent technologies, the present application also intends to include these modifications and variations.

[0152] References:

[0153] [1]Guo C, Wu H, Tan K, et al. Dcell: a scalable and fault-tolerant network structure for data centers[C]. Conference on Data Communication (SIGCOMM). ACM, 2008: 75 - 86.

[0154] [2]Xie K, Wang L, Wang X, et al. Sequential and adaptive sampling for matrix completion in network monitoring systems[C]. Conference on Computer Communications (INFOCOM). IEEE, 2015: 2443 - 2451.

[0155] [3]Mahoney M W, Drineas P. CUR matrix decompositions for improved data analysis[J]. Proceedings of the National Academy of Sciences, 2009, 106(3): 697 - 702.

[0156] [4]Frieze A, Kannan R, Vempala S. Fast Monte-Carlo algorithms for finding low-rank approximations[J]. Journal of the ACM (JACM), 2004, 51(6): 1025 - 1041.

[0157] [5] Ida Y, Kanai S, Fujiwara Y, et al. Fast deterministic CUR matrix decomposition with accuracy assurance[C]. International Conference on Machine Learning. PMLR, 2020:4594-4603.

[0158] [6] Bien J, Xu Y, Mahoney M W. CUR from a sparse optimization viewpoint[J]. Advances in Neural Information Processing Systems, 2010, 23.

[0159] [7] Fu W J. Penalized regressions: the bridge versus the lasso[J]. Journal of Computational and Graphical Statistics, 1998, 7(3):397-416.

[0160] [8] Tibshirani R, Bien J, Friedman J, et al. Strong rules for discarding predictors in lasso-type problems[J]. Journal of the Royal Statistical Society: Series B(Statistical Methodology), 2012, 74(2):245-266.

Claims

1. A method for compressing network monitoring data, characterized in that, It includes the following steps: S1. Construct a monitoring matrix P by using the collected network monitoring data; Let X = P or X = P T ; Generate a sequence of regularization constants N + 1 is the length of the sequence of regularization constants, λ q > 0; S2. Construct the column selection model of X: where W is a parameter matrix, and W i is the i-th row parameter vector of W, i = 1, …, n, and n is the number of nodes in the network; S3. Establish a set I for recording the indexes of non-zero parameter vectors, and update the parameter vectors according to the upper bounds of the optimal condition scores of the parameter vectors corresponding to the indexes recorded in the set I; where the set I is an empty set initially; S4. Increment the value of q by 1, and return to step S3 until q = N; Filter the columns indicated by the indices in set I to obtain matrix C or matrix R; where matrix C corresponds to the output matrix for X = P, and matrix R corresponds to the output matrix for X = P T ; S5. Let matrix U = C + PR + , to obtain the compressed network monitoring data C, U, R; where C + , R + are the pseudo-inverses of matrices C and R respectively. In step S1, the implementation process of constructing the monitoring matrix P by using the collected network monitoring data includes: 1) Construct the end-to-end monitoring data of all network paths with i as the source node into an n-dimensional vector where P i j represents the monitoring data obtained by sending probe packets with node i as the source node and node j as the destination node; 2) Model the end-to-end performance monitoring data of n nodes as an n×n monitoring matrix P = (P1, P2, …, P n ) T ; where the i-th row of P represents the end-to-end performance monitoring data with node i as the source node, from the source node to all network nodes; the j-th column of P represents the end-to-end performance monitoring data with node j as the destination node, at the destination node from all network nodes; In step S1, a sequence of regularization constants is generated The regularization constant λ q has the following expression: where γ and N are constants, and λ max = max i ||X iT X||₂, where X iT is the transpose of the i-th column of matrix X; In step S3, the parameter vector W i is updated according to the following formula: Among them, 2. The network monitoring data compression method according to claim 1, wherein In step S3, the process of establishing the set I that records the indices of non-zero parameter vectors includes: in ascending order of the indices, using the update formula of the parameter vector W i to calculate the updated parameter vectors in sequence. If the parameter vector W i ≠ 0, then add the index i to the set I.

3. The network monitoring data compression method according to claim 1 or 2, characterized in that, In step S3, according to the upper bound of the optimal condition score of each parameter vector corresponding to the index recorded in set I The process of updating the parameter vector includes: A. Calculate the upper bound of the optimal condition score of the non-zero parameter vectors indicated by the indices of the records in set I in ascending order of the index values. According to Update the parameter vectors. For each updated parameter vector, perform a synchronous update on δ. Among them, the update method of the parameter vectors is: calculate If Calculate the updated parameter vector according to the update formula of the parameter vector W i ; otherwise, set W i = 0; Where is the ‖z i ‖2 used in the previous update of W i , and ‖G i ‖2 represents the 2-norm of G i . G i represents the i-th row of the matrix G, G = X T X. The initial value of δ is zero, and its update formula is Among them ΔW i ′ represents the difference between the updated parameter vector and , represents the i-th row of the matrix , and the initial value of the matrix is a zero matrix of size n×n; B. Delete the indexes corresponding to the zero vectors indicated by the set I; C. If δ is less than the threshold, then all parameter vectors converge and end; otherwise, let And repeat steps A to C.

4. The network monitoring data compression method according to claim 1, wherein λ0 > λ1 > … > λ N 。 5. The network monitoring data compression method according to claim 1, characterized in that In step S4, the specific implementation process of screening the columns indicated by the indexes in the set I to obtain the matrix C or the matrix R includes: a) If is not the maximum value in the i-th column of the parameter matrix W, then the index i is deleted from the set I; represents the element in the i-th row and i-th column of the parameter matrix W; b) Establish a two-row and m-column array M, store the indexes recorded in the set I in the first row of M, and store the 2-norms of the rows of W corresponding to the indexes stored in the first row in the second row of M; m represents the number of indexes in the set I; c) Sort M column by column in descending order of the values in the second row of M, and delete the indexes with lower rankings from the set I; d) Select the columns of the matrix X indicated by the set I updated in step c) to construct the matrix C or the matrix R.

6. An electronic device, characterized in that, It includes: One or more processors; A memory storing one or more programs, which when executed by the one or more processors, cause the one or more processors to implement the steps of the method according to any one of claims 1 to 5.

7. A computer-readable storage medium, characterized in that, It stores a computer program, which when executed by a processor, implements the steps of the method according to any one of claims 1 to 5.