APP traffic anomaly detection method and system based on multi-scale time series
Through the multi-scale time series APP traffic anomaly detection method, combined with wavelet transformation, empirical modal decomposition and deep learning technology, the problem that the existing technology cannot effectively detect APP traffic anomaly is solved, and comprehensive detection and accurate identification of short-term and long-term abnormal patterns are achieved, which improves detection accuracy and model generalization capabilities.
Patent Information
- Application Number
- CN202510188973.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-20
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-02-20
AI Technical Summary
The existing technology cannot effectively capture the short-term and long-term abnormal patterns of APP traffic, and cannot distinguish between bursty, persistent and trendy abnormalities, has low detection accuracy, and lacks special detection methods for APP traffic.
The traffic anomaly detection method based on multi-scale time series is adopted. Through traffic data acquisition, preprocessing, multi-scale time series decomposition and abnormal detection modules, combined with wavelet transformation, empirical modal decomposition, isolated forest algorithm and long and short-term memory network, the abnormal characteristics of high-frequency and low-frequency components are captured and the traffic risk level and abnormal type are judged.
A comprehensive detection of medium-, short-term and long-term abnormal patterns of APP traffic is realized, and the rapid, persistent and trend abnormalities are accurately distinguished, detection accuracy is improved, historical information is used to analyze current data, and special detection methods are designed based on the characteristics of APP traffic, which improves the generalization ability and interpretability of the model.
Smart Images

Figure CN120050083A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of mobile Internet traffic monitoring, and relates to a method and system for detecting APP traffic anomalies based on multi-scale time series. Background Art
[0002] With the rapid development of network technology, network systems have become increasingly complex, the types of network devices have increased day by day, and network traffic anomalies have occurred frequently. Network traffic anomalies refer to the deviation of network traffic behavior from the normal mode, which will bring great harm to the network and the computers on the network. Therefore, monitoring network traffic behavior and detecting anomalies in a timely manner are of great significance for improving the reliability and availability of the network.
[0003] Traditional network traffic anomaly detection methods mainly analyze the overall network traffic in the PC or server environment, and rarely consider the uniqueness of the traffic of applications (APPs) on mobile devices. However, APP traffic often exhibits characteristics different from network traffic, such as the frequency of scene-driven application programming interface (API) calls, request patterns at different time scales, and various types of resource loading. Traditional methods are difficult to fully capture these unique characteristics, resulting in poor detection effects.
[0004] APP traffic has unique behavior patterns, such as periodic access to the server, specific API call frequencies, etc., and these characteristics may be more prominent in multi-scale time series analysis. Compared with network traffic detection, APP traffic detection can more sensitively capture abnormal traffic at the application layer, such as sudden requests within a short period of time or the abuse of specific APIs, which are often potential signs of fraud or malicious behavior.
[0005] Regarding the problems that existing network traffic anomaly detection technologies cannot effectively capture short-term and long-term anomaly patterns and cannot distinguish between sudden, persistent, and trend anomalies, there are already some invention patents. Patent CN106411597A discloses a network traffic anomaly detection method and system. This method uses the time series composed of traffic data samples extracted as samples for model training and classification detection. Considering that the changes in network traffic have temporal continuity and correlation, time information is introduced into the detection and classification of abnormal traffic. However, this patent still has the problem of further optimizing the structure of the neural network model to improve the generalization ability and interpretability of the model. Patent CN116232761A discloses a network abnormal traffic detection method and system based on shapelet. This method uses shapelet time series data processing technology to obtain a shapelet sequence for network traffic anomaly detection, which can improve the detection rate and reduce the false alarm rate and missed alarm rate. However, this patent still has the problem of further optimizing the generation method of the shapelet sequence to improve the interpretability and reproducibility of the shapelet sequence.
[0006] The existing technologies have the following disadvantages: They cannot effectively capture short-term and long-term anomaly patterns in network traffic and cannot comprehensively detect abnormal situations; they cannot distinguish between sudden, persistent, and trend anomalies and cannot accurately identify the types of anomalies; the detection accuracy is relatively low and they cannot make full use of historical information to analyze current data; there is a lack of a dedicated detection method for APP traffic and they cannot detect anomalies according to the characteristics of APP traffic; there is room for optimization in the model structure and parameter selection, and the generalization ability and interpretability need to be improved. Summary of the Invention
[0007] In view of the problems in the existing technologies that it is impossible to comprehensively detect anomalies in APP traffic, impossible to capture short-term and long-term anomaly patterns simultaneously, and impossible to distinguish between sudden, persistent, and trend anomalies. Therefore, in order to solve the above problems, the present invention provides a method and system for detecting APP traffic anomalies based on multi-scale time series.
[0008] To achieve the above object, the present invention provides the following technical solutions:
[0009] A method for detecting APP traffic anomalies based on multi-scale time series, comprising the following steps:
[0010] Step 1, traffic data collection: Capture the network traffic of the target application (APP) and classify it according to API calls, resource loading, and advertisement request types;
[0011] Step 2, Traffic data preprocessing: Identify the APP running scenarios according to the request content and the Uniform Resource Locator (URL), group the traffic data by scenarios, and divide time windows based on scenario characteristics. Calculate the call frequency of the Application Programming Interface (API), the request size distribution, the response time, the request type distribution, and the request type switching frequency within each time window;
[0012] Step 3, Multi-scale time series decomposition: Combine the statistical features of each time window into a feature vector sequence, decompose it into high-frequency components and low-frequency components through wavelet transform, and perform empirical mode decomposition on the low-frequency components to obtain multiple intrinsic mode functions;
[0013] Step 4, Calculate the anomaly score for the high-frequency components using the Isolation Forest algorithm, calculate the anomaly score for the intrinsic mode functions of the low-frequency components using the Long Short-Term Memory network, and use the high- and low-frequency anomaly scores to judge the traffic risk level and anomaly type.
[0014] Further, the identification of the APP running scenarios in Step 2 is as follows: Match the preset scenario tags according to the keywords or domain names in the request URL, and the scenario tags are registration, login, payment, advertisement, or resource loading;
[0015] The rule for dividing the time windows in Step 2 is: Use a smaller time window for high-density scenarios and a longer time window for low-density scenarios. The high-density scenarios include resource loading or advertisement requests, and the low-density scenarios include payment or login;
[0016] In Step 3, use the db4 wavelet basis to perform discrete wavelet transform on the feature vector sequence, decomposing it into high-frequency components and low-frequency components;
[0017] The method of using the high- and low-frequency anomaly scores in Step 4 is: If both the high-frequency anomaly score and the low-frequency anomaly score exceed the preset threshold, it is determined as a persistent anomaly; if only the high-frequency score exceeds the threshold, it is determined as a sudden anomaly; if only the low-frequency score exceeds the threshold, it is determined as a trend anomaly.
[0018] An APP traffic anomaly detection system based on multi-scale time series, including:
[0019] A traffic data collection module that captures the APP network traffic for packet capture and classifies it according to the request type;
[0020] A traffic data preprocessing module that identifies the APP running scenarios and divides time windows, and extracts the traffic statistical features of each window;
[0021] The multi-scale time series decomposition module decomposes the feature vector sequence into high-frequency components and low-frequency components through wavelet transform and empirical mode decomposition;
[0022] The abnormal traffic detection module includes a high-frequency component abnormal detection module, a low-frequency component abnormal detection module, and an abnormal judgment module. The isolation forest algorithm and the long short-term memory (LSTM) model are used to calculate the abnormal score, and the abnormal judgment module comprehensively judges the output abnormal type and risk level of the module.
[0023] Furthermore, the traffic data preprocessing module further includes a scene recognition unit, and the scene recognition unit associates a preset scene label by matching URL keywords or domain names;
[0024] In the multi-scale time series decomposition module, the decomposition of the low-frequency component uses empirical mode decomposition to generate multiple intrinsic mode functions;
[0025] The isolation forest algorithm of the high-frequency component abnormal detection module determines the abnormal score by calculating the path length of the data point. The shorter the path length, the higher the abnormal score;
[0026] The LSTM model of the low-frequency component abnormal detection module calculates the abnormal score through the residual between the predicted value and the actual value. The larger the residual, the higher the abnormal score.
[0027] Furthermore, the traffic data preprocessing module divides the traffic data into different APP usage scenarios. Let the traffic of each scenario be C k , where k represents the scenario number, and the traffic data set D containing n different scenario identifications is expressed as the following formula:
[0028] D = {C 1 , C 2 , …, C n}
[0029] Among them, C 1 represents the traffic data of the registration scenario; C 2 represents the traffic data of the login scenario; C n represents the traffic data of other recognizable scenarios;
[0030] For each scenario C k , the time window is divided according to its characteristics to capture the traffic characteristics within the scenario. Let the time window length of scenario C k be Δt k , then scenario C k is divided into m time windows {W k,1 , W k,2 ,..., W k,m}:
[0031]
[0032] Among them, W k,j represents the j-th time window in scenario C k , and the length Δt of the time window k can be adjusted according to the requirements of the scenario;
[0033] Calculate various traffic statistical features for each time window W k,j .
[0034] Furthermore, the calculation of various traffic statistical features for each time window W k,j is specifically as follows:
[0035] Calculate the API call frequency f API , request size distribution μ s and σ s 2 , response time μ r and σ r 2 , request type distribution: API request P(T 1 ), resource loading request P(T 2 ), advertisement request P(T 3 ), other requests P(T 4 ), request type switching frequency f 切换 ;
[0036] Combine the traffic features in the W k time window of scenario C k,j into a feature vector x k,j and express it as the following formula:
[0037]
[0038] Combine the feature vectors of all time windows of scenario C k into a feature vector sequence X k (t):
[0039] X k (t) = [x 1 , x 2 ,..., x m
[0040] Combine the feature sequences of multiple time windows under all scenarios into a time series X(t):
[0041] X(t) = [X 1 (t), X 2 (t),..., X n (t)] = [x 1 , x 2 ,..., x T
[0042] where T is the total number of time windows, n represents the number of classified scenario categories, and this time series provides input for the subsequent multi-scale time series decomposition module.
[0043] Furthermore, the multi-scale time series decomposition module combines the flow characteristics in the time window W k,j into a feature vector; combines the feature vectors of all time windows into a feature vector sequence; arranges the feature vectors of each time window in sequence to form a feature vector sequence X: performs wavelet decomposition on the above feature vector sequence X to obtain the high-frequency component and low-frequency component of the flow data; performs discrete wavelet transform on the flow feature sequence X(t), uses the db4 wavelet basis for one-level decomposition, and obtains the low-frequency component L(t) and high-frequency component H(t). The calculation formula of wavelet transform is as follows:
[0044]
[0045] where N represents the number of sample points of the flow feature sequence; n represents the discrete time index, that is, the sample point position of the input flow feature sequence X(n); X(n) represents the input flow feature sequence; g(t - n) represents the high-pass filter coefficient in the db4 wavelet basis, h(t - n) represents the low-pass filter coefficient in the wavelet transform; t is the time index of the output signal.
[0046] Furthermore, the high-frequency component anomaly detection module determines the anomaly S H through the following formula:
[0047]
[0048] where AnomalyScore(H i ) is the anomaly score of the data point H i by the Isolation Forest algorithm, and n is the number of data points.
[0049] Furthermore, the low-frequency component anomaly detection module uses empirical mode decomposition to decompose the low-frequency component L(t) into multiple Intrinsic Mode Functions (IMFs) to capture detailed trends and patterns;
[0050] Let the input of the model be IMF k (t), and the output is the anomaly score of this IMF The anomaly score S L is weighted by the anomaly scores of each IMF, and the formula is as follows:
[0051]
[0052] where α k is the weight of each IMF, is the anomaly score of the IMF, k is the number of IMFs, and K is the number of all IMFs.
[0053] Furthermore, based on the anomaly detection results of the high-frequency component and the low-frequency component, the anomaly judgment module further evaluates the traffic risk of each time window; for each time window t, the high-frequency component anomaly score and the low-frequency component anomaly score are comprehensively judged and processed, and the high-frequency anomaly threshold T H and the low-frequency anomaly threshold T L are obtained through the training model in the dataset. Different risk levels are divided according to the degree to which the score exceeds the threshold; the anomaly types of high frequency and low frequency are integrated to generate traffic anomalies, persistent anomaly patterns, and long-term trend anomaly types.
[0054] The beneficial effects of the present invention are as follows:
[0055] It can capture both short-term and long-term anomaly patterns in APP traffic, realizing comprehensive detection of abnormal situations; it can distinguish sudden, persistent, and trend anomalies, accurately identify anomaly types, and provide knowledge support for network security situation assessment and immune decision-making; it makes full use of historical information to analyze current data, improving the accuracy of anomaly detection; it designs an anomaly detection method specifically for the characteristics of APP traffic, and can more accurately detect APP traffic anomalies; it adopts advanced time series analysis methods such as wavelet transform and empirical mode decomposition, improving the generalization ability and interpretability of the model.
[0056] Other advantages, objectives, and features of the present invention will be described to some extent in the subsequent description, and to some extent, will be obvious to those skilled in the art based on the study of the following text, or can be taught from the practice of the present invention. The objectives and other advantages of the present invention can be realized and obtained through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be described in detail with reference to the accompanying drawings. Among them,
[0058] Figure 1 is the flowchart of the APP traffic anomaly detection method based on multi-scale time series according to the embodiment of the present invention;
[0059] Figure 2 is the architecture diagram of the APP traffic anomaly detection system based on multi-scale time series according to the embodiment of the present invention;
[0060] Figure 3 This is the data processing flowchart of the APP traffic anomaly detection system based on multi-scale time series according to the embodiments of the present invention. Specific embodiments
[0061] The following uses specific specific examples to illustrate the embodiments of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the drawings provided in the following embodiments only illustrate the basic concept of the present invention in a schematic manner. Without conflict, the following embodiments and the features in the embodiments can be combined with each other.
[0062] Among them, the drawings are only for illustrative purposes, showing only schematic diagrams, not physical diagrams, and should not be construed as limiting the present invention; in order to better illustrate the embodiments of the present invention, some components in the drawings will be omitted, enlarged or reduced, and do not represent the size of the actual product; for those skilled in the art, it is understandable that some well-known structures and their descriptions in the drawings may be omitted.
[0063] In the drawings of the embodiments of the present invention, the same or similar reference numerals correspond to the same or similar components; in the description of the present invention, it should be understood that if there are terms such as "upper", "lower", "left", "right", "front", "rear", etc. indicating the orientation or positional relationship, it is based on the orientation or positional relationship shown in the drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, the terms describing the positional relationship in the drawings are only for illustrative purposes and should not be construed as limiting the present invention. For those of ordinary skill in the art, the specific meanings of the above terms can be understood according to specific circumstances.
[0064] Please refer to Figure 1 , which is the flowchart of the APP traffic anomaly detection method based on multi-scale time series according to the embodiments of the present invention;
[0065] The method includes the following steps:
[0066] Step 1: The traffic data acquisition module captures the APP traffic and classifies the traffic data according to different types such as API calls, resource loading, and advertisement requests.
[0067] Step 2: The traffic data preprocessing module uses the information recorded in the data collection module to identify the APP scenario to which each request belongs according to the relevant information in the request (such as specific characters, request content, etc.), group the data according to different scenarios. Different application scenarios may have different traffic behaviors and temporal characteristics. Therefore, adjust the time window according to the scenario category to more accurately capture these characteristics. Then segment the requests belonging to each scenario according to the defined time window. A smaller time window is used for scenarios with higher density to capture rapidly changing behaviors; a larger time window is used for low-density scenarios to reduce data fragmentation. Calculate statistical features such as API call frequency, request size distribution, response time, request type distribution, and request type switching frequency within each time window segment.
[0068] Step 3: The multi-scale time series decomposition module combines the statistical features calculated by the traffic data preprocessing module into a feature vector sequence. A feature vector sequence is generated for each time window, and the feature vector sequences of all time windows are combined into a time series. The traffic characteristics of different application scenarios may be reflected on different time scales. Wavelet transform can clearly separate these characteristics for further analysis. Therefore, perform wavelet transform on the time series to decompose it into high-frequency components reflecting short-term fluctuations and low-frequency components reflecting global trends. The low-frequency components still contain change patterns on multiple time scales. Therefore, the low-frequency components are further decomposed using empirical mode decomposition to obtain multiple intrinsic mode functions. Each intrinsic mode function represents the change of the sequence on a specific time scale, facilitating the capture of detailed trends and patterns.
[0069] Step 4: The abnormal traffic detection module is divided into two sub-modules, the high-frequency component abnormal detection module and the low-frequency component abnormal detection module: The high-frequency component abnormal detection module receives the high-frequency components after wavelet decomposition and uses the isolation forest to give the abnormal score of the high-frequency components by calculating the "isolation degree" of the samples. The low-frequency component abnormal detection module receives the multiple intrinsic mode functions obtained from empirical mode decomposition and uses the long short-term memory network to give the abnormal score of the low-frequency components. Then, combine the high-frequency and low-frequency abnormal scores to give the risk level and abnormal type. For each time window, if both the high-frequency abnormal score and the low-frequency abnormal score exceed the set threshold, the traffic of this window is considered abnormal.
[0070] Please refer to Figure 2 for the architecture diagram of the APP traffic abnormal detection system based on multi-scale time series according to the embodiment of the present invention;
[0071] The system includes: a traffic data collection module that captures the network traffic of the APP, performs packet capture, and classifies it according to the request type; a traffic data preprocessing module that identifies the running scenarios of the APP, divides time windows, and extracts the traffic statistical features of each window; a multi-scale time series decomposition module that decomposes the feature vector sequence into high-frequency components and low-frequency components through wavelet transform and empirical mode decomposition; an abnormal traffic detection module, including a high-frequency component abnormal detection module, a low-frequency component abnormal detection module, and an abnormal judgment module, which calculates the abnormal score using the isolation forest algorithm and the Long Short-Term Memory (LSTM) model, and comprehensively judges the abnormal type and risk level through the abnormal judgment module. Further, in step 1, the specific method for capturing and classifying APP traffic includes:
[0072] Capture the network traffic of the target APP, obtain the network request data packets, and perform preprocessing.
[0073] Classify the data according to different types of requests (such as API calls, resource loading, advertising requests, etc.) for subsequent processing; API requests usually access the API endpoints of the background server, and the Uniform Resource Locator (URL) contains a specific API path (such as / api / v1 / user); resource loading requests usually refer to requests for resources such as pictures, videos, Cascading Style Sheets (CSS), etc., and the URL contains a static resource path (such as.jpg,.css,.mp4, etc.); advertising requests can be identified by the domain name or a specific URL pattern (such as the domain name of the advertising service provider, keywords like ads, track, etc.); requests that do not belong to the previous three request classifications are classified as other requests.
[0074] The processing process of the traffic data preprocessing module includes:
[0075] While capturing the traffic data, identify and record the running status information of the APP according to the request content and the request URL, including but not limited to scenario information such as registration, login, payment, advertising, resource loading, etc. For example, the login status is usually accompanied by login requests (such as keywords like auth, login, token, etc.), and the request URL is generally fixed on a specific authentication server; the content loading status will have multiple resource requests, such as URL requests for pictures, videos, texts, etc.; mark and store the collected traffic data according to different application scenarios and timestamps to ensure that the time series data in subsequent analysis has accurate time and scenario information.
[0076] In step 2, the scenario recognition and grouping method includes:
[0077] Step 2.1: The traffic data preprocessing module receives the original traffic data and the corresponding context information provided by the traffic data collection module. Through these context tags, the traffic data is divided into different APP usage scenarios. Let the traffic volume of each scenario be C k (where k represents the scenario number), and a traffic data set D containing n different scenario identifications is obtained:
[0078] D = {C 1 , C 2 ,..., C n}
[0079] Among them, C 1 represents the traffic data of the registration scenario; C 2 represents the traffic data of the login scenario; C n represents the traffic data of other identifiable scenarios;
[0080] The traffic data within each scenario will be analyzed subsequently according to independent characteristics and patterns;
[0081] Step 2.2: For each scenario C k , the traffic data preprocessing module divides the time window according to the characteristics of this scenario. The purpose of the time window design is to group the requests within the scenario by an appropriate time period to capture the traffic characteristics within this scenario. Let the time window length of scenario C k be Δt k , then scenario C k can be divided into m time windows {W k,1 , W k,2 ,…, W k,m}:
[0082]
[0083] Among them, W k,j represents the j-th time window in scenario C k . The length Δt k of the time window can be adjusted according to the requirements of the scenario. For example, for the loading scenario, a shorter window is selected to capture frequent requests; while for the payment scenario, a longer window is selected to cover the entire payment process.
[0084] Step 2.3: In each time window W k,j , calculate various traffic statistical characteristics within this time period:
[0085] In each time window W k,j , calculate the API call frequency f API , the request size distribution μ s and σ s 2, Response time μ r and σ r 2 , Request type distribution P(T 1 ), P(T 2 ), P(T 3 ), P(T 4 ), which represent API requests, resource loading requests, advertisement requests, and other requests respectively. The request type switching frequency is f 切换 .
[0086] Combine the traffic characteristics in the W k time window of scenario C k,j into a feature vector x k,j :
[0087]
[0088] Combine the feature vectors of all time windows of scenario C k into a feature vector sequence X k (t):
[0089] X k (t) = [x 1 , x 2 ,..., x m
[0090] The time windows corresponding to different scenarios are different. This feature vector sequence represents the change in traffic characteristics of the scenario under its specific time window. Combine the feature sequences of multiple time windows under all scenarios into a time series X(t):
[0091] X(t) = [X 1 (t), X 2 (t),..., X n (t)] = [x 1 , x 2 ,..., x T
[0092] Among them, T is the total number of time windows, n represents the number of divided scenario categories. This time series represents the change in traffic characteristics of the captured traffic under different time windows and provides input for the subsequent multi-scale time series decomposition module.
[0093] API call frequency: Calculate the number of API calls in the time window W k,j . Suppose there are n API calls in the time window W k,j , then the API call frequency is:[[]]
[0094]
[0095] Request size distribution: Count the sizes of each request within the statistical time window. Use the mean μ s and the variance σ s 2 to represent the distribution characteristics of the request sizes:
[0096]
[0097] where s i represents the size of the i-th request;
[0098] Response time: Calculate the response time of each request within the time window and count its mean μ r and the variance σ r 2
[0099]
[0100] where r i represents the response time of the i-th request;
[0101] Request type distribution: The request type distribution represents the proportion of different request types (such as API requests, resource loading requests, advertisement requests) within a certain time window. Assume that within a time window, there are p types of request types (such as API requests, resource loading, advertisement requests, etc.). For the k-th type of request, its proportion P(T k ) is calculated as follows:
[0102]
[0103] where n k is the number of the k-th type of request, P(T 1 ) represents API requests, P(T 2 ) represents resource loading requests, P(T 3 ) represents advertisement requests, and N is the total number of requests within the time window.
[0104] Request type switching frequency: Calculate the number of switches of the request type within the time window, defined as the number of switches in the sequence of different types of requests. Assume there are m switches within the window, then the switching frequency is:
[0105]
[0106] Furthermore, in step 3, the functional implementation process of the multi-scale time series decomposition module includes:
[0107] Step 3.1: Combine the traffic characteristics in the time window W k,j into a feature vector. Let the feature vector of the j-th window be x k,j, including the above features, such as API call frequency, request size distribution, response time, request type distribution, and request type switching frequency, then:
[0108]
[0109] Combine the feature vectors of all time windows into a feature vector sequence, representing the traffic feature changes of this scenario under different time windows. Arrange the feature vectors of each time window in sequence to form the feature vector sequence X:
[0110] X = [x 1 , x 2 ,..., x T
[0111] where T is the total number of time windows, and the multi-scale time series decomposition module performs wavelet decomposition on the above feature vector sequence X to obtain the high-frequency component and low-frequency component of the traffic data.
[0112] X = L(t) + H(t)
[0113] where L(t) represents the low-frequency component; H(t) represents the high-frequency component.
[0114] Step 3.2: Perform discrete wavelet transform (Discrete Wavelet Transform, DWT) on the traffic feature sequence X(t), use the db4 (Daubechies 4) wavelet basis for one-level decomposition to obtain the low-frequency component L(t) and the high-frequency component H(t). The calculation formula of wavelet transform is as follows:
[0115]
[0116] where N represents the number of sample points of the traffic feature sequence; n represents the discrete time index, that is, the sample point position of the input traffic feature sequence X(n); X(n) represents the input traffic feature sequence; g(t - n) represents the high-pass filter coefficient in the db4 wavelet basis, and h(t - n) represents the low-pass filter coefficient in the wavelet transform; t is the time index of the output signal.
[0117] Step 3.3. Analyze the low-frequency component L(t) extracted from the flow data by wavelet transform using Empirical Mode Decomposition (EMD): Extract the local extreme values of L(t) and construct the upper and lower envelope lines; calculate the mean of the upper and lower envelope lines and subtract it from L(t) to obtain a new signal h(t); determine whether h(t) meets the conditions of the Intrinsic Mode Function (IMF). If it meets the conditions, extract it as the first IMF; if h(t) does not meet the IMF conditions, continue to extract its extreme values, calculate the envelope lines and subtract them until an IMF that meets the conditions is obtained; repeat this process until all IMFs are obtained; finally, the low-frequency component L(t) is decomposed into multiple IMFs, representing the flow changes at different time scales.
[0118] The functional implementation process of the abnormal flow detection module in Step 4 includes:
[0119] Step 4.1. Receive the high-frequency component data H(t) after wavelet transform; adopt the Isolation Forest algorithm for high-frequency component anomaly detection. H(t) is the high-frequency component data after wavelet transform. The Isolation Forest algorithm constructs multiple decision trees and calculates the anomaly score based on the splitting path of each data point in the tree. The anomaly score S H is calculated by the following formula:
[0120]
[0121] where AnomalyScore(H i ) is the anomaly score of the data point H i by the Isolation Forest algorithm, and n is the number of data points.
[0122] Step 4.2. Receive the low-frequency component L(t) obtained from wavelet decomposition, and further decompose the low-frequency component L(t) using Empirical Mode Decomposition. The low-frequency component L(t) is decomposed into multiple Intrinsic Mode Functions (IMFs), namely IMF 1 (t), …, IMF k (t); representing the flow changes at a specific time scale to capture detailed trends and patterns.
[0123] For each IMF, adopt the Long Short-Term Memory (LSTM) network for time series anomaly detection. The input of the model is IMF k (t), and the output is the anomaly score of this IMF The anomaly score S L is obtained by weighting the anomaly scores of each IMF. The formula is as follows:
[0124]
[0125] where α k is the weight of each IMF, is the anomaly score of the IMF, k is the number of IMFs, and K is the total number of all IMFs.
[0126] The processing procedures of the high- and low-frequency component anomaly detection modules in the abnormal traffic detection module include:
[0127] 1) High-frequency component anomaly detection module: Receive the high-frequency component data H(t) from wavelet transform; Use the Isolation Forest algorithm for high-frequency component anomaly detection. Specifically, assume that H(t) is the high-frequency component data after wavelet transform. The Isolation Forest algorithm constructs multiple decision trees. In a tree, the smaller the number of splits h(t) required for a data point t to be isolated (i.e., split into a separate leaf node), the more abnormal t is. The anomaly score formula is calculated as follows:
[0128] First, calculate the average value of the path length h(t) of x in all trees:
[0129]
[0130] where n t is the total number of trees, and h i (t) is the path length of point t in the i-th tree;
[0131] After that, calculate the anomaly score S H (t) by normalizing the path length:
[0132]
[0133] where c(n) is the normalization factor, representing the expected path length of the tree:
[0134]
[0135] H(n) is the n-th harmonic number, and the value range of S H (t) is [0, 1]. The closer it is to 1, the faster point t is isolated and the more likely it is to be abnormal.
[0136] 2) Low-frequency component anomaly detection module: Receive the low-frequency component L(t) obtained from wavelet decomposition. The low-frequency component L(t) represents the traffic change on a long time scale and is the overall trend of the traffic data. Abnormal behaviors are manifested as large deviations in the low-frequency component, such as gradually accumulating changes or systematic anomalies; Then use EMD to decompose the low-frequency component L(t) to obtain multiple intrinsic mode functions, namely IMF 1 (t), IMF 2(t), …, IMF k (t), EMD decomposes the low-frequency component L(t) into IMFs of different time scales. These IMFs do not depend on preset parameters and can adapt to the complexity of traffic data; for each IMF k (t), a long short-term memory network (LSTM) is used for anomaly detection of time series. LSTM has the ability to capture long-term dependencies in time series and is suitable for processing anomaly detection scenarios. LSTM can learn its time dynamic features and give the anomaly score S of this IMF IMFK :
[0137]
[0138] Among them, is the value predicted by LSTM, and IMF k (t) is the actual value. The larger the score, the more abnormal the behavior at this time point. Then, by synthesizing the anomaly scores of each IMF, a global anomaly score S L (t) is obtained:
[0139]
[0140] Based on the anomaly detection results of high-frequency and low-frequency components, the comprehensive module further evaluates the traffic risk of each time window.
[0141] Step 4.3. For each time window t, the high-frequency component anomaly score S H and the low-frequency component anomaly score S L are comprehensively judged and processed. The high-frequency anomaly threshold TH and the low-frequency anomaly threshold TL are obtained through the training model in the dataset, and different risk levels are divided according to the degree of the score exceeding the threshold.
[0142] Step 4.4. Integrate the high-frequency and low-frequency anomaly types to generate specific anomaly types: the high-frequency anomaly score is high and the low-frequency anomaly score is low: judged as a sudden traffic anomaly, such as a short-term distributed denial of service (DDoS) attack or an instantaneous traffic spike; both the high-frequency and low-frequency scores are high: judged as a persistent anomaly pattern, such as an application being hijacked or malicious code continuously active; the high-frequency anomaly score is low and the low-frequency anomaly score is high: judged as a long-term trend anomaly, such as data leakage or resource abuse.
[0143] Please refer to Figure 3, is the data processing flowchart of the APP traffic anomaly detection system based on multi-scale time series according to the embodiments of the present invention; the system first captures the network traffic of the target APP through the traffic data collection module, classifies it according to API calls, resource loading, and advertisement request types, and records the APP running status as context information. Then, the traffic data preprocessing module groups the collected traffic data by scenario and extracts statistical features based on a set time window. Subsequently, the multi-scale time series decomposition module performs wavelet transform and empirical mode decomposition on the feature vector sequence to extract the high-frequency component and low-frequency component of the traffic respectively, capturing short-term bursts and long-term trend changes. The abnormal traffic detection module uses the isolation forest algorithm to detect anomalies in the high-frequency component, and at the same time detects anomalies in the low-frequency component based on the long short-term memory network (LSTM). After that, the detection results of the two are comprehensively evaluated and the abnormal type is given.
[0144] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the purpose and scope of the present technical solution, and they should all be covered within the scope of the claims of the present invention.
Claims
1. A method for detecting APP traffic anomaly based on multi-scale time series, characterized in that: The following steps are involved: Step 1: Traffic data collection: Capture the network traffic of the target application (Application, APP) and classify it by API call, resource loading, and advertising request type; Step 2: Traffic data preprocessing: Identify the APP running scenario based on the request content and Uniform Resource Locator (URL), group the traffic data by scenario, divide the time window based on the scenario characteristics, and calculate the (Application Programming Interface, API) call frequency, request size distribution, response time, request type distribution, and request type switching frequency in each time window; Step 3, multi-scale time series decomposition: combine the statistical features of each time window into a feature vector sequence, decompose it into high-frequency components and low-frequency components through wavelet transform, and perform empirical mode decomposition on the low-frequency components to obtain multiple intrinsic mode functions; Step 4: Use the isolation forest algorithm to calculate the anomaly score for the high-frequency component, and use the long short-term memory network to calculate the anomaly score for the intrinsic mode function of the low-frequency component. Use the high and low frequency anomaly scores to determine the traffic risk level and anomaly type.
2. According to claim 1, a method for detecting APP traffic anomaly based on multi-scale time series is characterized in that: The identification of the APP running scene in step 2 is: matching a preset scene tag according to a keyword or domain name in the request URL, and the scene tag is registration, login, payment, advertising or resource loading; The time window division rule in step 2 is: a small time window is used for a high-density scenario, and a long time window is used for a low-density scenario. The high-density scenario includes resource loading or advertisement request, and the low-density scenario includes payment or login. In step 3, the db4 wavelet basis is used to perform discrete wavelet transform on the feature vector sequence to decompose it into high-frequency components and low-frequency components; The method of using high and low frequency anomaly scores in step 4 is: if both the high frequency anomaly score and the low frequency anomaly score exceed the preset threshold, it is determined to be a persistent anomaly; if only the high frequency score exceeds the threshold, it is determined to be a sudden anomaly; if only the low frequency score exceeds the threshold, it is determined to be a trend anomaly.
3. An APP traffic anomaly detection system based on multi-scale time series, characterized in that: include: Traffic data collection module captures APP network traffic and classifies it by request type; The traffic data preprocessing module identifies the APP operation scenarios and divides the time windows to extract the traffic statistics characteristics of each window; Multi-scale time series decomposition module, which decomposes the feature vector sequence into high-frequency components and low-frequency components through wavelet transform and empirical mode decomposition; The abnormal traffic detection module, including the high-frequency component anomaly detection module, the low-frequency component anomaly detection module and the anomaly judgment module, uses the isolation forest algorithm and the long short-term memory network (Long Short-Term Memory, LSTM) model to calculate the anomaly score, and outputs the anomaly type and risk level through the anomaly judgment module.
4. The APP traffic anomaly detection system based on multi-scale time series according to claim 3 is characterized by: The traffic data preprocessing module further includes a scene recognition unit, which associates preset scene tags by matching URL keywords or domain names; In the multi-scale time series decomposition module, the decomposition of low-frequency components uses empirical mode decomposition to generate multiple intrinsic mode functions; The isolation forest algorithm of the high-frequency component anomaly detection module determines the anomaly score by calculating the path length of the data point, and the shorter the path length, the higher the anomaly score; The LSTM model of the low-frequency component anomaly detection module calculates the anomaly score through the residual between the predicted value and the actual value. The larger the residual, the higher the anomaly score.
5. The APP traffic anomaly detection system based on multi-scale time series according to claim 4 is characterized by: The traffic data preprocessing module divides the traffic data into different APP usage scenarios, assuming that the traffic of each scenario is C k , k represents the scene number, and the traffic data set D containing n different scene recognitions is expressed as the following formula: D={C1,C2,…,C n } Among them, C1 represents the traffic data of the registration scenario; C2 represents the traffic data of the login scenario; C n Indicates other identifiable scene traffic data; For each scenario C k The characteristics of the time window are divided to capture the traffic characteristics in the scene. Suppose scene C k The time window length is Δt k , then scene C k Divided into m time windows {W k,1 ,W k,2 ,…,W k,m }: Among them, W k,j Indicates scene C k The jth time window in the time window, the length of the time window Δt k It can be adjusted according to the needs of the scene; Calculate each time window W k,j Various traffic statistics characteristics.
6. The APP traffic anomaly detection system based on multi-scale time series according to claim 5 is characterized by: The calculation of each time window W k,j The traffic statistics characteristics are as follows: Calculate the API call frequency f within the time period API , Request size distribution μ s and σ s 2 , response time μ r and σ r 2 , Request type distribution: API request P(T1), resource loading request P(T2), advertising request P(T3), other request P(T4), request type switching frequency f 切换 ; Scene C k W k,j The traffic features in the time window are combined into a feature vector x k,j It is expressed as the following formula: Scene C k The feature vectors of all time windows are combined into a feature vector sequence X k (t): X k (t)=[x1,x2,…,x m ] Combine the feature sequences of multiple time windows in all scenarios into a time series X(t): X(t)=[X1(t),X2(t),...,X n (t)]=[x1,x2,…,x T ] Among them, T is the total number of time windows, n represents the number of divided scene categories, and this time series provides input for the subsequent multi-scale time series decomposition module.
7. The APP traffic anomaly detection system based on multi-scale time series according to claim 4 is characterized by: The multi-scale time series decomposition module divides the time window W k,j The various flow characteristics in are combined into a feature vector; the feature vectors of all time windows are combined into a feature vector sequence; the feature vectors of each time window are arranged in sequence to form a feature vector sequence X: the above feature vector sequence X is subjected to wavelet decomposition to obtain the high-frequency component and low-frequency component of the flow data; the flow feature sequence X(t) is subjected to discrete wavelet transform, and the db4 wavelet basis is used for primary decomposition to obtain the low-frequency component L(t) and the high-frequency component H(t). The calculation formula of the wavelet transform is as follows: Among them, N represents the number of sample points of the flow characteristic sequence; n represents the discrete time index, that is, the sample point position of the input flow characteristic sequence X(n); X(n) represents the input flow characteristic sequence; g(tn) represents the high-pass filter coefficient in the db4 wavelet basis, h(tn) represents the low-pass filter coefficient in the wavelet transform; t is the time index of the output signal.
8. The APP traffic anomaly detection system based on multi-scale time series according to claim 4 is characterized by: The high frequency component anomaly detection module determines the anomaly S H The calculation is done by the following formula: Among them, AnomalyScore (H i ) is the isolation forest algorithm for the data point H i The anomaly score of , n is the number of data points.
9. The APP traffic anomaly detection system based on multi-scale time series according to claim 4 is characterized by: The low-frequency component anomaly detection module uses empirical mode decomposition to decompose the low-frequency component L(t) into multiple intrinsic mode functions (IMFs) to capture detailed trends and patterns; Assume that the input of the model is IMF k (t), the output is the anomaly score of the IMF Abnormal score S L It is obtained by weighting the anomaly score of each IMF, and the formula is as follows: Among them, α k is the weight of each IMF, is the anomaly score of the IMF, k is the number of IMFs, and K is the number of all IMFs.
10. The APP traffic anomaly detection system based on multi-scale time series according to claim 4 is characterized by: The anomaly judgment module further evaluates the traffic risk of each time window based on the anomaly detection results of high-frequency components and low-frequency components; for each time window t, the high-frequency component anomaly score and the low-frequency component anomaly score are comprehensively judged and processed, and the high-frequency anomaly threshold T is obtained by training the model in the data set. H and low frequency anomaly threshold T L , divided into different risk levels according to the degree to which the score exceeds the threshold; High-frequency and low-frequency anomaly types are integrated to generate flow anomalies, persistent anomaly patterns, and long-term trend anomaly types.
Citation Information
Patent Citations
Network traffic abnormality detection method and system
CN106411597A
Shape-based network abnormal flow detection method and system
CN116232761A
System load prediction method integrating isolated forest and long-term and short-term memory network
CN111738520A
Network traffic abnormality online detection method and system, computer equipment and storage medium
CN112329713A
Dam safety monitoring data anomaly detection method based on unsupervised learning
CN113076975A
Cited By
Network access anomaly detection method and device, equipment and storage medium
CN122069120A
A network access anomaly detection method, device, equipment and storage medium
CN122069120B