A SSR flow identification system and method based on machine learning
Through the machine learning-based SSR traffic recognition system, the random forest algorithm is used for identification, and the problems of high complexity and reduced accuracy of SSR traffic recognition in the prior art are solved, and real-time identification and supervision of SSR traffic under large-scale gateways are realized.
Patent Information
- Application Number
- CN202111370935.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-18
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2041-11-18
AI Technical Summary
The prior art is difficult to effectively identify and monitor SSR traffic, especially in the problem of high algorithm complexity and reduced cross-device identification accuracy, and it is difficult to apply to actual SSR traffic monitoring systems.
The SSR traffic recognition system based on machine learning is adopted to extract data flow statistical features through packet capture, processing and analysis, and identify them using a random forest algorithm to achieve real-time acquisition and recognition.
Real-time identification of SSR traffic under large-scale gateways improves the network security department's ability to supervise network traffic, and the identification model is strongly robust.
Smart Images

Figure CN114091602B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of information security technology, and further relates to traffic identification, and specifically to a SSR traffic identification system and method based on machine learning, which can be used for the detection and review of SSR traffic by public security or enterprise network security departments. Background Art
[0002] The anonymous proxy system based on virtual private servers not only protects user privacy and data security, but also facilitates illegal and criminal activities. As a typical and widely used anonymous proxy system SSR (ShadowsocksR), the traffic generated by this proxy system is called SSR traffic, and the network protocol involved is called SSR protocol. The SSR system has the characteristics of easy deployment, good communication quality, strong anonymity, high security, and not easy to be monitored. It is often used to penetrate firewalls and bypass supervision and review, which makes it convenient for criminals to engage in illegal network activities. In order to effectively supervise, trace and collect evidence of cybercrime activities, an effective traffic identification method for SSR is needed.
[0003] Due to the special mechanisms of the SSR protocol, including double encryption mechanism, transparent data transmission mechanism and invisible key negotiation mechanism, the existing common encrypted traffic identification methods and the virtual private network VPN (Virtual Private Network) traffic identification methods based on the IPSec protocol cannot effectively identify the SSR and its proxy App traffic. In addition, the existing SSR traffic identification methods also have problems such as low efficiency due to high algorithm complexity and low cross-device identification accuracy, making them difficult to directly apply to the actual SSR traffic monitoring system.
[0004] Traditional VPNs based on IPSec protocol have certain characteristics due to key negotiation and other processes. With the continuous development of encrypted traffic identification technology based on machine learning, VPNs with obvious traffic fingerprints are becoming increasingly difficult to use. With the continuous development of anonymous proxy technology, anonymous proxies such as SSR are becoming more and more widely used due to their stronger anonymity. At present, there are few studies on the identification of SSR traffic. In 2017, Deng et al. also used the random forest algorithm to realize the identification of SSR traffic. They extracted 3,000 features to form a 3,000-dimensional vector, and trained and classified 1GB of SSR traffic and 10GB of ordinary traffic. The results improved with the increase of the size of the training set and the test set, with the highest accuracy of 92%; however, their description of feature extraction is relatively vague, and only a partial feature list (9) is given. The evaluation criteria for the results are only the accuracy rate, which cannot accurately reflect the various indicators, and the cross-device identification effect is unknown. In 2019, Zeng et al. proposed a method for identifying SS (ShadowSocks) traffic. SS is the predecessor of the SSR anonymous system. SSR adds traffic obfuscation features on its basis, which greatly increases the difficulty of identification. This method analyzes the operating mechanism of the SS proxy system, and finds that SS traffic and ordinary traffic differ in flow context, flow host behavior, and host behavior on DNS, thereby extracting features in these aspects, and then using the random forest algorithm for model training. The final results show that its recognition accuracy has reached 93.43%. However, this method uses a sliding window to extract the features of all traffic within a period of time, which is a relatively complex method. On the other hand, the extracted domain name resolution DNS features are based on the DNS leak in a certain version of SS, but the current commonly used SSR software version does not have this vulnerability, so this feature is generally facing failure and cannot effectively identify SSR traffic. Summary of the invention
[0005] The purpose of the present invention is to address the deficiencies of the above-mentioned prior art and propose a SSR traffic identification system and method based on machine learning, which is used to solve the problem that the existing protocol identification cannot adapt to the constantly changing and unpredictable network traffic in the actual network environment. First, all the data packet information of the traffic is obtained by capturing the traffic flowing through the gateway network card, and the data packets are integrated and filtered to obtain pure data flow information. Then, the statistical characteristics of the data flow are extracted, and the data flow feature vector is judged by machine learning to derive the identification result. The present invention can achieve real-time collection and identification of SSR traffic under larger-scale gateways, which improves the network security department's ability to supervise network traffic.
[0006] The specific scheme of the present invention to achieve the above purpose is as follows:
[0007] The system of the present invention comprises: a data acquisition and identification unit, an identification information storage module and a data analysis and display unit; wherein the data acquisition and identification unit is composed of a data packet capture module, a data packet processing module, a data packet analysis module and a data packet identification module which are connected in sequence in a unidirectional manner, and the data analysis and display unit is composed of an identification result analysis module and a web interface; the identification information storage module is respectively connected to the data acquisition and identification unit and the data analysis and display unit;
[0008] The data packet capture module is used to obtain network data traffic;
[0009] The data packet processing module is used to extract basic information of the data packet from the network data flow obtained by the data packet capture module;
[0010] The data packet analysis module is used to pre-process the data packet according to the basic information obtained by the data packet processing module to obtain the pre-processed flow information;
[0011] The data packet identification module is used to identify the pre-processed flow information obtained by the data packet analysis module to obtain an identification result;
[0012] The identification information storage module is used to store the identification results obtained by the data packet identification module in the data acquisition and identification unit, and to be called by the identification result analysis module in the data analysis and display unit;
[0013] The identification result analysis module is used to perform real-time analysis on the information stored in the identification information storage module, and display the analysis results on a web interface for analysts to query.
[0014] Furthermore, the basic information of the above data packet at least includes payload characteristics, length and time.
[0015] Furthermore, the above-mentioned data packet analysis module pre-processes the data packet according to the basic information obtained by the data packet processing module, specifically performing traffic grouping and filtering operations; filtering includes: filtering out data packets of all protocols except TCP protocol, and filtering out data packets that are retransmitted due to network connection anomalies.
[0016] Furthermore, the data packet identification module identifies the pre-processed traffic information obtained by the data packet analysis module, specifically extracts features from the packet data stream in the pre-processed traffic information, and then uses machine learning to complete the identification.
[0017] The steps of the method of the present invention include:
[0018] (1) Capture data traffic based on the traffic arrival of the device network card:
[0019] (1.1) Estimate the gateway traffic scale, set the single capture order of magnitude and initial queuing time based on the evaluation results, and ensure that the single round of data capture time is within 30-45 seconds;
[0020] (1.2) Design a real-time system redundancy mechanism, that is, set a dynamic waiting time, which is calculated in real time based on the system's internal memory usage ratio, processor computing workload, and the number of capture file queues;
[0021] (1.3) According to the pipeline method, the data packet capture module is called cyclically to obtain network data traffic;
[0022] (2) extracting basic information of data packets from network data traffic through the data packet processing module to obtain data traffic load information including load characteristics, length, and time;
[0023] (3) Preprocessing data packets using data traffic load information:
[0024] (3.1) The data packet analysis module filters the data packets according to the load characteristics of the data flow, filters out the data packets of all protocols except the TCP protocol, retains only the TCP data packets, and filters out the data packets that are retransmitted due to abnormal network connection, and obtains the data packet set R:
[0025] R = {pkg 1 ,pkg 2 ,...,pkg i ,...,pkg r},
[0026] Among them, pkg i represents the i-th data packet in the set R, i = 1, 2, ..., r, and r represents the total number of filtered data packets;
[0027] (3.2) The data packet analysis module groups data packets according to the following rules:
[0028] (3.2.1) Extract data package pkg i Source IP address src-i , Source Port src-i 、Destination IP address dst-i , Destination Port dst-i and transport layer protocol proto i Five types of information, and form a data packet pkg i Head i :
[0029] h i =(IP src-i ,Port src-i ,IPdst-i ,Port dst-i ,proto i ),
[0030] pkg i ={h i ,Len(pkg i ),stime i};
[0031] Among them, Len(pkg i ) indicates data packet pkg i The length of stime i Indicates data packet pkg i Arrival time;
[0032] (3.2.2) In the data packet set R, for the data packet pkg i The same or opposite packet, its header and pkg i Constitute a packet data stream;
[0033] (3.2.3) Take i = 1, 2, ..., r and follow steps (3.2.1)-(3.2.2) to obtain the packet data stream corresponding to each data packet in the data packet set R. All packet data streams together constitute the grouped data stream set D, that is, the preprocessed traffic information:
[0034] D = {flow 1 ,flow 2 ,...,flow k ,...,flow d},
[0035] Among them, flow k represents the kth packet data flow, k=1,2,...,d, d represents the total number of packet data flows;
[0036] (4) The data packet identification module extracts features from the packet data streams in the data stream set D and screens them, using machine learning to identify them:
[0037] (4.1) Statistics of packet data flow k The number of all packets in the flow is recorded as total(flow k ), record all the data packets with the same sending direction as the first data packet as output packets, and the remaining data packets as input packets;
[0038] (4.2) Calculate flow separately kStatistics of all input packets, all output packets, and all data packet lengths: mean, minimum, maximum, absolute deviation, absolute median, standard deviation, variance, skew, kurtosis, 10%-90% percentiles;
[0039] (4.3) The statistical values obtained in step (4.2) are combined into flow k The statistical eigenvector of PLS k , the statistical characteristic vectors corresponding to all packet data flows together constitute the packet length statistical characteristic matrix PLS;
[0040] (4.4) Perform forward search and combined feature screening on the features in the packet length statistical feature matrix PLS, divide the features into two categories: positive features and negative features, and perform forward search again until the result is optimal, and obtain the optimized packet length statistical feature matrix PLS';
[0041] (4.5) Inputting the matrix PLS' into the model trained based on the random forest algorithm for recognition, obtaining the recognition result, and storing the result in the recognition information storage module;
[0042] (5) The identification information storage module divides the identification results into two categories: SSR results and all results, and stores them in a specific database mysql with the data stream start time as the index;
[0043] (6) The recognition result analysis module performs real-time analysis on the recorded information in the MySQL database and outputs the analysis results:
[0044] (6.1) For the recognition results in the MySQL database over a period of time, statistics are performed and the scores are calculated:
[0045]
[0046] Among them, Num ssr Indicates the number of SSR traffic identified, Num all Indicates the total number of data streams, Num dst Indicates the number of communication destination addresses;
[0047] (6.2) Rank the SSR traffic used by different devices according to the score, and dynamically set different confidence levels to obtain multi-dimensional traffic analysis results for a single user;
[0048] (6.3) Display the analysis results on the web interface.
[0049] Compared with the prior art, the present invention has the following advantages:
[0050] First, compared with the prior art, the present invention proposes a real-time identification system for SSR traffic for the first time, and puts the identification technology into practice in large-scale gateways;
[0051] Second, the present invention optimizes the existing SSR recognition method by screening the features through forward search and combined search, extracting stable machine learning features, so that the recognition model has strong robustness in different network environments;
[0052] Third, since the present invention adopts a stream processing data processing mode as a whole, new data is continuously merged to calculate results, which significantly improves the system's data processing speed and reduces the impact of complex machine learning time-consuming calculation steps on the overall system operation, thereby achieving the design goal of real-time computing. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] Figure 1 It is a schematic diagram of the overall architecture of the system of the present invention;
[0054] Figure 2 Flow chart for realizing the method of the present invention;
[0055] Figure 3 This is a schematic diagram of a flow data collection scenario of the present invention;
[0056] Figure 4 Schematic diagram of the flow processing calculation mode in the method of the present invention. DETAILED DESCRIPTION
[0057] The present invention will be further described below in conjunction with the accompanying drawings.
[0058] Embodiment 1: Refer to the attached Figure 1 The present invention proposes an SSR traffic identification system based on machine learning, comprising: a data acquisition and identification unit, an identification information storage module and a data analysis and display unit; wherein the data acquisition and identification unit is composed of a data packet capture module, a data packet processing module, a data packet analysis module and a data packet identification module which are connected in sequence in a unidirectional manner, and the data analysis and display unit is composed of an identification result analysis module and a web interface; the identification information storage module is respectively connected to the data acquisition and identification unit and the data analysis and display unit;
[0059] The data packet capture module is used to obtain network data traffic;
[0060] The data packet processing module is used to extract basic information of the data packet from the network data flow obtained by the data packet capturing module, and the basic information includes at least load characteristics, length and time.
[0061] The data packet analysis module is used to pre-process the data packet according to the basic information obtained by the data packet processing module to obtain the pre-processed traffic information; specifically, it performs traffic grouping and filtering operations; filtering includes: filtering out data packets of all protocols except the TCP protocol, and filtering out data packets that are retransmitted due to abnormal network connections. The facility of this module takes into account that ShadowsocksR basically uses the TCP protocol, so it is more appropriate to filter out data packets of all other protocols and only retain the TCP data packets in RawData. In addition, abnormal packets such as TCP retransmission packets may interfere with the results of application identification, so they cannot be used; when an abnormality occurs in the network connection, various control measures in the TCP protocol will be activated, and a series of data packet retransmission and other behaviors will be performed. The data packets generated by this behavior often carry more redundant information and abnormal information, and cannot be used for training and identification, so it is also necessary to filter out these data packets that are retransmitted due to abnormal network connections.
[0062] The data packet identification module is used to identify the pre-processed flow information obtained by the data packet analysis module to obtain an identification result; specifically, it extracts features from the packet data flow in the pre-processed flow information and then uses machine learning to complete the identification.
[0063] The identification information storage module is used to store the identification results obtained by the data packet identification module in the data acquisition and identification unit, and to be called by the identification result analysis module in the data analysis and display unit; this module stores the identification results in the database according to their attribution type and with the data flow start time as the index.
[0064] The identification result analysis module is used to perform real-time analysis on the information stored in the identification information storage module, and display the analysis results on a web interface for analysts to query.
[0065] Embodiment 2: Refer to the attached Figure 2 The present invention proposes a method for flow identification using an SSR flow identification system based on machine learning, and the specific implementation steps are as follows:
[0066] Step 1. Capture data traffic based on the traffic arrival of the device network card:
[0067] (1.1) Estimate the gateway traffic scale, set the single capture order of magnitude and initial queuing time based on the evaluation results, and ensure that the single round of data capture time is within 30-45 seconds;
[0068] (1.2) Design a real-time system redundancy mechanism, that is, set a dynamic waiting time, which is calculated in real time based on the system's internal memory usage ratio, processor computing workload, and the number of capture file queues;
[0069] (1.3) According to the pipeline method, the data packet capture module is called cyclically to obtain network data traffic;
[0070] Step 2. Extract basic information of data packets from network data traffic through the data packet processing module to obtain data traffic load information including load characteristics, length, and time;
[0071] Step 3. Pre-process the data packets using the data traffic load information:
[0072] (3.1) The data packet analysis module filters the data packets according to the load characteristics of the data flow, filters out the data packets of all protocols except the TCP protocol, retains only the TCP data packets, and filters out the data packets that are retransmitted due to abnormal network connection, and obtains the data packet set R:
[0073] R = {pkg 1 ,pkg 2 ,...,pkg i ,...,pkg r},
[0074] Among them, pkg i represents the i-th data packet in the set R, i = 1, 2, ..., r, and r represents the total number of filtered data packets;
[0075] (3.2) The data packet analysis module groups data packets according to the following rules:
[0076] (3.2.1) Extract data package pkg i Source IP address src-i , Source Port src-i 、Destination IP address dst-i , Destination Port dst-i and transport layer protocol proto i Five types of information, and form a data packet pkg i Head i , that is, the five-tuple of data packets:
[0077] h i =(IP src-i ,Port src-i ,IP dst-i ,Port dst-i ,proto i ),
[0078] pkg i ={h i ,Len(pkg i ),stime i};
[0079] Among them, Len(pkg i ) indicates data packet pkg i The length of stime i Indicates data packet pkg i Arrival time;
[0080] (3.2.2) In the data packet set R, for the data packet pkg i The same or opposite packet, its header and pkg i Constitute a packet data stream;
[0081] (3.2.3) Take i = 1, 2, ..., r and follow steps (3.2.1)-(3.2.2) to obtain the packet data stream corresponding to each data packet in the data packet set R. All packet data streams together constitute the grouped data stream set D, that is, the preprocessed traffic information:
[0082] D = {flow 1 ,flow 2 ,...,flow k ,...,flow d},
[0083] Among them, flow k represents the kth packet data flow, k=1,2,...,d, d represents the total number of packet data flows;
[0084] Step 4. The data packet identification module extracts features from the packet data streams in the data stream set D and screens them, using machine learning for identification:
[0085] (4.1) Statistics of packet data flow k The number of all packets in the flow is recorded as total(flow k ), record all the data packets with the same sending direction as the first data packet as output packets, and the remaining data packets as input packets;
[0086] (4.2) Calculate flow separately k The statistical values of all input packets, all output packets and all data packet lengths: mean, minimum, maximum, absolute difference, absolute median, standard deviation, variance, skew, kurtosis, 10%-90% percentile. In this embodiment, the above 19 statistical values are counted for the above three types of data packets respectively, and a total of 57 dimensions of packet length statistical features are obtained. The 10%-90% percentiles that need to be counted here can be calculated as follows:
[0087] (4.2.1) Let per% denote any percentile from 10% to 90%;
[0088] (4.2.2) Packet data flow k total(flow k ) data packet lengths are arranged in ascending order to obtain the length of the sorted data packet;
[0089] (4.2.3) Select the λth packet length from the sorted packet lengths and calculate per% according to the following formula:
[0090]
[0091] in, Represents round up.
[0092] (4.3) The statistical values obtained in step (4.2) are combined into flow k The statistical eigenvector of PLS k , the statistical feature vectors corresponding to all packet data flows together form the packet length statistical feature matrix PLS. The length statistical features of the obtained data packets can reflect the distinctive characteristics of SSR traffic and non-SSR traffic from the perspectives of the average size of the traffic and the difference in the length of the data packets, so that the classification model established by the machine learning algorithm can more accurately identify the traffic type.
[0093] (4.4) Perform forward search and combined feature screening on the features in the packet length statistical feature matrix PLS, divide the features into two categories: positive features and negative features, and perform forward search again until the result is optimal, and obtain the optimized packet length statistical feature matrix PLS';
[0094] (4.5) Inputting the matrix PLS' into the model trained based on the random forest algorithm for recognition, obtaining the recognition result, and storing the result in the recognition information storage module;
[0095] Step 5. The identification information storage module divides the identification results into two categories: SSR results and all results, and stores them in a specific database mysql with the data stream start time as the index;
[0096] Step 6. The recognition result analysis module performs real-time analysis on the recorded information in the MySQL database and outputs the analysis results:
[0097] (6.1) For the recognition results in the MySQL database over a period of time, statistics are performed and the scores are calculated:
[0098]
[0099] Among them, Num ssr Indicates the number of SSR traffic identified, Numall Indicates the total number of data streams, Num dst Indicates the number of communication destination addresses;
[0100] According to the above formula, the score and the recognition result have the following relationship:
[0101] a. The ratio of the number of SSR data streams used by the device in the identification result to the total number of data streams of the device is proportional to the score;
[0102] b. The number of SSR data streams used by the device in the recognition result is proportional to the score;
[0103] c. The number of devices with foreign IP addresses in the identification results is inversely proportional to the score.
[0104] (6.2) SSR traffic used by different devices is ranked according to the score, and different confidence levels are dynamically set to obtain multi-dimensional traffic analysis results for a single user; here, the SSR traffic used by different devices is ranked from high to low according to the score, where the higher the score, the greater the probability that the device uses SSR traffic. For different confidence levels, dynamic settings can be used to obtain multi-dimensional traffic identification results for a single user, which can help determine user behavior to a certain extent.
[0105] (6.3) The analysis results are displayed on the web interface. In this step of the present embodiment, for the results identified as SSR traffic, the JavaScript-based visualization chart library Echarts is dynamically loaded, and the data is updated in real time to the front-end web page, so that the monitoring reaches a visualization level.
[0106] Embodiment 3: Refer to the attached Figure 3 and Figure 4 To further describe the method of the present invention, this embodiment is based on the method steps of the second embodiment. First, the data of the network card in the server where the system is located is captured to obtain a data flow pcap file, and the information (IP, port, packet timestamp, packet size) is parsed. Then, the data packets are divided according to the five-tuple, and the information belonging to the data flow is counted and identified by machine learning, and finally the identification results are stored in the database. The specific implementation method is as follows:
[0107] Step A. Refer to the attached Figure 3, users connect to the wireless access point AP (Access Point) through mobile devices and use SSR proxy for data transmission. Data is transmitted through the gateway and sent to the smart device through the AP. The generated data will also be sent by the AP to the gateway and then to the target server. Among them, a server is deployed at the gateway, which can mirror and copy all the traffic of the campus gateway, so all the traffic can be captured in the mirror gateway, which is the deployment location of this system.
[0108] Step B. Use the shell program to specify the traffic source network card in the server. After the program is started, it will capture the network card data in real time. The duration of a single traffic capture is estimated based on the network card traffic scale and the system memory size (usually the default setting is 30 seconds). When a single data capture is completed, the data is saved as a pcap file and resides in the memory (it has not been synchronously refreshed to the hard disk at this time). The program starts the asynchronous recognition module group to identify the data file. When the asynchronous task submission is completed, data capture can continue, such as Figure 3 As shown, the design goal of the real-time system is achieved through asynchronous calling.
[0109] Step C. The shell program monitors and manages the submitted asynchronous tasks and the pcap data files that have not been processed. When the estimated time and the occupied memory exceed the expected set range (that is, it may affect the integrity of the real-time captured data), the program will adjust through certain strategies (delaying the start of the next stage of data capture, forcibly stopping the asynchronous task with the longest estimated remaining time) to ensure that the program runs in a reasonable state.
[0110] Step D. The data packet analysis module extracts the data packets in the memory, disassembles the IP, port, data packet timestamp, data packet size and part of the payload information, and after disassembly, the memory occupied by the data packet can be released for use by the data capture module.
[0111] Step E: The data packet analysis module diverts the data, arranges the IP information of the data packet in lexicographic order, calculates the md5 of the five-tuple to identify the data flow information, and then summarizes the data flow information (forward, reverse, and bidirectional).
[0112] Step F. Calculate the data stream feature sequence, and calculate each piece of data according to the algorithm and sequence of the recognition model feature group, including time statistical features, time distribution features, length statistical features, length distribution features, traffic behavior features, etc., and then input them into the random forest model for recognition.
[0113] Step G: Store the recognition results in the database, divide the storage into different databases and tables according to the recognition date, and maintain the data time range through a time sliding window.
[0114] Step H. Read data from the database, count the real-time quantity of SSR traffic, calculate device scores, query device details, calculate confidence, etc., and output visual charts on the web page.
[0115] The present invention provides a method for identifying SSR traffic based on machine learning and a system that can use the above method for identification in a real network environment. All network data in the corresponding network is obtained by capturing the data of the network card, parsing the data information therein, and then diverting it according to the five-tuple of the data packet, calculating each data, including time statistical characteristics, time distribution characteristics, length statistical characteristics, length distribution characteristics, traffic behavior characteristics, etc., and then using machine learning to count the information belonging to the data flow, identify the SSR traffic therein, and associate it with user information. Not only does it ensure a high SSR identification accuracy, but also by optimizing the calculation process in the system, it can achieve real-time collection and identification under a larger-scale gateway, effectively improving the network security department's ability to supervise network traffic.
[0116] Parts of the present invention that are not described in detail belong to common knowledge among those skilled in the art.
[0117] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. It is obvious that for professionals in this field, after understanding the content and principles of the present invention, they may make various modifications and changes in form and details without departing from the principles and structures of the present invention. However, these modifications and changes based on the ideas of the present invention are still within the scope of protection of the claims of the present invention.
Claims
1. A method for traffic identification using an SSR traffic identification system based on machine learning, It is characterized in that The steps include: (1) Capture data traffic based on the traffic arrival of the device network card: (1.1) Estimate the gateway traffic scale, set the single capture order of magnitude and initial queuing time based on the evaluation results, and ensure that the single round of data capture time is within 30-45 seconds; (1.2) Design a real-time system redundancy mechanism, that is, set a dynamic waiting time, which is calculated in real time based on the system's internal memory usage ratio, processor computing workload, and the number of capture file queues; (1.3) According to the pipeline method, the data packet capture module is called cyclically to obtain network data traffic; (2) extracting basic information of data packets from network data traffic through the data packet processing module to obtain data traffic load information including load characteristics, length, and time; (3) Preprocessing data packets using data traffic load information: (3.1) The data packet analysis module filters the data packets according to the load characteristics of the data flow, filters out the data packets of all protocols except the TCP protocol, retains only the TCP data packets, and filters out the data packets that are retransmitted due to abnormal network connection, and obtains the data packet set R: R={pkg 1 ,pkg 2 ,...,pkg i ,...,pkg r }, Among them, pkg i represents the i-th data packet in the set R, i = 1, 2, ..., r, and r represents the total number of filtered data packets; (3.2) The data packet analysis module groups data packets according to the following rules: (3.2.1) Extract data package pkg i Source IP address src-i , Source Port src-i 、Destination IP address dst-i , Destination Port dst-i and transport layer protocol proto i Five types of information, and form a data packet pkg i Head i : h i =(IP src-i ,Port src-i ,IP dst-i ,Port dst-i ,proto i ), pkg i }{h i ,Len(pkg i ),estimate i }; Among them, Len(pkg i ) indicates data packet pkg i The length of stime i Indicates data packet pkg i Arrival time; (3.2.2) In the data packet set R, for the data packet pkg i The same or opposite packet, its header and pkg i Constitute a packet data stream; (3.2.3) Take i = 1, 2, ..., r and follow steps (3.2.1)-(3.2.2) to obtain the packet data stream corresponding to each data packet in the data packet set R. All packet data streams together constitute the grouped data stream set D, that is, the preprocessed traffic information: D = {flow 1 , flow 2 ,..., flow k ,..., flow d}, Among them, flow k represents the kth packet data flow, k=1,2,...,d, d represents the total number of packet data flows; (4) The data packet identification module extracts features from the packet data streams in the data stream set D and screens them, using machine learning to identify them: (4.1) Statistics of packet data flow k The number of all packets in the flow is recorded as total(flow k ), record all the data packets with the same sending direction as the first data packet as output packets, and the remaining data packets as input packets; (4.2) Calculate flow separately k Statistics of all input packets, all output packets, and all data packet lengths: mean, minimum, maximum, absolute deviation, absolute median, standard deviation, variance, skew, kurtosis, 10%-90% percentiles; (4.3) Compose the statistical values obtained in step (4.2) into flow k to form the statistical feature vector PLS k ; the statistical feature vectors corresponding to all grouped data flows together form the packet length statistical feature matrix PLS; (4.4) Perform forward search and combined feature screening on the features in the packet length statistical feature matrix PLS, divide the features into two categories: positive features and negative features, and perform forward search again until the result is optimal, and obtain the optimized packet length statistical feature matrix PLS'; (4.5) Inputting the matrix PLS' into the model trained based on the random forest algorithm for recognition, obtaining the recognition result, and storing the result in the recognition information storage module; (5) The identification information storage module divides the identification results into two categories: SSR results and all results, and stores them in a specific database mysql with the data stream start time as the index; (6) The recognition result analysis module performs real-time analysis on the recorded information in the MySQL database and outputs the analysis results: (6.1) For the recognition results in the MySQL database over a period of time, statistics are performed and the scores are calculated: Among them, Num ssr Indicates the number of SSR traffic identified, Num all Indicates the total number of data streams, Num dst Indicates the number of communication destination addresses; (6.2) Rank the SSR traffic used by different devices according to the score, and dynamically set different confidence levels to obtain multi-dimensional traffic analysis results for a single user; (6.3) Display the analysis results on the web interface.
2. The method according to claim 1, Features: The 10%-90% percentiles described in step (4.2) are calculated as follows: (4.2.1) Let per% represent any percentile from 10% to 90%; (4.2.2) Packet data flow k total(flow k ) data packet lengths are arranged in ascending order to obtain the length of the sorted data packet; (4.2.3) Select the λth packet length from the sorted packet lengths and calculate per% according to the following formula: in, Represents round up.
3. The method according to claim 1, characterized in that: In step (6.2), the SSR traffic used by different devices is ranked according to the score score, specifically ranked from high to low according to the score, where the higher the score, the greater the probability that the device uses SSR traffic.
4. The method according to claim 1, characterized in that: In step (6.3), the analysis result is displayed on the web interface. For the result identified as SSR traffic, it is dynamically loaded through the visualization chart library Echarts based on JavaScript, and the data is updated in real time to the front-end web page to make the monitoring reach the visualization level.
5. An SSR traffic identification system for implementing the method as claimed in claim 1, characterized in that, comprising: A data acquisition and identification unit, an identification information storage module, and a data analysis and display unit; wherein, the data acquisition and identification unit is composed of a packet capture module, a packet processing module, a packet analysis module, and a packet identification module that are connected unidirectionally in sequence, and the data analysis and display unit is composed of an identification result analysis module and a web interface; the identification information storage module is respectively connected to the data acquisition and identification unit and the data analysis and display unit; The packet capture module is used to obtain network data traffic; The packet processing module is used to extract the basic information of the packet from the network data traffic obtained by the packet capture module; The packet analysis module is used to preprocess the packet according to the basic information obtained by the packet processing module to obtain preprocessed traffic information; The packet identification module is used to identify the preprocessed traffic information obtained by the packet analysis module to obtain an identification result; The identification information storage module is used to store the identification result obtained by the packet identification module in the data acquisition and identification unit and is called by the identification result analysis module in the data analysis and display unit; The identification result analysis module is used to perform real-time analysis on the information stored in the identification information storage module and display the analysis result on the web interface for analysts to query.
6. The system according to claim 5, characterized in that: The basic information of the packet includes at least payload characteristics, length, and time.
7. The system according to claim 5, characterized in that: The packet analysis module preprocesses the packet according to the basic information obtained by the packet processing module, specifically performing traffic grouping and filtering operations; the filtering includes: filtering out packets of all other protocols except the TCP protocol, and filtering out packets retransmitted due to abnormal network connections.
8. The system according to claim 5, characterized in that: The packet identification module identifies the preprocessed traffic information obtained by the packet analysis module, specifically extracting features from the grouped data streams in the preprocessed traffic information, and then using machine learning to complete the identification.
9. The system according to claim 5, Features: The identification information storage module stores the identification results obtained by the data packet identification module in the data acquisition and identification unit, and stores them in the database according to the attribution type of the identification results and the data flow start time as an index.