A traffic identification method and device, a storage medium and a computer program product

By extracting the packet payload and congestion window class feature vectors, and combining passive and active identification methods, the problem of low identification accuracy caused by differences in the design of anonymous proxy protocol obfuscation plugins is solved, achieving higher traffic identification accuracy and identification rate.

CN119788380BActive Publication Date: 2025-11-04CHINA MOBILE (SUZHOU) SOFTWARE TECH CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411944124.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-26
Publication Date
2025-11-04
Estimated Expiration
2044-12-26

AI Technical Summary

Technical Problem

Existing traffic identification methods fail to effectively consider the differences in design and implementation principles of anonymous proxy protocol obfuscation plugins, resulting in insufficient feature dimensions and low identification accuracy.

Method used

Data traffic is identified by extracting feature vectors of packet payload and/or congestion window. By combining passive and active identification methods and utilizing machine learning algorithms and preset anonymous proxy domain information, the feature dimensions are increased to enhance identification accuracy.

Benefits of technology

It improves the accuracy of identifying anonymous proxy traffic, especially in large-scale network environments, reducing the false alarm rate and increasing the identification rate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119788380B_ABST
    Figure CN119788380B_ABST
Patent Text Reader

Abstract

The application provides a traffic identification method and device, a storage medium and a computer program product, which include: performing feature extraction on first data traffic to obtain a first type of feature vector; the first type of feature vector includes a data packet payload type of feature vector and / or a congestion window type of feature vector; and the first type of feature vector is used to identify the first data traffic to obtain a traffic type of the first data traffic; the traffic type includes an anonymous proxy type and / or a general type. The accuracy of traffic identification can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of electronic applications, and in particular to a traffic identification method and device, a storage medium and a computer program product. BACKGROUND

[0002] With the user's attention to personal network information security, anonymous communication system emerges as the times require, and the original intention of anonymous communication technology is to protect the privacy of communication content. It hides the identity information of the communication parties through various ways such as data forwarding, data encryption, data confusion and other technical means, so that the communication content is difficult to be tracked and located. Among them, the anonymous proxy system is a typical application of anonymous communication technology, and its related protocol provides encryption confusion plug-ins, which has the characteristics of high security, strong anonymity and difficult to be monitored, and is usually used to penetrate WAF firewall, bypass IP block and regulatory review to access internal network. It provides convenience for engaging in illegal network activities, endangering the interests of government and citizens. Therefore, how to identify the proxy traffic generated by the anonymous proxy system has become a problem to be solved at present.

[0003] At present, the identification based on machine learning is combined with deep flow detection (Deep / Dynamic Flow Inspection, DFI) to perform. DFI identifies the connection behavior of different applications based on the traffic behavior of data. The traffic behavior of data includes the length statistical characteristics of data packets, time statistical characteristics, or other features from the perspective of data report (such as context traffic information, host behavior information, fingerprint information, etc.). Different types of traffic can be identified based on the above traffic behavior.

[0004] However, the current traffic type identification method only considers the traffic behavior, and does not consider the influence of the differences in the design and implementation principles of the anonymous proxy protocol confusion plug-ins. The feature dimension for traffic identification is small, and the accuracy of traffic identification is low. SUMMARY

[0005] The present application provides a traffic identification method and device, a storage medium and a computer program product. The accuracy of traffic identification can be improved.

[0006] The technical solution of the present application is realized as follows:

[0007] In a first aspect, the present application provides a traffic identification method, which comprises:

[0008] performing feature extraction on the first data traffic to obtain a first type of feature vector; the first type of feature vector includes a data packet payload type of feature vector and / or a congestion window type of feature vector;

[0009] The first data flow is identified by using the first type of feature vector, and a flow type of the first data flow is obtained; the flow type includes an anonymous proxy type and / or a general type.

[0010] In a second aspect, the present application provides a flow identification device, which comprises:

[0011] a feature extraction unit configured to extract features of the first data flow to obtain a first type of feature vector; the first type of feature vector includes a data packet payload type of feature vector and / or a congestion window type of feature vector;

[0012] a flow identification unit configured to identify the first data flow by using the first type of feature vector to obtain a flow type of the first data flow; the flow type includes an anonymous proxy type and / or a general type.

[0013] In a third aspect, the present application provides a flow identification device, which comprises a processor, a memory and a communication bus; the communication bus is configured to realize connection communication between the processor and the memory; the processor realizes the flow identification method when executing the running program stored in the memory.

[0014] In a fourth aspect, the present application provides a storage medium, which stores a computer program; the computer program is executed by a processor to realize the flow identification method.

[0015] In a fifth aspect, the present application provides a computer program product, which comprises a computer program; the computer program is executed by a processor to realize the flow identification method.

[0016] This application provides a traffic identification method and apparatus, storage medium, and computer program product. The method includes: extracting features from a first data traffic to obtain a first type of feature vector; the first type of feature vector includes: a data packet payload type feature vector and / or a congestion window type feature vector; using the first type of feature vector to identify the first data traffic and obtain the traffic type of the first data traffic; the traffic type includes: anonymous proxy type and / or general type. Using the above implementation scheme, a proxy traffic identification method is proposed that addresses the design characteristics of anonymous proxy protocol obfuscation plugins. From the perspective of the congestion window algorithm used in proxy service deployment and the change in payload entropy caused by protocol design, the following two features are proposed: data packet payload type features and congestion window type features; by extracting the data packet payload type feature vector and / or congestion window type feature vector of the first data traffic to identify the first data traffic, the method considers the impact of differences in the design and implementation principles of anonymous proxy protocol obfuscation plugins, increases the feature dimensions for traffic identification, and thus improves the accuracy of traffic identification. Attached Figure Description

[0017] Figure 1 A flow chart of a traffic identification method provided in this application embodiment Figure 1 ;

[0018] Figure 2 A flow chart of a traffic identification method provided in this application embodiment Figure 2 ;

[0019] Figure 3 A schematic diagram illustrating an exemplary active identification method provided in an embodiment of this application;

[0020] Figure 4 A structural composition diagram of an exemplary traffic identification device provided in an embodiment of this application;

[0021] Figure 5 A schematic flowchart illustrating an exemplary traffic identification method provided in this application embodiment;

[0022] Figure 6 A schematic diagram of the structure of a flow identification device provided in this application embodiment. Figure 1 ;

[0023] Figure 7 A schematic diagram of the structure of a flow identification device provided in this application embodiment. Figure 2 . Detailed Implementation

[0024] In order to enable more detailed understanding of the features and technical contents of the embodiments of the present application, the implementation of the embodiments of the present application is described in detail below, and the attached drawings are used for reference only and are not intended to limit the embodiments of the present application.

[0025] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the present application belongs. The terms used herein are only for the purpose of describing the embodiments of the present application and are not intended to limit the present application.

[0026] In the following description, "some embodiments" are described, which describe a subset of all possible embodiments, but it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict. It should be noted that the terms "first, second, third" involved in the embodiments of the present application are only used to distinguish similar objects, and do not represent a specific order of the objects. It can be understood that "first, second, third" can be interchanged in a specific order or sequence as allowed, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.

[0027] The embodiments of the present application provide a traffic identification method, as shown in the method can include: Figure 1

[0028] S101, feature extraction is performed on the first data traffic to obtain a first type of feature vector; the first type of feature vector includes: a data packet payload type of feature vector and / or a congestion window type of feature vector.

[0029] The traffic identification method proposed in the embodiments of the present application is applied to the scene of identifying anonymous proxy traffic designed for ShadowSocksR proxy software.

[0030] In the present application, the first type of feature vector can include at least one of the following: a time sequence statistical feature vector, a length statistical feature vector, a data packet payload type of feature vector, a congestion window type of feature vector, etc. The specific selection can be made according to the actual situation, and the embodiments of the present application are not limited.

[0031] S102, using the first type of feature vector to identify the first data traffic, obtaining the traffic type of the first data traffic; the traffic type includes: anonymous proxy type and / or general type.

[0032] ​In the embodiment of the present application, the first type of feature vector is input into the first model for passive flow identification to obtain the flow type of the first data flow; if the flow type of the first data flow is a general type, the second model and / or preset anonymous proxy domain name information are used to perform active flow identification on the first data flow to determine whether to adjust the flow type of the first data flow from the general type to an anonymous proxy type.

[0033] In the embodiment of the present application, the first model is a mixed type model.

[0034] It should be noted that the basic four types of data flows are divided according to different obfuscation methods used according to the ShadowSocksR protocol, including pseudo Hypertext Transfer Protocol (HTTP) type flow http , pseudo Transport Layer Security (TLS) type flow tls , plain (unmasked) type flow plain , and mixed type flow mix . The model used to identify the mixed type flow mix is a mixed type model.

[0035] In the embodiment of the present application, first, the first type of feature vector is input into the first model for passive flow identification to obtain the passive identification result of the first data flow, i.e., whether the flow type of the first data flow is a general type or an anonymous proxy type.

[0036] It should be noted that if the passive identification result is that the flow type of the first data flow is a general type, the second model and / or preset anonymous proxy domain name information are used to determine whether the flow type of the first data flow is indeed a general type.

[0037] It should be noted that the first type of feature vector can be input into the second model for active flow identification, and / or the preset anonymous proxy domain name information can be used to perform active flow identification on the first type of feature vector. The specific selection can be made according to actual conditions, and the embodiment of the present application does not make specific limitations.

[0038] It can be understood that a distributed prototype system is designed by using the combination of passive identification and active identification, the passive identification result is further optimized, the high ShadowSocksR identification rate and the low false positive rate are ensured, the system has high usability in a large-scale network environment and a scenario with a particularly large amount of data, and the system is easy to deploy. The identification accuracy can be further improved, and the identification result is closer to the actual situation in the real scenario.

[0039] In an embodiment, the process of determining whether to adjust the traffic type of the first data traffic from the general type to the anonymous proxy type by actively identifying the first data traffic by using the second model comprises: inputting the first data traffic into the second model to obtain a second traffic type; the second traffic type comprises at least one of the following: a pseudo HTTP type, a pseudo TLS type, a plain type, and a general type; if the second traffic type is at least one of the pseudo HTTP type, the pseudo TLS type, and the plain type, it is determined to adjust the traffic type of the first data traffic from the general type to the anonymous proxy type; and if the second traffic type is the general type, it is determined not to adjust the traffic type of the first data traffic.

[0040] In an embodiment of the present application, the second model comprises at least one of the following models: a plain type model, an HTTP protocol type model, and a TLS protocol type model; wherein the plain type model is used to identify plain type data traffic and general type data traffic, the HTTP protocol type model is used to identify pseudo HTTP type data traffic and general type data traffic, and the TLS protocol type model is used to identify pseudo TLS type data traffic and general type data traffic. In particular, the selection can be made according to actual conditions, and the embodiment of the present application does not make specific limitations.

[0041] It should be noted that the obfuscation method of the plain type is to not perform obfuscation on the traffic, and to directly send data packets using the encrypted results of the protocol.

[0042] It should be noted that the obfuscation method of the pseudo HTTP type is to obfuscate and disguise as an HTTP protocol. The plug-in can customize almost complete HTTP headers and fill them in the obfuscation parameters. This obfuscation is not to reduce features, but rather to provide a strong feature to try to deceive the protocol detection of the firewall.

[0043] It should be noted that the obfuscation method of the pseudo TLS type is to obfuscate and disguise as a TLS protocol, and to simulate tls1.2 to avoid sending complex certificates and other steps when the client has a session ticket, so that the firewall cannot determine through the certificate. In addition, the protocol also has certain anti-replay attack ability and packet length obfuscation ability. However, due to the additional handshake process, the connection time will be slightly longer than the original protocol.

[0044] It can be understood that after being deployed in front of the cloud database MySQL gateway, the HTTP and TLS traffic obfuscated and disguised by the anonymous proxy system can be accurately identified, intercepted, and traced in a timely manner, and a blacklist can be set for filtering.

[0045] In another embodiment, the process of actively identifying the first data flow by using preset anonymous proxy domain name information to determine whether to adjust the traffic type of the first data flow from the general class to the anonymous proxy class includes: determining the first domain name information corresponding to the first data flow; if the first domain name information is found in the preset anonymous proxy domain name information, it is determined that the traffic type of the first data flow is adjusted from the general class to the anonymous proxy class; if the first domain name information is not found in the preset anonymous proxy domain name information, it is determined that the traffic type of the first data flow is not adjusted.

[0046] It should be noted that the common ShadowSocksR service provider domain name library is collected, and the service provider commonly used domain name or some keywords are crawled to establish the preset anonymous proxy domain name information.

[0047] It should be noted that the ShadowSocksR proxy service provider divides the domain name configured by SSR-Server into the above six categories. The network type, network channel (purpose), region, airport abbreviation, number, and other service provider custom class are divided. Through this rule, the abbreviations of each region, service provider airport name, network type, network channel type, etc. are collected to establish the preset anonymous proxy domain name information, and are periodically updated and supplemented. The network type, network channel, and region can match any number and order, followed by a number or no number to form a third-level domain name, an abbreviation or other types to form a second-level domain name. At the same time, it is observed that the top-level domain name commonly used by the proxy service provider is usually a non-traditional domain name such as xyz, dog, cat, gay, etc.

[0048] For example, the storage form of the preset anonymous proxy domain name information can be: [network type][network channel][region]+[number].[abbreviation][other].[top-level domain name]. Specific selection can be made according to actual conditions, and the embodiments of the present application are not limited specifically.

[0049] It should be noted that for the first data flow, the IP is reverse checked for the domain name, and the valid domain name still in use for nearly half a year is retained.

[0050] For an unknown domain name, the preset anonymous proxy domain name information is searched through public substring matching and regular expression matching. If the spliced domain name is in the same style as the preset anonymous proxy domain name information, it can be determined that it is an anonymous proxy class data flow generated by the ShadowSocksR proxy node.

[0051] It can be understood that active identification has less restrictions than passive identification. First, the user is not required to connect with the server. Second, it is not necessary to find a packet with proxy traffic characteristics in a bunch of packets. Active identification attempts to build a packet similar to the authentication packet structure of the proxy traffic and sends it to the proxy software server. Machine learning is then used to find the response packet rules of the proxy software server. This mainly takes advantage of the vulnerabilities of related proxy systems, and some proxy protocols do not have the structure to check the encryption.

[0052] It should be noted that before using the first model and the second model, the first model and the second model are first subjected to model training and model verification. For details, see Figure 2 .

[0053] S201, obtaining a data flow sample; the data flow sample includes an anonymous proxy class data flow sample and a general class data flow sample.

[0054] In the embodiments of the present application, the data flow sample can be collected at the network exit. The anonymous proxy class data flow sample is from a built ShadowSocksR server and a plurality of different ShadowSocksR proxy service providers commonly available on the market.

[0055] S202, dividing the data flow sample according to a first obfuscation manner and a second obfuscation manner to obtain a first training set, a second training set and a test set.

[0056] In the embodiments of the present application, the first obfuscation manner is a mixed class, and the second obfuscation manner includes at least one of the following: a pseudo-HTTP class obfuscation manner, a pseudo-TLS class obfuscation manner and a plain class obfuscation manner. The specific selection can be made according to the actual situation, and the embodiments of the present application do not make specific limitations.

[0057] In an embodiment, the anonymous proxy type data flow samples and the general type data flow samples are divided into 4 groups. First, part of the samples collected in a time period in the anonymous proxy type data flow samples are selected as the to-be-trained positive samples, and are filtered according to the configured confusion mode, and three confusion types are divided, including using the http_simple mode (labeled as "http positive sample 1"), using the tls1.2_ticket_auth mode (labeled as "tls positive sample 2"), and using the plain mode (labeled as "plain positive sample 3"). Then, part of the samples collected in a time period in the general type data flow samples are selected as the to-be-trained negative samples, and half of them are selected as the negative samples corresponding to the mixed type (labeled as "mixed negative sample 4"), and the remaining to-be-trained negative samples are divided according to the proportion of the three different confusion types of the to-be-trained positive samples (labeled as "http negative sample 1", "tls negative sample 2", and "plain negative sample 3"), and then combined with the to-be-trained positive samples of the corresponding type. Finally, for the mixed type, the total amount of the corresponding to-be-trained positive samples of the three confusion types is taken half (labeled as "mixed positive sample 4") and combined with the "mixed negative sample 4" labeled before. The specific division set is shown in Table 1.

[0058] Sample name Combination way Pseudo http class {http-positive sample 1, http-negative sample 1} Pseudo tls class {tls-positive sample 2, tls-negative sample 2} Plain class {plain-positive sample 3, plain-negative sample 3} Mixed class {mixed-positive sample 4, http-negative sample 4}

[0059] wherein (http positive sample 1):(tls positive sample 2):(plain positive sample 3)=x:y:z, x, y, z satisfy formula (1).

[0060] x + y + z = 1 (1)

[0061] After obtaining the above 4 group sets, the training-test set is divided. It is divided into 5 groups G1, G2, G3, G4, and G5, and the following scheme is adopted, that is, in G1, the pseudo tls type is the training set, and the pseudo http type is the test set; in G2, the pseudo http type is the training set, and the pseudo tls type is the test set; in G3, the plain type is the training set, and the mixed type is the test set; in G4, the pseudo tls type and the pseudo http type are the training set, and the mixed type is the test set; and in G5, the training set and the test set are both the mixed type.

[0062] In the embodiment of the present application, the anonymous proxy type data flow samples and the general type data flow samples are divided into five-tuple data flows flow src-i according to source IP addresses IP dst-i , destination IP addresses IP src-i , source ports Port dst-i , destination ports Port i , and used protocols Proto i .

[0063] S203, training a first model using the first training set and training a second model using the second training set.

[0064] In the embodiments of the present application, the data flow samples in the first training set and the second training set are feature extracted to obtain a time series statistical feature vector TCV i , a length statistical feature vector LCV i , a data packet payload class feature vector ECV i , and a congestion window special class feature vector WCV i , and the feature vectors of each adjacent flow are spliced into TM i . Then, the above feature vectors of all flows are combined to obtain a feature matrix TM, which is input into the selected machine learning algorithm to obtain a classifier model.

[0065] In the embodiments of the present application, the random forest algorithm is used by the algorithm analysis module to perform binary classification training on positive samples (anonymous proxy class data flow samples) and negative samples (general class data flow samples), and to retain the ShadowSocksR multiple obfuscation mode training model. Specifically, the following steps are included:

[0066] First, the positive and negative sample detection feature data set is input into the random forest algorithm binary classification model of the algorithm analysis module; the model training parameters are set, including the model division standard, the maximum tree depth d, the minimum number of samples S required for splitting the internal node, the minimum number of samples L on the leaf node, the number of trees e in the forest, and the number of cross-validation k.

[0067] Then, the positive samples are marked as True, and the negative samples are marked as False; the sample feature data set is randomly shuffled, and then divided into a sample detection feature training set with N train =N*p feature data, and a sample detection feature test set with N test =N*(1-p) feature data, 0<p<1, indicating the training set ratio.

[0068] After that, the random forest algorithm model is trained using the sample detection feature training set, and the training set N train is randomly selected with replacement at least S samples, m features are selected from M feature sets, and a decision tree is established for the selected samples; the e decision trees with maximum depth d and minimum leaf node number L are generated e times to form a random forest recognition model.

[0069] Then, four types of sample recognition models are trained, including a first model and three second models, which are: a mixed class model, a plain class model, an HTTP protocol class model, and a TLS protocol class model.

[0070] Finally, the model performance test is performed. Specifically, a probability threshold t is set, the samples in the test set are respectively input into the four types of sample recognition models, and whether it is a positive sample is determined by using the threshold t. If the probability value output by the model is greater than the probability threshold t, the sample is determined as a positive sample. The black and white sets are defined, wherein the white set is the set of true positive samples, i.e., the ShadowSocksR flow set, and the black set is the set of true negative samples, i.e., the normal general flow. The false positive rate, accuracy, precision, recall rate and F1 value of the model are calculated as the recognition evaluation criteria. The recognition results are stored in the distributed recognition model storage module for calling by the recognition result analysis module in the data analysis and display unit.

[0071] Based on the above embodiment, the process of the passive recognition algorithm is shown in Table 2.

[0072] Table 2

[0073]

[0074] S204, the first model and the second model are verified by using the test set and the preset anonymous proxy domain name information.

[0075] In the embodiment of the present application, the process of verifying the first model and the second model by using the test set and the preset anonymous proxy domain name information includes: using the first model to divide the test set into a first flow set and a second flow set; and performing a first scoring on the second flow set to obtain first scoring data; the flow type of the first flow set is general type, and the flow type of the second flow set is anonymous proxy type; using the second model to re-divide the data flow in the first flow set to obtain an updated first flow set and an updated second flow set; and performing a second scoring on the updated second flow set to obtain second scoring data; using the preset anonymous proxy domain name information to re-divide the data flow in the updated first flow set to obtain a second updated first flow set and a second updated second flow set; and performing a third scoring on the second updated second flow set to obtain third scoring data; and verifying the first model and the second model according to the first scoring data, the second scoring data and the third scoring data.

[0076] In the embodiment of the present application, under the premise of meeting the generalization and relative accuracy of the algorithm to the four types of models, the new sample detection feature set is collected for effect recognition, and the passive recognition result is Accuracy mixThe new samples are verified by using the hybrid class model, a threshold value η is set for the first time for the output result, the correctness of the first identification result can be initially believed, the samples in the white set are named as white set 1 (the first flow set), and the samples in the black set are named as black set 1 (the second flow set). Scoring is performed for all white set 1 samples, and the score is w1; scoring is performed for all black set 1 samples, and the score is v1.

[0077] Let η = Accuracy mix * 100, then w1 = η, v1 = (1 - Accuracy mix ) * 100.

[0078] In the embodiment of the application, the black set 1 output above is subjected to secondary model performance test by using the remaining three-class contrast model (the second model), secondary scoring is performed, and the white set is expanded. Specifically, the black set 1 output above is subjected to secondary verification by using the remaining three-class identification model, and secondary scoring is performed; the samples that are collectively judged as ShadowSocksR flow samples are scored, and the samples with a score higher than the threshold value η are classified into the white set 2 (the first flow set updated once), and the remaining samples are classified into the black set 2 (the second flow set updated once); this step considers that if the samples in the black set 1 are judged as non-ShadowSocksR flow samples by the hybrid class model, under the premise of trusting the identification result, most of the samples must be normal flow, but it cannot be excluded that the samples are identified as ShadowSocksR flow samples by other three-class models, after all, there are differences between ShadowSocksR protocols and confusion, and if the samples in the black set 1 are not sensitive to the secondary screening of the other three-class models, the final score will not exceed the threshold value η, and then the samples will not be classified into the white set 2; the black set 2 is retained. The calculation process of the second scoring data w2 of the white set 2 is shown in formula (2).

[0079] w2 = v1 + max (Accuracy http , Accuracy tls , Accuracy origin ) * 100 (2)

[0080] In the embodiment of the present application, the ShadowSocksR field name keyword pair black set is filtered three times, and the third score is obtained. First, a common ShadowSocksR service provider domain name library is collected, and the service provider commonly used domain name or some keywords are crawled to establish a keyword library; then, the non-target network export IP in the black set 2 is extracted to obtain a to-be-detected set SSRIPset; the IP address in the SSRIPset is domain name reverse lookup, and the valid domain name still in use in the past half year is reserved for the third score. If the domain name is included in the keyword library, the corresponding sample in the SSRIPset is added to the white set 3 (the first traffic set of the second update) and supplemented into the white set 1. The score of all samples in the white set 3 is w3; the remaining samples are added to the black set 3 (the second traffic set of the second update). The calculation process of the third score data w3 of the white set 3 is shown in formula (3).

[0081] w3 = v1 + w2 * 100 (3)

[0082] In the embodiment of the present application, the third score result is counted. Specifically, the score of all samples in the current white set 1 is greater than or equal to the threshold η, and the score of all samples in the black set 3 is less than the threshold η; with the change of the scale of new samples and the proportion of black and white samples, the values of w2 and η will change; with the expansion of the keyword library, the value of w3 will change, that is, the credibility weight changes.

[0083] For example, the active identification process is shown in Figure 3 The second score of the black set 1 is determined, and it is determined whether the second score is higher than the threshold η. If yes, it is divided into the white set; if no, it is divided into the black set 2; it is determined whether the black set 2 is in the keyword library. If yes, it is divided into the white set; if no, it is divided into the black set 3.

[0084] It should be noted that the first model and the second model proposed in the embodiment of the present application are traffic identification models with the above scoring rules.

[0085] It can be understood that a proxy traffic identification method for the design characteristics of the anonymous proxy protocol obfuscation plug-in is proposed. From the perspective of the congestion window algorithm used by the proxy service deployment and the change of the payload entropy caused by the protocol design, the following two features are proposed: packet payload feature and congestion window feature. The first data traffic is identified by extracting the packet payload feature vector and / or the congestion window feature vector of the first data traffic. The influence of the difference between the design and the implementation principle of the anonymous proxy protocol obfuscation plug-in is considered. The feature dimension for traffic identification is increased, and the accuracy of traffic identification is improved.

[0086] Based on the above embodiment, a traffic identification device is proposed, which is shown inFigure 4 The traffic identification device comprises a data collection and identification unit, a distributed identification model storage module, an active identification unit and a data analysis and display unit. The data collection and identification unit is composed of a data packet capture collection module, a data packet feature module and an algorithm analysis module; the active identification unit is composed of a scoring module and a keyword library; the distributed identification model storage module is composed of an Api Server, an MQ and a DataServer; and the data analysis and display unit is composed of an identification result analysis module and a Kibana display web page using ElasticSearch.

[0087] The data packet capture collection module is used to acquire network data traffic and divide it into positive and negative sample sets, wherein the positive sample set is ShadowSocksR proxy traffic and the negative sample set is normal general traffic; the data packet feature module is used to extract basic information of data packets from the two types of network data traffic acquired by the data packet capture collection module and perform data packet preprocessing to obtain preprocessed traffic information; the algorithm analysis module is used to perform machine learning algorithm identification according to the preprocessed traffic information of the data packet feature module to obtain a passive identification result; the distributed identification model storage module is used to store the identification result obtained by the algorithm analysis module in the data collection and identification unit in a distributed manner to improve the availability of the entire system and provide the identification result analysis module in the data analysis and display unit with a calling function; the scoring module is used for active secondary screening and scoring; the keyword library is a common ShadowSocksR service provider domain name library established by crawling commonly used domain names of service providers or some keywords; more than 80 ShadowSocksR proxy service providers of different sizes are collected and some services are purchased to obtain nearly 1000 available target nodes. It is observed that service providers will name nodes according to the region where VPS is deployed, and the third-level domain name usually contains the English abbreviation of the region, such as hk, tw, jp and uk; the network channel name is, for example, game for games and iepl for special lines; and the abbreviation or abbreviation of the service provider's own name is added in the second-level domain name. The identification result analysis module is used to perform real-time analysis on the information stored in the distributed identification model storage module and input to the web page for display.

[0088] Based on the above traffic identification device, a traffic identification method is proposed, as shown in Figure 5 .

[0089] 1. Collecting mobile phone traffic at the network exit and constructing positive and negative sample data sets.

[0090] 2. Constructing five-tuple data flow and dividing it into four categories according to the confusion method.

[0091] 3. Constructing data packet features.

[0092] 4. Using a machine learning algorithm model to perform traffic identification.

[0093] 5. Using the identification result to perform performance testing and divide the black and white sets.

[0094] 6. Using the black set to re-divide the black and white sets, and performing secondary scoring on the re-divided black set.

[0095] 7. Re-dividing the black and white sets according to a keyword library, and performing tertiary scoring on the re-divided black set.

[0096] 8. Counting the final identification result.

[0097] 9. Storing the final identification result in a distributed identification model storage module.

[0098] Embodiments of the present application provide a traffic identification device 1. As shown in the figure, the traffic identification device 1 comprises: Figure 6

[0099] a feature extraction unit 10 configured to perform feature extraction on first data traffic to obtain a first type of feature vector; the first type of feature vector comprises a data packet payload type of feature vector and / or a congestion window type of feature vector;

[0100] a traffic identification unit 11 configured to perform traffic identification on the first data traffic using the first type of feature vector to obtain a traffic type of the first data traffic; the traffic type comprises an anonymous proxy type and / or a general type.

[0101] Optionally, the device further comprises a passive identification unit and an active identification unit.

[0102] The passive identification unit is configured to input the first type of feature vector into a first model to perform passive traffic identification, and obtain the traffic type of the first data traffic.

[0103] The active identification unit is configured to, if the traffic type of the first data traffic is the general type, perform active traffic identification on the first data traffic using a second model and / or preset anonymous proxy domain name information, and determine whether to adjust the traffic type of the first data traffic from the general type to the anonymous proxy type.

[0104] ​Optionally, the active identification unit is further configured to input the first data flow into the second model to obtain a second flow type; the second flow type includes at least one of the following: a pseudo HTTP type, a pseudo TLS type, an unmasked type, and a general type; if the second flow type is at least one of the pseudo HTTP type, the pseudo TLS type, and the unmasked type, it is determined that the flow type of the first data flow is adjusted from the general type to the anonymous proxy type; if the second flow type is the general type, it is determined that the flow type of the first data flow is not adjusted.

[0105] Optionally, the active identification unit is further configured to determine first domain name information corresponding to the first data flow; if the first domain name information is found in the preset anonymous proxy domain name information, it is determined that the flow type of the first data flow is adjusted from the general type to the anonymous proxy type; if the first domain name information is not found in the preset anonymous proxy domain name information, it is determined that the flow type of the first data flow is not adjusted.

[0106] Optionally, the device further includes an acquisition unit, a division unit, a training unit, and a verification unit.

[0107] The acquisition unit is configured to acquire data flow samples; the data flow samples include anonymous proxy type data flow samples and general type data flow samples.

[0108] The division unit is configured to divide the data flow samples according to a first confusion manner and a second confusion manner to obtain a first training set, a second training set, and a test set.

[0109] The training unit is configured to respectively train the first model by using the first training set and train the second model by using the second training set.

[0110] The verification unit is configured to perform model verification on the first model and the second model by using the test set and the preset anonymous proxy domain name information.

[0111] Optionally, the verification unit is further configured to divide the test set into a first traffic set and a second traffic set by using the first model, and perform first scoring on the second traffic set to obtain first scoring data; the traffic type of the first traffic set is a general type, and the traffic type of the second traffic set is an anonymous proxy type; perform re-division on the data traffic in the first traffic set by using the second model to obtain an updated first traffic set and an updated second traffic set; perform second scoring on the updated second traffic set to obtain second scoring data; perform re-division on the data traffic in the updated first traffic set by using the preset anonymous proxy domain name information to obtain a second updated first traffic set and a second updated second traffic set; perform third scoring on the second updated second traffic set to obtain third scoring data; and complete model verification on the first model and the second model according to the first scoring data, the second scoring data, and the third scoring data.

[0112] Optionally, the first model is a mixed model; and the second model includes at least one of the following models: an uncamouflaged model, an HTTP protocol model, and a TLS protocol model; the uncamouflaged model is used to identify uncamouflaged data traffic and general data traffic, the HTTP protocol model is used to identify pseudo HTTP data traffic and general data traffic, and the TLS protocol model is used to identify pseudo TLS data traffic and general data traffic.

[0113] The traffic identification device provided in the embodiment of the application performs feature extraction on the first data traffic to obtain a first feature vector; the first feature vector includes a data packet payload feature vector and / or a congestion window feature vector; the first data traffic is identified by using the first feature vector to obtain the traffic type of the first data traffic; and the traffic type includes an anonymous proxy type and / or a general type. It can be seen that the traffic identification device provided in the embodiment proposes a proxy traffic identification method for the design features of an anonymous proxy protocol obfuscation plug-in, proposes two features, a data packet payload feature and a congestion window feature, from the perspective of the changes in payload entropy caused by the congestion window algorithm and protocol design used by the proxy service deployment, identifies the first data traffic by extracting the data packet payload feature vector and / or the congestion window feature vector of the first data traffic, considers the influence of the differences in the design and implementation principles of the anonymous proxy protocol obfuscation plug-in, increases the feature dimension for traffic identification, and thus improves the accuracy of traffic identification.

[0114] Figure 7 The composition structure of the traffic identification device 1 provided in the embodiment of the application is shown in the figure Figure 2In actual applications, based on the same disclosure concept of the above embodiments, as shown in the figure, the flow recognition device 1 of the present embodiment comprises a processor 12, a memory 13 and a communication bus 14. Figure 7

[0115] The processor 12 can be at least one of an application specific integrated circuit (ASIC), a digital signal processor (DSP), a digital signal processing device (DSPD), a programmable logic device (PLD), a field programmable gate array (FPGA), a CPU, a controller, a microcontroller, and a microprocessor. It can be understood that for different devices, the electronic device for realizing the function of the processor can also be other, and the present embodiment is not limited specifically.

[0116] In the present embodiment, the communication bus 14 is used to realize the connection and communication between the processor 12 and the memory 13; and the processor 12 realizes the following flow recognition method when executing the running program stored in the memory 13.

[0117] The first data flow is subjected to feature extraction to obtain a first type of feature vector; the first type of feature vector includes a data packet payload type feature vector and / or a congestion window type feature vector; the first data flow is subjected to flow recognition by using the first type of feature vector to obtain the flow type of the first data flow; and the flow type includes an anonymous proxy type and / or a general type.

[0118] Further, the processor 12 is further used to input the first type of feature vector into a first model to perform passive flow recognition to obtain the flow type of the first data flow; if the flow type of the first data flow is the general type, a second model and / or preset anonymous proxy domain name information are used to perform active flow recognition on the first data flow to determine whether to adjust the flow type of the first data flow from the general type to the anonymous proxy type.

[0119] ​Further, the processor 12 is further configured to input the first data flow into the second model to obtain a second flow type, the second flow type including at least one of a pseudo HTTP type, a pseudo TLS type, an unmasked type, and a general type; if the second flow type is at least one of the pseudo HTTP type, the pseudo TLS type, and the unmasked type, determining to adjust the flow type of the first data flow from the general type to the anonymous proxy type; and if the second flow type is the general type, determining not to adjust the flow type of the first data flow.

[0120] Further, the processor 12 is further configured to determine first domain name information corresponding to the first data flow; if the first domain name information is found in the preset anonymous proxy domain name information, determining to adjust the flow type of the first data flow from the general type to the anonymous proxy type; and if the first domain name information is not found in the preset anonymous proxy domain name information, determining not to adjust the flow type of the first data flow.

[0121] Further, the processor 12 is further configured to obtain data flow samples, the data flow samples including anonymous proxy type data flow samples and general type data flow samples; divide the data flow samples according to a first confusion manner and a second confusion manner to obtain a first training set, a second training set, and a test set; train the first model by using the first training set and train the second model by using the second training set; and perform model verification on the first model and the second model by using the test set and the preset anonymous proxy domain name information.

[0122] Further, the processor 12 is further configured to divide the test set into a first flow set and a second flow set by using the first model, and perform first scoring on the second flow set to obtain first scoring data; the flow type of the first flow set is the general type, and the flow type of the second flow set is the anonymous proxy type; redivide data flow in the first flow set by using the second model to obtain an updated first flow set and an updated second flow set, perform second scoring on the updated second flow set to obtain second scoring data; redivide data flow in the updated first flow set by using the preset anonymous proxy domain name information to obtain a second updated first flow set and a second updated second flow set, and perform third scoring on the second updated second flow set to obtain third scoring data; and complete model verification on the first model and the second model according to the first scoring data, the second scoring data, and the third scoring data.

[0123] Further, the first model is a mixed class model; the second model includes at least one of the following models: a plain class model, an HTTP protocol class model, and a TLS protocol class model; the plain class model is used to identify plain data traffic and general data traffic, the HTTP protocol class model is used to identify pseudo HTTP data traffic and general data traffic, and the TLS protocol model is used to identify pseudo TLS data traffic and general data traffic.

[0124] The embodiment of the present application provides a storage medium, which stores a computer program; the computer readable storage medium stores one or more programs; the one or more programs can be executed by one or more processors, and are applied to a traffic identification device; and the computer program implements the traffic identification method.

[0125] Based on the above embodiment, the embodiment of the present application provides a computer program product, which comprises a computer program; the computer program can be executed by one or more processors; and the computer program implements the traffic identification method.

[0126] It should be noted that, in this document, the term "comprising" or "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that processes, methods, articles or devices including a series of elements not only include those elements, but also include other elements not explicitly listed, or further include elements inherent to such processes, methods, articles or devices. Without more limitations, the element defined by the statement "comprising a" does not exclude the presence of additional identical elements in the process, method, article or device including the element.

[0127] From the above description of the embodiments, those skilled in the art can clearly understand that the above-mentioned embodiment method can be realized by means of software and necessary general hardware platform, of course, it can also be realized by hardware, but in many cases, the former is a better embodiment. Based on such understanding, the technical solutions of the present disclosure can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a plurality of instructions for making an image display device (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) execute the methods described in various embodiments of the present disclosure.

[0128] The above description is only a preferred embodiment of the present application, and is not intended to limit the protection scope of the present application.

Claims

1. A traffic identification method, characterized by, The method comprises: characteristic extraction is performed on the first data flow to obtain a first type of characteristic vector; the first type of characteristic vector comprises a data packet payload type of characteristic vector and / or a congestion window type of characteristic vector; flow identification is performed on the first data flow by using the first type of characteristic vector to obtain a flow type of the first data flow; the flow type comprises an anonymous proxy type and / or a general type; wherein the flow identification performed on the first data flow by using the first type of characteristic vector to obtain the flow type of the first data flow comprises: the first type of characteristic vector is input into a first model for passive flow identification to obtain the flow type of the first data flow; if the flow type of the first data flow is the general type, active flow identification is performed on the first data flow by using a second model and / or preset anonymous proxy domain name information to determine whether to adjust the flow type of the first data flow from the general type to the anonymous proxy type; the method further comprises: obtaining data flow samples; the data flow samples comprise anonymous proxy type data flow samples and general type data flow samples; the data flow samples are divided according to a first confusion manner and a second confusion manner to obtain a first training set, a second training set and a test set; the first model is trained by using the first training set, and the second model is trained by using the second training set; the first model and the second model are verified by using the test set and the preset anonymous proxy domain name information.

2. The method of claim 1, wherein, the active flow identification performed on the first data flow by using the second model to determine whether to adjust the flow type of the first data flow from the general type to the anonymous proxy type comprises: the first data flow is input into the second model to obtain a second flow type; the second flow type comprises at least one of a pseudo hypertext transfer protocol (HTTP) type, a pseudo transport layer security (TLS) type, an unmasked type and a general type; if the second flow type is at least one of the pseudo HTTP type, the pseudo TLS type and the unmasked type, it is determined that the flow type of the first data flow is adjusted from the general type to the anonymous proxy type; if the second flow type is the general type, it is determined that the flow type of the first data flow is not adjusted.

3. The method of claim 1, wherein, the active flow identification performed on the first data flow by using the preset anonymous proxy domain name information to determine whether to adjust the flow type of the first data flow from the general type to the anonymous proxy type comprises: first domain name information corresponding to the first data flow is determined; if the first domain name information is found in the preset anonymous proxy domain name information, it is determined that the flow type of the first data flow is adjusted from the general type to the anonymous proxy type; if the first domain name information is not found in the preset anonymous proxy domain name information, it is determined that the flow type of the first data flow is not adjusted.

4. The method of claim 1, wherein, the model verification performed on the first model and the second model by using the test set and the preset anonymous proxy domain name information comprises: The first model is used to divide the test set into a first traffic set and a second traffic set, and to score the second traffic set for the first time to obtain first score data; the traffic type of the first traffic set is a general type, and the traffic type of the second traffic set is an anonymous proxy type; The second model is used to redivide the data traffic in the first traffic set to obtain an updated first traffic set and an updated second traffic set, and to score the updated second traffic set for the second time to obtain second score data; The preset anonymous proxy domain name information is used to redivide the data traffic in the updated first traffic set to obtain a second updated first traffic set and a second updated second traffic set, and to score the second updated second traffic set for the third time to obtain third score data; The first model and the second model are verified according to the first score data, the second score data, and the third score data.

5. The method of claim 1, wherein, The first model is a mixed model; the second model includes at least one of the following models: an unmasked model, an HTTP protocol model, and a TLS protocol model; the unmasked model is used to identify unmasked data traffic and general data traffic, the HTTP protocol model is used to identify pseudo HTTP data traffic and general data traffic, and the TLS protocol model is used to identify pseudo TLS data traffic and general data traffic.

6. A traffic identification device, characterized by The traffic identification device includes: A feature extraction unit is configured to extract features from first data traffic to obtain a first feature vector; the first feature vector includes a data packet payload feature vector and / or a congestion window feature vector; A traffic identification unit is configured to identify the traffic type of the first data traffic by using the first feature vector; the traffic type includes an anonymous proxy type and / or a general type; A passive identification unit is configured to input the first feature vector into a first model to perform passive traffic identification and obtain the traffic type of the first data traffic; An active identification unit is configured to perform active traffic identification on the first data traffic by using a second model and / or preset anonymous proxy domain name information, and determine whether to adjust the traffic type of the first data traffic from the general type to the anonymous proxy type; A obtaining unit is configured to obtain data traffic samples; the data traffic samples include anonymous proxy type data traffic samples and general type data traffic samples; A dividing unit is configured to divide the data traffic samples according to a first confusion method and a second confusion method to obtain a first training set, a second training set, and a test set; A training unit is configured to train the first model by using the first training set and train the second model by using the second training set. A verification unit is configured to perform model verification on the first model and the second model by using the test set and the preset anonymous proxy domain name information.

7. A traffic identification device, characterized by The traffic identification device comprises a processor, a memory and a communication bus, the communication bus is configured to realize connection communication between the processor and the memory, and the processor is configured to execute a running program stored in the memory to realize the method according to any one of claims 1-5.

8. A storage medium having stored thereon a computer program, characterized in that The computer program, when executed by a processor, realizes the method according to any one of claims 1-5.

9. A computer program product comprising a computer program, characterized in that, The computer program, when executed by a processor, realizes the method according to any one of claims 1-5. The computer program, when executed by a processor, realizes the method according to any one of claims 1-5.

Citation Information

Patent Citations

  • Anonymous protocol classification method based on statistical feature classification

    CN106330611A

  • Anonymous service flow association identification method and system nested in encryption tunnel

    CN111224940A