Intranet illegal external connection detection method based on traffic load frequent items

By building a machine learning model based on frequent item features, the problems of insufficient detection coverage and easy bypass of existing methods for detecting illegal external connections within the intranet are solved, and the seamless detection and efficient identification of encrypted traffic are achieved.

CN120768636APending Publication Date: 2025-10-10SOUTHEAST UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511005154.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-21
Publication Date
2025-10-10

AI Technical Summary

Technical Problem

Existing methods for detecting illegal external connections within the intranet rely on specific protocols, have insufficient detection coverage, are easily bypassed, and are complex to deploy, making them difficult to apply on a large scale in intranet environments with high security requirements.

Method used

By collecting and labeling encrypted protocol traffic samples, extracting frequent item features, and building a machine learning model, we can achieve seamless detection of encrypted traffic.

Benefits of technology

It achieves accurate detection of private encryption protocols, is suitable for complex intranet environments, is easy to deploy, does not affect normal communications, and has strong applicability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120768636A_ABST
    Figure CN120768636A_ABST
Patent Text Reader

Abstract

The invention discloses a traffic load frequent item-based intranet illegal external connection detection method, which comprises two stages of off-line training and on-line detection, in the off-line training stage, a large number of public encryption protocol traffic samples and private encryption protocol traffic samples are collected and marked, and the public encryption protocol traffic samples and the private encryption protocol traffic samples are marked; performing flow grouping on the data packet according to the quintuple information of the network flow to form a data flow session; then, extracting a'frequent item 'in the load of the aggregation flow; then, by counting the occurrence frequency of the frequent item in the load, the load length, the ratio of the occurrence frequency to the load length and other characteristics, a characteristic vector is constructed; and finally, training by using a specific machine learning method to obtain an illegal external connection detection model. In the online detection stage, flow grouping and feature extraction are carried out on real-time network flow, an obtained feature vector is input into a pre-trained detection model, and the flow with a prediction result being a private encryption protocol is marked as an illegal external connection behavior. According to the method, the identification of the private encryption protocol in the intranet environment can be completed, so that the illegal external connection behavior can be effectively detected.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application relates to a method for detecting illegal external connection of an internal network based on frequent items of traffic load, and belongs to the technical field of cyberspace security. BACKGROUND

[0002] In the current digital era, the internal networks of governments, enterprises and large organizations bear a large amount of core businesses and sensitive data. In order to protect information security and prevent data leakage, a core security principle is "internal-external separation and boundary controllability", that is, strict access control and security isolation are performed on the communication between the internal network and the external Internet. However, with the development of network technology and the evolution of attack means, the internal host may illegally bypass the security policy and establish a connection with the external Internet through an unauthorized channel, that is, illegal external connection. Such behavior not only may lead to leakage of core data, but also may cause the internal host to be implanted with a malicious remote control program, thus constituting a major security risk, and has become a very threatening and very hidden security problem.

[0003] At present, there are invention patents for detecting illegal external connection of an internal network. The existing invention patent "Method for detecting existence of illegal external connection channel of terminal based on middleware injection mode" injects detection code in the middleware of an internal network Web server. When a user accesses a specific internal website, the server will dynamically insert a piece of JavaScript code in the returned page, and the code functions to let the user's browser try to access a public network monitoring server. If the monitoring server records the access, it proves that the user terminal has external connection capability. However, the limitation of this method is very obvious, it can only cover the user traffic accessing specific internal Web applications through browsers, and cannot detect a large number of non-browser applications (such as desktop clients, background services) and other protocol traffic. At the same time, injecting code into normal web pages may cause compatibility problems or affect user experience. In addition, "A method for detecting illegal external connection based on proxy configuration file" configures DHCP / DNS services in the internal network, and issues a specific proxy configuration file (wpad.dat) containing a unique host domain name for each internal network device. When the device connects to the Internet, it will automatically try to resolve the unique domain name. Once the DNS server deployed on the public network receives the resolution request of the domain name, it can confirm that the device has illegally connected to the external network. The effect of this method depends on the support and configuration of the specific protocol (such as WPAD) by the client operating system, and modern operating systems have gradually limited or closed such automatic discovery functions for security reasons. In addition, experienced users or malicious software can easily bypass this proxy-based detection mechanism by modifying system settings.

[0004] In summary, although existing research on detecting illegal external connections within the intranet has achieved certain results, it generally has one or more of the following defects: (1) It relies on specific protocols or applications, resulting in insufficient detection coverage and being unable to cope with non-standard and non-browser traffic; (2) It is easy for users to bypass it, whether by modifying the client configuration or circumventing specific applications. Experienced users can easily avoid detection; (3) It is complex or invasive to deploy, making it difficult to apply on a large scale in complex, heterogeneous, and security-intensive intranet environments.

[0005] The method for detecting illegal external connections in an intranet based on frequent items of traffic load proposed in the present invention is a new detection technology that is simple to deploy, does not rely on specific protocols, and is not perceived by users. Summary of the Invention

[0006] To solve the above problems, the present invention discloses a method for detecting illegal external connections in the intranet based on frequent items of traffic load. The method proposed by the present invention can complete the detection of illegal external connections in the intranet. The method of the present invention includes two stages: offline training and online detection. In the offline training stage, a large number of public encryption protocol traffic (such as SSL / TLS, SSH) samples and private encryption protocol traffic (such as Vmess, Shadowsocks) samples are collected and labeled, and the data packets are grouped according to the five-tuple information of the network traffic to form a data flow session; then, the "frequent items" (such as "AB", "AXB", "AXXB") in the payload (Payload) of the aggregated flow are extracted; then, by counting the number of occurrences of the "frequent items" in the payload, the payload length, the ratio of the number of occurrences and the payload length, etc., a feature vector is constructed; finally, a specific machine learning method is used for training to obtain an illegal external connection detection model. In the online detection stage, the real-time network traffic is grouped and feature extracted, and the obtained feature vector is input into the pre-trained detection model, and the traffic predicted to be a private encryption protocol is marked as illegal external connection behavior. The present invention can effectively detect private encryption protocols without decrypting traffic, and is of great value in discovering illegal external connection behaviors such as data leakage and malicious remote control in the intranet. It is suitable for intranet monitoring scenarios with high security and confidentiality requirements.

[0007] To achieve the purpose of the present invention, the specific technical steps of this solution are as follows: A method for detecting illegal external connections in an intranet based on frequent items of traffic load, the method comprising the following steps:

[0008] Step (1) Build a virtual environment to automatically and efficiently collect sample data of public encryption protocol traffic (such as SSL / TLS, SSH) and private encryption protocol traffic (such as Vmess, Shadowsocks), and add labels to form a training dataset. The private encryption traffic is considered as potential illegal external contact behavior;

[0009] Step (2) aggregating the traffic data in the training data set obtained in step (1) based on its five-tuple information to form a complete session aggregation flow;

[0010] Step (3) extracts the “frequent items” in the aggregated flow load obtained in step (2) and constructs the corresponding feature vector;

[0011] Step (4) inputs the feature vector constructed in step (3) into a specific machine learning algorithm for training to generate an illegal outbound connection detection model;

[0012] Step (5) deploys the illegal external connection detection model trained in step (4) on the intranet monitoring node;

[0013] Step (6) performs the same aggregation and feature extraction operations as steps (2) and (3) on the real-time traffic in the intranet, and inputs the generated feature vector into the deployed illegal outbound connection detection model to detect illegal outbound connection behavior.

[0014] Furthermore, in step (1), the steps of constructing the data set used for training are as follows:

[0015] (1.1) Build a virtual environment and simulate the client through automated scripts to communicate with public encryption protocols such as HTTPS and SSH, and private encryption protocols such as VMess and Shadowsocks proxy services, and capture and store the generated network data packets;

[0016] (1.2) According to the known types of communication targets, the traffic data captured in step (1.1) are associated and labels are added to the traffic data in step (1.1) to form two types of training sets: "normal encrypted traffic" and "illegal external traffic".

[0017] Furthermore, in step (2), in the process of grouping the traffic data of the training set, the network data packets collected in step (1) need to be combined into traffic according to the five-tuple (srcIP, srcPort, dstIP, dstPort, Protocol) information to form a complete session traffic;

[0018] Furthermore, in step (3), the steps of constructing the feature vector used for training are as follows:

[0019] (3.1) This method finds that the public encryption protocol and the private encryption protocol can be distinguished by the occurrence of the “frequent items” in the encrypted traffic payload. The frequent item structure is defined as (A, B, X1, ..., X k ,A,B), this structure specifically refers to two adjacent and repeated byte pairs (i.e., A, B), separated by k arbitrary bytes;

[0020] (3.2) For each aggregated flow's payload data, the present invention uses a sliding window scanning algorithm to extract and calculate the number of "frequent items" in step (3.1). The specific process of the algorithm is as follows: First, a window of M bytes is used to slide byte by byte from the beginning of the payload to the end to extract the frequent items in the current window; then, a hash table is used to record the frequent items as indexes and calculate the number of times they appear in the entire payload. After the scan is completed, the hash table stores all the frequent items that have appeared in the payload and their corresponding count values;

[0021] (3.3) Convert the relevant features of the frequent items into feature vectors for machine learning. For example, the total number of occurrences of the frequent items, the total length of the payload of the aggregated stream, and the ratio of the total number of occurrences of the frequent items to the total length of the payload of the aggregated stream.

[0022] Furthermore, in step (4), the steps of performing the machine learning training process of the feature vector are as follows:

[0023] (4.1) Divide the feature vector dataset obtained after processing in step (3) into a training set and a test set according to a predetermined ratio (e.g., 8:2). The training set is used for model learning, and the test set is used to evaluate the final performance of the model.

[0024] (4.2) During the machine learning training of feature data, considering the algorithm's processing speed, classification accuracy, and the inherent noise in network traffic, the random forest algorithm, as an ensemble learning method, is insensitive to outliers and noise and is very suitable for this application scenario. Therefore, the random forest algorithm is selected to train the training set in step (4.1);

[0025] (4.3) Using the Bootstrap sampling method (random sampling with replacement), create multiple different training subsets from the original training set in step (4.1). Based on each training subset, independently construct a decision tree. Finally, combine the generated multiple decision trees into a random forest model.

[0026] (4.4) Use the test set to evaluate the performance of the models under different hyperparameters, and select the model with the best overall performance in terms of accuracy, recall rate, and F1 score as the final illegal outbound connection detection model.

[0027] Furthermore, in the step (5), when deploying the violation detection model, it is necessary to comprehensively consider the management complexity and deployment cost, and select the mirror port of the egress gateway or the intranet core switch to deploy the violation external connection detection model trained in step (4), and deploy it on the key monitoring nodes of the intranet, which will greatly reduce the management complexity and deployment cost, and achieve excellent results;

[0028] Furthermore, in step (6), the steps for detecting illegal external connection behavior of intranet traffic are as follows:

[0029] (6.1) The traffic captured from the egress gateway or the intranet core switch is grouped and frequent items are calculated according to the process of steps (2) and (3);

[0030] (6.2) Inputting the feature vector extracted in step (6.1) into the illegal external connection detection model trained in step (4) to obtain the prediction result;

[0031] (6.3) When the prediction result obtained in step (6.2) is "illegal external connection", a security alert is generated and the specific five-tuple information is recorded to facilitate subsequent tracing and inspection by management personnel.

[0032] Compared with the prior art, the technical solution of the present invention has the following beneficial technical effects.

[0033] (1) The present invention proposes a method for detecting illegal external connections in the intranet based on frequent items of traffic load. This method analyzes the side channel characteristics in the encrypted traffic load and trains a machine learning intelligent detection model, so that it can accurately and efficiently identify the covert channels established using private encryption protocols, providing a reliable technical basis for network security managers to discover and deal with illegal external connection behaviors.

[0034] (2) This invention addresses the problem that existing detection methods cannot effectively identify private encryption protocol traffic without fixed features. It innovatively proposes a frequent item characterization method. This method utilizes the difference in the statistical randomness of payload byte sequences between standard encryption protocols (such as SSL / TLS) and private encryption protocols (such as VMess), providing a new and effective technical path to distinguish between these two types of traffic without decrypting the traffic.

[0035] (3.) Compared to intrusive detection methods that require client installation or network configuration modification, the proposed method has the advantages of being user-insensitive, simple to deploy, and highly versatile. This method can be deployed on mirrored traffic at network edges or core switches without affecting normal network communications. Its core algorithm involves only payload scanning and counting, resulting in low computational overhead and suitability for high-traffic environments. It also does not rely on specific ports or IP addresses, effectively countering various types of protocol spoofing and possessing strong applicability in complex, real-world network environments. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] Figure 1 This is a system framework diagram of the intranet illegal external connection detection method based on frequent items of traffic load. DETAILED DESCRIPTION

[0037] The technical solutions provided by the present invention will be described in detail below with reference to specific embodiments. It should be understood that the following specific embodiments are only used to illustrate the present invention and are not used to limit the scope of the present invention.

[0038] Example: The present invention provides a method for detecting illegal external connections in an intranet based on frequent items of traffic load, and its overall system structure is as follows: Figure 1 As shown, the following steps are included:

[0039] Step (1) Build a virtual environment to automatically and efficiently collect sample data of public encryption protocol traffic (such as SSL / TLS, SSH) and private encryption protocol traffic (such as Vmess, Shadowsocks), and add labels to form a training dataset. The private encryption traffic is considered as potential illegal external contact behavior;

[0040] In one embodiment of the present invention, the steps of constructing a data set for training are as follows:

[0041] (1.1) Build a virtual environment and simulate the client through automated scripts to communicate with public encryption protocols such as HTTPS and SSH, and private encryption protocols such as VMess and Shadowsocks proxy services, and capture and store the generated network data packets;

[0042] (1.2) Based on the known type of the communication target, the traffic data captured in step (1.1) is correlated and labeled to form two types of training sets: "normal encrypted traffic" and "illegal external traffic". For example, when collecting Vmess traffic, build the server yourself, make sure the configured uuid is consistent with the client, turn off other interfering traffic, and ensure the traffic is as pure as possible. After the collection is completed, classify it as illegal external traffic;

[0043] Step (2) aggregating the traffic data in the training data set obtained in step (1) based on its five-tuple information to form a complete session aggregation flow;

[0044] In one embodiment of the present invention, when grouping the traffic data of the training set, the network data packets collected in step (1) need to be grouped according to the five-tuple (srcIP, srcPort, dstIP, dstPort, Protocol) information to form a complete session traffic. For example, a specific five-tuple information is (192.168.1.108,7890,192.168.1.109,16823,TCP)

[0045] Step (3) extracts the “frequent items” in the aggregated flow load obtained in step (2) and constructs the corresponding feature vector;

[0046] In one embodiment of the present invention, the steps of constructing a feature vector for training are as follows:

[0047] (3.1) The present invention finds that the public encryption protocol and the private encryption protocol can be distinguished by the occurrence of the "frequent items" in the encrypted traffic load. The frequent item structure is defined as (A, B, X1, ..., X k ,A,B), this structure specifically refers to two adjacent and repeated byte pairs (i.e., A, B), separated by k arbitrary bytes. For example, this method sets the value of the interval distance k to {0,1,2}. This set can effectively capture the frequent item features, thus serving as the key basis for distinguishing public encryption protocols from private encryption protocols;

[0048] (3.2) For each aggregated flow payload, the present invention uses a sliding window scanning algorithm to extract and calculate the number of "frequent items" in step (3.1). The specific process of the algorithm is as follows: First, a window of M (for example, when k = 2, M = 4) bytes is used to slide from the beginning of the payload to the end, extracting the frequent items in the current window (for example, extracting the frequent item (P) at position i). i ,P i+k+1 )); Then, a hash table is used to record frequent items as indexes and count the number of times they appear in the entire payload. After the scan is completed, the hash table stores all the frequent items that have appeared in the payload and their corresponding count values;

[0049] (3.3) Convert the relevant features of frequent items into feature vectors for machine learning. For example, the total number of occurrences of frequent items, the total length of the payload of the aggregated stream, and the ratio of the total number of occurrences of frequent items to the total length of the payload of the aggregated stream. Table 1 lists some feature names and their meanings.

[0050] Table 1: Some feature names and their meanings

[0051]

[0052] Step (4) inputs the feature vector constructed in step (3) into a specific machine learning algorithm for training to generate an illegal outbound connection detection model;

[0053] In one embodiment of the present invention, the steps of performing a machine learning training process for a feature vector are as follows:

[0054] (4.1) Divide the feature vector dataset obtained after processing in step (3) into a training set and a test set according to a predetermined ratio (e.g., 8:2). The training set is used for model learning, and the test set is used to evaluate the final performance of the model.

[0055] (4.2) During the machine learning training of feature data, considering the algorithm's processing speed, classification accuracy, and the inherent noise in network traffic, the random forest algorithm, as an ensemble learning method, is insensitive to outliers and noise and is very suitable for this application scenario. Therefore, the random forest algorithm is selected to train the training set in step (4.1);

[0056] (4.3) Using the Bootstrap sampling method (random sampling with replacement), create multiple different training subsets from the original training set in step (4.1). Based on each training subset, independently construct a decision tree. Finally, combine the generated multiple decision trees into a random forest model.

[0057] (4.4) Use the test set to evaluate the performance of the models under different hyperparameters. Select the model with the best overall performance in terms of accuracy, recall, and F1 score as the final illegal outbound connection detection model. For example, the number of decision trees (n_estimators) in this method is set to 150 and the maximum depth of the decision tree (max_depth) is set to 5;

[0058] Step (5) deploys the illegal external connection detection model trained in step (4) on the intranet monitoring node;

[0059] In one embodiment of the present invention, in the process of deploying the violation detection model, it is necessary to comprehensively consider the management complexity and deployment cost, and select the mirror port deployment step (4) of the egress gateway or the intranet core switch to train the violation external connection detection model, and deploy it at the key monitoring node of the intranet, which will greatly reduce the management complexity and deployment cost, and the effect is excellent.

[0060] Step (6) performs the same aggregation and feature extraction operations as steps (2) and (3) on the real-time traffic in the intranet, and inputs the generated feature vector into the deployed illegal outbound connection detection model to detect illegal outbound connection behavior.

[0061] In one embodiment of the present invention, the steps for detecting illegal external connection behavior of intranet traffic are as follows:

[0062] (6.1) The traffic captured from the egress gateway or the intranet core switch is grouped and frequent items are calculated according to the process of steps (2) and (3).

[0063] (6.2) Inputting the feature vector extracted in step (6.1) into the illegal external connection detection model trained in step (4) to obtain the prediction result;

[0064] (6.3) When the prediction result obtained in step (6.2) is "illegal external connection", a security alert is generated and the specific five-tuple information is recorded to facilitate subsequent tracing and inspection by management personnel.

[0065] The technical means disclosed in the solutions of the present invention are not limited to those disclosed in the above-mentioned embodiments, but also include technical solutions composed of any combination of the above-mentioned technical features. It should be noted that those skilled in the art may make various improvements and modifications without departing from the principles of the present invention, and such improvements and modifications are also considered to be within the scope of protection of the present invention.

Claims

1. A method for detecting illegal external connections in an intranet based on frequent items of traffic load, characterized in that: The method comprises the following steps: Step (1) building a virtual environment, automatically and efficiently collecting public encryption protocol traffic and private encryption protocol traffic sample data, and adding labels to form a training data set, wherein the private encryption traffic is regarded as potential illegal external contact behavior; Step (2) aggregating the traffic data in the training data set obtained in step (1) based on its five-tuple information to form a complete session aggregation flow; Step (3) extracts the "frequent items" in the aggregated flow obtained in step (2) and constructs the corresponding feature vector; Step (4) inputs the feature vector constructed in step (3) into a specific machine learning algorithm for training to generate an illegal outbound connection detection model; Step (5) deploys the illegal external connection detection model trained in step (4) on the intranet monitoring node; Step (6) performs the same aggregation and feature extraction operations as steps (2) and (3) on the real-time traffic in the intranet, and inputs the generated feature vector into the deployed illegal outbound connection detection model to detect illegal outbound connection behavior.

2. The method for detecting illegal external connections in an intranet based on frequent items of traffic load according to claim 1 is characterized in that: The step (1) specifically includes the following sub-steps: (1.1) Build a virtual environment and simulate the client through automated scripts to communicate with the HTTPS and SSH services of public encryption protocols and the VMess and Shadowsocks proxy services of private encryption protocols, and capture and store the generated network data packets; (1.2) According to the known types of communication targets, the traffic data captured in step (1.1) are associated and labeled to form two types of training sets: "normal encrypted traffic" and "illegal external traffic".

3. The method for detecting illegal external connections in an intranet based on frequent items of traffic load according to claim 1 is characterized in that: In the step (2), the network data packets collected in the step (1) need to be combined into traffic according to the five-tuple (srcIP, srcPort, dstIP, dstPort, Protocol) information to form a complete session traffic.

4. The method for detecting illegal external connections in an intranet based on frequent items of traffic load according to claim 1 is characterized in that: The step (3) specifically includes the following sub-steps: (3.1) This method finds that the public encryption protocol and the private encryption protocol can be distinguished by the appearance of the "frequent items" in the encrypted traffic payload. The frequent item structure is defined as (A, B, X1, ..., X k ,A,B), this structure specifically refers to two adjacent and repeated byte pairs (i.e., A, B), separated by k arbitrary bytes; (3.2) For each aggregated flow's payload data, the present invention uses a sliding window scanning algorithm to extract and calculate the number of "frequent items" in step (3.1). The specific process of the algorithm is as follows: First, a window of M bytes is used to slide byte by byte from the beginning of the payload to the end to extract the frequent items in the current window; then, a hash table is used to record the frequent items as indexes and calculate the number of times they appear in the entire payload. After the scan is completed, the hash table stores all the frequent items that have appeared in the payload and their corresponding count values; (3.3) Convert the relevant features of frequent items into feature vectors for machine learning.

5. The method for detecting illegal external connections in an intranet based on frequent items of traffic load according to claim 1 is characterized in that: The step (4) specifically includes the following sub-steps: (4.1) Divide the feature vector dataset obtained after processing in step (3) into a training set and a test set according to a predetermined ratio. The training set is used for learning the model, and the test set is used to evaluate the final performance of the model. (4.2) During the machine learning training of feature data, considering the algorithm's processing speed, classification accuracy, and the inherent noise in network traffic, the random forest algorithm, as an ensemble learning method, is insensitive to outliers and noise and is very suitable for this application scenario. Therefore, the random forest algorithm is selected to train the training set in step (4.1); (4.3) Using the Bootstrap sampling method (random sampling with replacement), create multiple different training subsets from the original training set in step (4.1). Based on each training subset, independently construct a decision tree. Finally, combine the generated multiple decision trees into a random forest model. (4.4) Use the test set to evaluate the performance of the models under different hyperparameters, and select the model with the best overall performance in terms of accuracy, recall rate, and F1 score as the final illegal outbound connection detection model.

6. The method for detecting illegal external connections in an intranet based on frequent items of traffic load according to claim 1, characterized in that: In the step (5), taking into account the management complexity and deployment cost, the mirror port of the egress gateway or the intranet core switch is selected to deploy the illegal external connection detection model trained in step (4), and deployed on the key monitoring nodes of the intranet, which will greatly reduce the management complexity and deployment cost and achieve excellent results.

7. The method for detecting illegal external connections in an intranet based on frequent items of traffic load according to claim 1, characterized in that: The step (6) specifically includes the following sub-steps: (6.1) The traffic captured from the egress gateway or the intranet core switch is grouped and frequent items are calculated according to the process of steps (2) and (3); (6.2) Inputting the feature vector extracted in step (6.1) into the illegal external connection detection model trained in step (4) to obtain the prediction result; (6.3) When the prediction result obtained in step (6.2) is "illegal external connection", a security alert is generated and the specific five-tuple information is recorded to facilitate subsequent tracing and inspection by management personnel.