A method and system for judging suspicious traffic in encrypted traffic
By collecting and analyzing the characteristic information of encrypted traffic, using machine learning models and relationship maps to identify suspicious traffic, and performing subsequent decryption analysis, the problem of difficult identification of suspicious traffic in encrypted traffic is solved, and network security protection capabilities are improved.
Patent Information
- Application Number
- CN202211070466.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-02
- Publication Date
- 2025-09-02
- Estimated Expiration
- 2042-09-02
AI Technical Summary
The prior art is difficult to effectively identify and distinguish suspicious traffic in encrypted traffic, resulting in malicious traffic being hidden and difficult to detect and protect.
By collecting encrypted traffic, extracting access feature information, protocol feature information and passing feature information, using machine learning models and relationship map analysis, determine traffic types, and conducting subsequent decryption analysis of suspicious traffic.
It improves the accuracy and efficiency of identifying suspicious traffic in encrypted traffic, reduces the workload of subsequent decryption analysis, and enhances network security protection capabilities.
Smart Images

Figure CN115514537B_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the field of network security, and in particular to a method and system for determining suspicious traffic in encrypted traffic. Background Art
[0002] With the development of the internet, people's awareness of privacy is also increasing, resulting in a growing demand for traffic encryption. However, while encrypted traffic protects privacy, it also facilitates the concealment of malicious traffic. Encrypted malicious traffic hides many known and unknown threats. A method for identifying suspicious traffic within encrypted traffic is needed to enhance network security protection capabilities. Summary of the Invention
[0003] One or more embodiments of this specification provide a method for determining suspicious traffic in encrypted traffic. The method includes: collecting encrypted traffic to be tested, extracting encrypted traffic characteristics of the encrypted traffic to be tested; wherein the encrypted traffic characteristics include first traffic characteristics, and the first traffic characteristics include access characteristic information, protocol characteristic information, and transmission characteristic information; based on the encrypted traffic characteristics of the encrypted traffic to be tested, determining the traffic type of the encrypted traffic to be tested, wherein the traffic type includes normal traffic and the suspicious traffic, and the suspicious traffic is used for subsequent decryption analysis of the encrypted traffic to be tested.
[0004] One or more embodiments of the present specification provide a system for determining suspicious traffic in encrypted traffic, the system comprising: a traffic collection module, configured to collect the encrypted traffic to be tested and extract encrypted traffic characteristics of the encrypted traffic to be tested; wherein the encrypted traffic characteristics include a first traffic characteristic, and the first traffic characteristic includes access characteristic information, protocol characteristic information, and transmission characteristic information; a type determination module, configured to determine the traffic type of the encrypted traffic to be tested based on the encrypted traffic characteristics of the encrypted traffic to be tested, the traffic type including normal traffic and the suspicious traffic, and the suspicious traffic is used for subsequent decryption analysis of the encrypted traffic to be tested.
[0005] One or more embodiments of the present specification provide a device for determining suspicious traffic in encrypted traffic, including a processor, wherein the processor is configured to execute at least part of the computer instructions to implement a method for determining suspicious traffic in encrypted traffic.
[0006] One or more embodiments of the present specification provide a computer-readable storage medium that stores computer instructions. When the computer instructions are executed by a processor, a method for determining suspicious traffic in encrypted traffic is implemented. BRIEF DESCRIPTION OF THE DRAWINGS
[0007] This specification will be further described in the form of exemplary embodiments, which will be described in detail with reference to the accompanying drawings. These embodiments are not limiting, and in these embodiments, like numbers represent like structures, wherein:
[0008] Figure 1 This is a schematic diagram of an application scenario of a system for determining suspicious traffic in encrypted traffic according to some embodiments of this specification;
[0009] Figure 2 is an exemplary module diagram of a system for determining suspicious traffic in encrypted traffic according to some embodiments of this specification;
[0010] Figure 3 is an exemplary flow chart of a method for determining suspicious traffic in encrypted traffic according to some embodiments of this specification;
[0011] Figure 4 is an exemplary schematic diagram of obtaining a second flow characteristic through a relationship map according to some embodiments of this specification;
[0012] Figure 5 This is a schematic diagram of a suspicious traffic identification model according to some embodiments of this specification. DETAILED DESCRIPTION
[0013] To more clearly illustrate the technical solutions of the embodiments of this specification, the following briefly describes the drawings required for describing the embodiments. Obviously, the drawings described below are merely examples or embodiments of this specification. Those skilled in the art can apply this specification to other similar scenarios based on these drawings without inventive effort. Unless otherwise apparent from the context or otherwise noted, the same reference numerals in the figures represent the same structure or operation.
[0014] It should be understood that the terms "system," "device," "unit," and / or "module" used herein are a method for distinguishing different components, elements, parts, portions, or assemblies at different levels. However, if other terms can achieve the same purpose, the terms may be replaced by other expressions.
[0015] As used in this specification and claims, unless the context clearly indicates otherwise, the words "a," "an," "an," and / or "the" do not refer to the singular but also include the plural. Generally speaking, the terms "comprises" and "include" only indicate the inclusion of the steps and elements specifically identified, and these steps and elements do not constitute an exclusive list. A method or apparatus may also include other steps or elements.
[0016] Flowcharts are used throughout this specification to illustrate the operations performed by systems according to embodiments of this specification. It should be understood that preceding or following operations do not necessarily need to be performed in exact order. Instead, the steps may be processed in reverse order or simultaneously. Furthermore, other operations may be added to these processes, or one or more operations may be removed from these processes.
[0017] Figure 1 This is a schematic diagram of an application scenario of a system for determining suspicious traffic in encrypted traffic according to some embodiments of this specification.
[0018] Deep Packet Inspection (DPI) involves inspecting and analyzing traffic and packet content at key points in the network. It can filter and control the inspected traffic based on predefined policies. This capability includes refined link service identification, traffic flow analysis, traffic share statistics, traffic shaping, application layer denial of service attacks, virus and Trojan filtering, and P2P abuse control. For example, if the decryption DPI module determines that the encrypted traffic under test is suspicious, it can then perform subsequent decryption analysis on the traffic.
[0019] like Figure 1 As shown, application scenario 100 may include network 110, router 120, processor 130, encrypted traffic 140, and traffic determination result 150. Router 120 may obtain encrypted traffic 140 to be tested from network 110, and processor 130 may copy encrypted traffic 140 in router 120 to collect encrypted traffic 140 and generate traffic determination result 150.
[0020] The network 110 may include any suitable network that provides information and / or data exchange that can facilitate the bandwidth application scenario 100. The router 120 of the application scenario 100 can exchange information and / or data with the network 110. For example, the network 110 can send traffic information generated by a user to the router 120. In some embodiments, the network 110 can be any one or more of a wired network or a wireless network. In some embodiments, the network 110 can include one or more network access points. For example, the network 110 can include a wired or wireless network access point. In some embodiments, the network can be a point-to-point, shared, centralized, or other topological structure, or a combination of multiple topological structures.
[0021] Router 120 can be a network device that reads addresses from data packets and then stores, groups, and forwards them. In some embodiments, router 120 can be used to connect two or more networks 110. In some embodiments, router 120 receives encrypted traffic 140 from network 110 and forwards the encrypted traffic 140 stored in router 120 to processor 130. Router 120 can be local or remote.
[0022] Processor 130 may include a device for executing the method for determining suspicious traffic in encrypted traffic 140. It may process data and / or information obtained from router 120 and, based on the relevant data, execute the method for determining suspicious traffic in encrypted traffic as provided in this specification to generate a traffic determination result 150. For example, processor 130 may determine traffic characteristics based on the encrypted traffic information received by router 120, and then determine whether encrypted traffic 140 is suspicious traffic based on the traffic characteristics, thereby generating traffic determination result 150. In some embodiments, processor 130 may be a single server or a server group. In some embodiments, processor 130 may be integrated into a suspicious traffic determination system (e.g., integrated within router 120). Processor 130 may be local or remote. Processor 130 may be implemented on a cloud platform.
[0023] Traffic can be traffic generated by users while surfing the Internet. In some embodiments, the traffic can be encrypted or unencrypted. Traffic encryption is used to mitigate various eavesdropping and man-in-the-middle attacks, making web pages virtually tamper-proof and ensuring user Internet security. However, some malicious traffic can still be hidden within encrypted traffic 140. In some embodiments, processor 130 determines whether malicious traffic is present. For example, after encrypted traffic 140 containing malicious traffic is transmitted to a router by network 110, processor 130 determines it is suspicious traffic.
[0024] The traffic determination result 150 may include that the encrypted traffic 140 is suspicious traffic or that the encrypted traffic 140 is normal traffic. In some embodiments, the traffic determination result 150 is executed by the processor 130 .
[0025] It should be noted that application scenario 100 is provided for illustrative purposes only and is not intended to limit the scope of this application. A person skilled in the art may make various modifications or variations based on the description of this specification. For example, application scenario 100 may further include an information source. However, these variations and modifications do not deviate from the scope of this application.
[0026] Figure 2 This is a module diagram of a system for determining suspicious traffic in encrypted traffic according to some embodiments of this specification.
[0027] like Figure 2 As shown, in some embodiments, the suspicious traffic judgment system 200 may include a traffic feature acquisition module 210 and a traffic type determination module 220 .
[0028] Traffic feature acquisition module 210 can be used to collect encrypted traffic to be tested and extract encrypted traffic features from the encrypted traffic to be tested. In some embodiments, the encrypted traffic features may include first traffic features, which may include access feature information, protocol feature information, and transmission feature information. For details on traffic feature acquisition, see step 310 and its related description.
[0029] Traffic type determination module 220 can be used to determine the traffic type of the encrypted traffic to be tested based on the encrypted traffic characteristics of the encrypted traffic to be tested. In some embodiments, the traffic type can include normal traffic and suspicious traffic. Suspicious traffic is used for subsequent decryption analysis of the encrypted traffic to be tested. For details on traffic type determination, see step 320 and its related description.
[0030] In some embodiments, the traffic type determination module 220 may further be used to: process the encrypted traffic characteristics of the encrypted traffic to be tested based on the suspicious traffic identification model to determine the traffic type of the encrypted traffic to be tested, wherein the suspicious traffic identification model is a machine learning model. For details about the suspicious traffic identification model, please refer to Figure 5 and its related descriptions.
[0031] In some embodiments, the suspicious traffic determination system 200 may further include a decryption DPI module 230 for performing subsequent decryption analysis on the encrypted traffic to be tested in response to the traffic type of the encrypted traffic to be tested being suspicious traffic.
[0032] It should be noted that the above description of the suspicious traffic judgment system 200 and its modules is for convenience only and does not limit this specification to the scope of the embodiments. It is understandable that those skilled in the art, after understanding the principles of the system, may arbitrarily combine the modules or form subsystems connected to other modules without deviating from the principles. In some embodiments, Figure 2 The traffic feature acquisition module 210 and traffic type determination module 220 disclosed in the disclosure may be different modules within a system, or a single module may implement the functions of two or more of the aforementioned modules. For example, the modules may share a storage module, or each module may have its own storage module. Such variations are within the scope of protection of this specification.
[0033] Figure 3This is an exemplary flow chart of a method for determining suspicious traffic in encrypted traffic according to some embodiments of this specification. In some embodiments, process 300 may be executed by processor 130. Figure 3 As shown, the process 300 includes the following steps.
[0034] Step 310: Collect the encrypted traffic to be tested and extract the encrypted traffic features of the encrypted traffic to be tested.
[0035] Encrypted traffic refers to network traffic that has been encrypted. This can be done by the user or by the service provider to protect privacy. For example, users who conduct online business transactions over the internet can rely on encryption mechanisms in mobile, cloud, and web applications. Through data encryption, keys, and certificates are used to ensure security and establish trust.
[0036] The basic process of data encryption involves algorithmically processing originally plaintext files or data (traffic), transforming it into unreadable code, commonly referred to as "ciphertext." This data encryption approach protects data from unauthorized theft and reading. In some embodiments, encrypted traffic can include both normal and suspicious traffic. Encrypted suspicious traffic often disguises or conceals malicious traffic characteristics. For example, encrypted suspicious traffic often disguises or conceals malicious Trojan horses, infectious viruses, worms, malicious downloaders, and other malicious programs, potentially attacking servers and causing them to crash.
[0037] Encrypted traffic characteristics are traffic characteristics related to the encrypted traffic to be measured. Traffic characteristics may include statistical information such as five-tuple information, encryption protocol information, average packet size, and average packet sending interval. The five-tuple information includes the source IP address, source port, destination IP address, destination port, and transport layer protocol. Encrypted protocol information refers to protocol-related messages for secure communication established between the server and client during the authentication process. The authentication process includes: the client sends a message to the server; the server responds to the client with a self-authentication message; the client and server complete a key exchange, ending the authentication process. In some embodiments, the encryption protocol information may include the TLS / SSL protocol version, extension fields, etc. The average packet size refers to the average length of the data in several packets, expressed in bytes. For example, the average length of ten IP packets is 1000 bytes. The average packet sending interval refers to the average time interval between the transmission of the current data frame and the transmission of the next data frame during the data transmission process, for example, a data frame is sent every 2 seconds on average.
[0038] In some embodiments, the encrypted traffic feature may include a first traffic feature, which is a feature related to the content of the encrypted traffic to be tested. In some embodiments, the first traffic feature may include access feature information, protocol feature information, and transmission feature information.
[0039] Access characteristic information refers to characteristic information related to access. Access can refer to the process in which a visitor actively searches for a specific purpose using a network platform, and traffic is generated during the access process. For example, encrypted traffic can be traffic generated by a visitor clicking on a website URL saved in a bookmark or traffic generated by a visitor directly entering a URL in the browser address bar. Access characteristic information can be used to distinguish different sessions, such as communications between different users. In some embodiments, access characteristic information may include information such as source IP address, source port, destination IP address, and destination port.
[0040] You can determine whether encrypted traffic is suspicious based on access characteristics. For example, if the number of visits from a browser's IP address increases by 500% in a single day, you might want to check whether suspicious traffic is causing the increase.
[0041] Protocol characteristic information refers to characteristic information related to a protocol. Protocol characteristic information can be used to distinguish the transmission method of network traffic, for example, whether it is encrypted transmission, etc. In some embodiments, protocol characteristic information may include transmission protocol information, encryption protocol information, etc.
[0042] The probability that encrypted traffic is suspicious can be determined based on protocol characteristics. For example, historically, malicious traffic often uses hidden encrypted transmission protocols. By identifying the encryption protocol used by encrypted traffic, such as Secure Socket Layer (SSL), it can be determined that the encryption protocol is more vulnerable to malicious traffic attacks.
[0043] The transmission characteristic information refers to characteristic information related to information transmission. In some embodiments, the transmission characteristic information may include an average size of a data packet, an average interval between sending data packets, and the like.
[0044] Suspicious encrypted traffic can be determined based on transmission characteristics. For example, if a network platform's normal access time is 8:00-18:00, with an average packet size of 512 bytes and an average packet transmission interval of 20ms, but a large number of intensive accesses occur between 02:00 and 04:00 on a given day, with an average packet size of 1500 bytes and an abnormally low average packet transmission interval of 6ms, this abnormal traffic can be identified as suspicious.
[0045] In some embodiments, the first traffic feature may further include a byte distribution probability vector.
[0046] In computer security, data is transmitted from a sender to a receiver in the form of packets. A packet contains a header. The data sent by the sender is called the payload. The receiver can determine the size of the packet's payload by subtracting the IP header length from the total length of the IP packet. The header is appended to the payload for transmission and then discarded upon successful arrival at the destination. Malicious traffic primarily spreads viruses through the payload. Payloads include corrupted data, messages containing insulting text, or bulk emails sent to a large number of people. Byte distribution refers to the count of each byte value in a packet's payload. For example, the byte distribution of a packet might be: in the packet's payload, the first byte "00000001" appears 10 times, the second byte "00000011" appears 15 times, and so on, and the Nth byte "111111111" appears 5 times. Byte distribution probability refers to the probability of occurrence of each byte value in the packet's payload. In some embodiments, the probability of occurrence of each byte value can be approximated by its frequency of occurrence. A byte distribution probability vector is a vector of the probabilities of each of the 256 possible values a byte can assume in a data stream. The byte distribution probability vector can provide a large amount of information about data encoding and data padding. Many illegal behaviors of malicious traffic are often hidden in this information. In some embodiments, the byte distribution count of each byte value can be divided by the total number of bytes in the payload to obtain the byte distribution frequency, which is used to represent the byte distribution probability. Finally, this feature is represented as a 1*256-dimensional byte distribution probability vector. For example, malicious traffic may use certain fields in the HTTP header (e.g., content-type, server, etc.) to initiate some malicious activities, which shows that HTTP fields can well indicate some malicious activities. HTTP context flow refers to all HTTP flows sent from the same source IP address within a 5-minute window of the secure transmission protocol TLS (Transport Layer Security). A feature vector of a binary variable is used to represent all observed HTTP header information. If any HTTP flow has a specific header value (i.e., a header containing malicious traffic), the feature will be 1 regardless of other HTTP flows. For the byte distribution probability vector P1, the processor 130 can count 100 flows of P1 in the network traffic within a preset time period, among which 60 HTTP header features are 1, which means that these 60 are malicious flows. The byte distribution probability vector of the encrypted traffic to be tested is P1, and the frequency of the flow being malicious traffic is 60%.
[0047] You can use various traffic collection methods to collect encrypted traffic to be tested, including but not limited to sniffer, SNMP (Simple Network Management Protocol), NetFlow, and sFlow.
[0048] In some embodiments, a sniffer can be used to collect encrypted traffic. As an example, a data collection point can be set on the mirror port of the switch to completely copy the data information in the network through the mirror port to collect the encrypted traffic information to be tested.
[0049] After collecting the encrypted traffic to be tested, extract its encrypted traffic features. This can be done through various methods, such as using an encrypted traffic basic information extraction library (e.g., Flowcontainer), encrypted traffic feature extraction tools (e.g., WireShark, QPA, Tstat), or other encrypted traffic extraction algorithms or machine learning models.
[0050] In some embodiments, encrypted traffic features may also include a second traffic feature. This second traffic feature is derived from the content of the encrypted traffic being measured. This second traffic feature (domain popularity) can be determined by extracting the destination domain name from the encrypted traffic being measured and then using other external content related to the second traffic feature (e.g., searching the internet or a knowledge graph for the domain name). In some embodiments, the second traffic feature may include domain popularity.
[0051] Domain popularity refers to the degree to which malicious traffic tends to access a domain name. In some embodiments, domain popularity can include the number of times (or frequency, probability, etc.) malicious traffic accesses the domain name. The more frequently a domain name is accessed by malicious traffic, the higher the popularity of the domain name. In some embodiments, if network traffic includes high-popularity domain names, the traffic type determination module 220 can determine that the network traffic is more likely to be suspicious traffic.
[0052] In some embodiments, the traffic feature acquisition module 210 can obtain a second traffic feature of the encrypted traffic to be tested. In some embodiments, the second traffic feature can be expressed as a score value. For example, when the second traffic feature is the popularity of a domain name, the higher the domain name popularity score value, the higher the domain name popularity, and vice versa. The score value can be obtained based on the historical access of the domain name, user reports on the domain name, etc. In some embodiments, the second traffic feature can be used to determine whether the domain name is susceptible to malicious traffic attacks, and further used to determine the traffic type.
[0053] In some embodiments, the second flow characteristic can be obtained through a relationship map.
[0054] Figure 4 This is an exemplary schematic diagram of obtaining the second flow characteristic through a relationship map according to some embodiments of this specification.
[0055] A relationship graph can include domain name nodes, entity nodes, edges connecting entity nodes, and edges connecting entity nodes with domain name nodes. Edge attributes can include communication-related data and traffic types. For example, a domain name node might be Jming.com, and an entity node might be the IP address 207.46.197.101 corresponding to that domain name. An IP address can correspond to multiple domain names, but a domain name only has one IP address. When a user enters a domain name, it first reaches a Domain Name System (DNS) server, which then resolves the domain name to the IP address of the corresponding website. This process is called domain name resolution. Domain name nodes and entity nodes work together to enable client hosts to access servers.
[0056] In some embodiments, the traffic feature acquisition module 210 may construct a relationship graph based on the encrypted traffic information to be measured, and obtain a second traffic feature based on the malicious neighbor value determined by the relationship graph. The malicious neighbor value represents the number of edges of a node (e.g., node A) that meet preset conditions. The preset conditions may include: the direction of the edge points to node A, i.e., the end point is node A. The traffic type of the edge is malicious traffic.
[0057] like Figure 4 As shown, the relationship graph 410 may include domain name nodes 420 (eg, node A, node B, node C), entity nodes 430 (eg, node 1, node 2, node 3), and edges 440 connecting the nodes, wherein the edges are directed edges.
[0058] In some embodiments, the traffic feature acquisition module 210 can construct edges of a relationship graph based on the communication between each node. The communication represented by the edge is an abstract communication. An abstract communication may include multiple information interactions in a short period of time. For example, nodes A and B are connected by a directed edge, which means that there are multiple information interactions between nodes A and B in a short period of time. These multiple information interactions can be regarded as an abstract communication. The direction of the edge can be determined by the initiator of the first information interaction. For example, in the above-mentioned multiple information interactions, the first information interaction is initiated by A to B, then the direction of the multiple information interactions can be correspondingly determined as A pointing to B. In some embodiments, in response to the existence of multiple communications between nodes A and B (for example, communications occurring at different times with a longer time span), nodes A and B may have multiple directed edges. In some embodiments, the processor 130 can count the number of edges between each node whose traffic type is "malicious traffic" based on the attributes of the edges. A malicious neighbor value of 0 indicates that there is no edge with a traffic type of "malicious traffic" between the two nodes; a malicious neighbor value of 1 indicates that there is one edge with a traffic type of malicious traffic between the two nodes; a malicious neighbor value of 2 indicates that there are two edges with a traffic type of malicious traffic between the two nodes, and so on. In some embodiments, the second traffic feature can be determined based on the malicious neighbor value. For more information on determining the second traffic feature based on the malicious neighbor value, please refer to Figure 4 and its related descriptions.
[0059] Schematic process 400 is an example of determining the second traffic feature through a relationship graph. For example, the second traffic feature 460 in the schematic process 400 is domain name popularity. Specifically, the traffic feature acquisition module 210 can find the corresponding node based on the encrypted traffic information to be tested (for example, IP address), and the processor 130 can obtain all edges 440 connected to the node based on the relationship graph 410, and count the number of edges whose traffic type is malicious traffic in the edge attributes corresponding to the node in the graph; determine the malicious neighbor value 450 based on the number of edges; and determine the domain name popularity based on the malicious neighbor value. Figure 4 As shown in the figure, the malicious neighbor value of node 1 and node B is 1, the malicious neighbor value of node 3 is 2, and the malicious neighbor value of node 2 is 3. Therefore, the domain name popularity of node 2 is the highest.
[0060] In some embodiments, the traffic feature acquisition module 210 can determine the popularity of a domain name based on the malicious neighbor value 450 and the preset neighbor rule. Among them, the preset neighbor rule can be to sort the malicious neighbor values corresponding to each node according to size, and the higher the ranking, the higher the domain name popularity. The preset neighbor rule can be set according to actual needs. For example, the domain name corresponding to the top three nodes is output, and the domain name with the highest domain name popularity is output. For domain names with high domain name popularity, the traffic corresponding to the domain name can directly enter the subsequent decryption DPI analysis without traffic type classification.
[0061] In some embodiments, the edge features of the relationship graph may also include the number of times the communication data between the two nodes has been reported by users. Among them, the user terminal is located at a node attacked by malicious traffic. When the communication data between the two nodes is reported by the user, the traffic type corresponding to the communication is recorded as malicious traffic. The processor 130 can count the number of edges with a traffic type of "malicious traffic" between each node. Further, based on the number of edges, a second traffic feature is determined. As an example only, when the second traffic feature is domain name popularity, the more edges with a traffic type of "malicious traffic", the higher the domain name popularity of the domain name corresponding to the node.
[0062] In the embodiments of this specification, by determining the second traffic feature through a relationship graph, the traffic types of network traffic can be effectively integrated, and associations between domain names and domain names, entity IPs and entity IPs, and domain names and entity IPs can be constructed, that is, edges between nodes are constructed according to the direction of traffic flow, thereby more efficiently supporting the mining and extraction of the second traffic feature; by determining the second traffic feature, the accuracy of judging whether the encrypted traffic feature is malicious traffic can be improved.
[0063] Step 320: Determine the traffic type of the encrypted traffic to be measured based on the encrypted traffic characteristics of the encrypted traffic to be measured.
[0064] In some embodiments, the traffic type may include normal traffic and suspicious traffic. The traffic type of the encrypted traffic to be tested may be determined in a variety of ways. In some embodiments, the traffic type may be determined based on historical data, preset rules, or a suspicious traffic identification model. In some embodiments, determining the traffic type based on historical data includes: obtaining historical suspicious traffic through the traffic type determination module 220, and comparing the historical suspicious traffic with the traffic characteristics of the encrypted traffic to be tested, and when the similarity is greater than a certain threshold (for example, greater than 0.8), determining that the traffic type of the encrypted traffic to be tested is suspicious traffic. In some embodiments, determining the traffic type based on preset rules includes determining that the traffic type of the encrypted traffic to be tested is suspicious traffic when the number of suspicious traffic features of the traffic characteristics of the encrypted traffic to be tested is greater than a certain value (for example, greater than 1). In some embodiments, the suspicious traffic identification model may be a machine learning model. For details about the suspicious traffic identification model, please refer to Figure 5 and its related descriptions.
[0065] Step 330: In response to the traffic type of the encrypted traffic to be tested being suspicious traffic, performing subsequent decryption analysis on the encrypted traffic to be tested.
[0066] In some embodiments, if the traffic type of the encrypted traffic to be tested is normal traffic, no subsequent decryption analysis is required.
[0067] Decryption analysis can include confirming the protocol type, splitting the protocol, splitting the protocol domain, SSL unloading, payload analysis, identifying the negotiation protocol, etc. Through decryption analysis of suspicious traffic, it can be further determined whether the suspicious traffic is malicious traffic. The traffic characteristics corresponding to the suspicious traffic can also be marked through the decryption DPI module 230, and the suspicious traffic characteristics are stored in the traffic type determination module 220 for identifying suspicious traffic in the encrypted traffic to be tested. At the same time, it is convenient to obtain more training samples for model training, making the judgment of encryption analysis more accurate.
[0068] The embodiments of this specification screen out normal traffic and suspicious traffic in encrypted traffic, and only perform subsequent decryption analysis on suspicious traffic, thereby reducing the load of subsequent analysis work and improving analysis efficiency.
[0069] Figure 5 This is a schematic diagram of a suspicious traffic identification model according to some embodiments of this specification.
[0070] In some embodiments, based on the encrypted traffic characteristics of the encrypted traffic to be tested, the traffic type of the encrypted traffic to be tested is determined, including: processing the encrypted traffic characteristics of the encrypted traffic to be tested based on a suspicious traffic identification model to determine the traffic type of the encrypted traffic to be tested, and the suspicious traffic identification model is a machine learning model.
[0071] like Figure 5 As shown in , an initial suspicious traffic identification model 550 can be trained based on a large number of labeled training samples 540 to obtain a trained suspicious traffic identification model 520. Specifically, labeled training samples 540 are input into the initial suspicious traffic identification model 550, and the initial suspicious traffic identification model is trained based on the labels. In some embodiments, the training samples 540 can include normal traffic and suspicious traffic.
[0072] In some embodiments, the identification of the training sample may be whether the training sample is suspicious traffic. For example, if the training sample is suspicious traffic, the identification is 1, otherwise it is 0.
[0073] In some embodiments, the initial suspicious traffic identification model 550 can be a binary classifier trained using suspicious traffic as positive samples and normal traffic as negative samples. In some embodiments, the binary classifier can be a logistic regression model, a support vector machine, a random forest, or other classification models.
[0074] In some embodiments, the suspicious traffic identification model 520 can be used to determine the category of traffic corresponding to the input traffic feature. In some embodiments, the input of the suspicious traffic identification model 520 can include the first traffic feature 510-1 and / or the second traffic feature 510-2, and the output of the suspicious traffic identification model 520 can include one of suspicious traffic 530-1 and normal traffic 530-2.
[0075] In some embodiments, training ends when the trained suspicious traffic identification model meets a preset condition. The preset condition may be that the accuracy rate is greater than or equal to a preset threshold. The preset threshold can be set based on actual needs, such as 90% or 95%.
[0076] In some embodiments, the accuracy of a trained suspicious traffic identification model can be determined using multiple test samples, where the test samples contain labels indicating whether the traffic is suspicious. After inputting the multiple test samples into the trained suspicious traffic identification model, the model can output a corresponding predicted category. If the predicted category matches the label, the prediction is correct; otherwise, the prediction is incorrect. The accuracy can be calculated by dividing the number of correctly predicted samples by the total number of test samples.
[0077] The embodiments of this specification use a machine learning model to identify traffic types, and can learn the inherent characteristics of malicious traffic based on a large amount of historical traffic data, thereby more accurately determining whether the encrypted traffic to be tested is suspicious traffic.
[0078] In some embodiments, the output of the suspicious traffic identification model may further include a classification vector 530 - 3 , where the classification vector 530 - 3 includes a confidence level that the encrypted traffic to be tested belongs to different categories of suspicious traffic.
[0079] In some embodiments, before using the suspicious traffic identification model to output a classification vector, the initial suspicious traffic identification model should be trained using a large number of multi-classification training samples to enable it to have a certain multi-classification capability. In some embodiments, the training samples can be normal traffic and different categories of malicious traffic. For example, malicious traffic can belong to "suspicious traffic for privacy leakage", "suspicious traffic for malicious attacks", etc. In some embodiments, the identifier of the training sample can be the category of the training sample. For example, the malicious traffic is identified as A, indicating that the category of the malicious traffic is "suspicious traffic for privacy leakage"; the malicious traffic is identified as B, indicating that the category of the malicious traffic is suspicious traffic for malicious attacks. In some embodiments, the classification vector output by the suspicious traffic identification model can represent the confidence that the suspicious traffic belongs to different malicious behaviors. In some embodiments, the classification vector output by the suspicious traffic model includes multiple numerical values between 0 and 1, which are used to represent the confidence that the sample belongs to the corresponding category. As an example, the suspicious traffic identification model can output a vector [0.2, 0.8, 0.1], where 0.2 indicates that the confidence that the sample belongs to class A is 0.2, 0.8 indicates that the confidence that the sample belongs to class B is 0.8, and 0.1 indicates that the confidence that the sample belongs to class C is 0.1. It can be determined that the sample belongs to class B.
[0080] In some embodiments, the input of the suspicious traffic identification model may further include a reference malicious value 510 - 3 of the byte distribution probability vector. The method for determining the byte distribution probability vector may refer to step 310 and its related description.
[0081] The reference malicious value refers to the possibility that the byte distribution probability vector is suspicious traffic.
[0082] In some embodiments, the reference malicious value of the byte distribution probability vector may be determined based on historical data or the like.
[0083] In some embodiments, determining a reference malicious value based on historical data includes obtaining a byte distribution probability vector of historical suspicious traffic using a suspicious traffic determination model, and comparing the byte distribution probability vector of the historical suspicious traffic with the byte distribution probability vector corresponding to the encrypted traffic to be tested. When the similarity is greater than a certain threshold (for example, greater than 0.8), the malicious value of the byte distribution probability vector of the historical suspicious traffic is determined as the reference malicious value of the current byte distribution probability vector.
[0084] In some embodiments, the edge attributes of the relationship graph also include a byte distribution probability vector.
[0085] In some embodiments, a reference malicious value can be obtained based on a relationship graph, including: based on the edges in the relationship graph that meet preset conditions, counting the frequency of edges whose traffic type in the edge attributes of the edges that meet the preset conditions is malicious traffic, and determining the reference malicious value based on the frequency.
[0086] In some embodiments, the preset condition is that the similarity between the byte distribution probability vector in the edge attribute and the byte distribution probability vector of the encrypted traffic to be tested is close to a preset range. The preset range can be one of a system default value, an empirical value, a manually preset value, etc. For example, for the byte distribution probability vector P2, the processor 130 can count 100 traffic flows of P2 in the network traffic within a preset time period, of which 40 are normal traffic flows and 60 are malicious traffic flows, indicating that the frequency of traffic flows with the byte distribution probability vector P2 of the encrypted traffic to be tested being malicious traffic is 60%.
[0087] The reference malicious value can be further calculated based on the aforementioned 60%, and the greater the frequency, the greater the reference malicious value. In some embodiments, the byte distribution probability vector of the current traffic to be tested is P2, and all vectors with close similarity to vector P2 are searched in the relationship graph. For example, all vectors with close similarity to vector P2 are: P3, P4, and P5, where the edges corresponding to P3 and P4 are malicious traffic; the edge corresponding to P5 is normal traffic. Then the frequency of malicious traffic is 2 / 3, and the reference malicious value of P2 is calculated based on the frequency of malicious traffic 2 / 3. The method of determining the reference malicious value according to the frequency may include determining the reference malicious value according to the rule table. For example, if the frequency of the byte distribution probability vector for malicious traffic is 60%, then the malicious value in the corresponding rule table is 80; if the frequency of the byte distribution probability vector for malicious traffic is 80%, then the malicious value in the corresponding rule table is 90.
[0088] In some embodiments of the present specification, the preset relationship map can be updated based on the correlation between the encrypted traffic information to be tested and the reference malicious value, including: comparing the frequency of the byte distribution probability vector corresponding to the encrypted traffic to be tested being malicious traffic with the frequency corresponding to the reference malicious value; if the frequency of the byte distribution probability vector corresponding to the encrypted traffic to be tested being malicious traffic is greater than the frequency corresponding to the current reference malicious value, then adding a child node to the encrypted traffic to be tested and associating the child node with the byte distribution probability vector representing malicious traffic to update the preset relationship map; if the frequency of the byte distribution probability vector corresponding to the encrypted traffic to be tested being malicious traffic is less than the frequency corresponding to the current reference malicious value, then adding a child node to the encrypted traffic to be tested and associating the child node with the byte distribution probability vector representing normal traffic to update the preset relationship map.
[0089] The embodiments of this specification obtain reference malicious values through a relationship graph, and can obtain more accurate reference malicious values based on the byte distribution probability vector obtained from a large amount of statistics. At the same time, the relationship graph is updated in real time, and the reference malicious values can be obtained in real time more accurately and efficiently.
[0090] While the basic concepts have been described above, it will be apparent to those skilled in the art that the detailed disclosure is merely illustrative and does not limit this specification. Although not explicitly stated herein, various modifications, improvements, and revisions to this specification may be made by those skilled in the art. Such modifications, improvements, and revisions are suggested in this specification and remain within the spirit and scope of the exemplary embodiments of this specification.
[0091] This specification also uses specific terms to describe the embodiments of this specification. For example, "one embodiment," "an embodiment," and / or "some embodiments" refer to a feature, structure, or characteristic associated with at least one embodiment of this specification. Therefore, it should be emphasized and noted that references to "one embodiment," "an embodiment," or "an alternative embodiment" two or more times in different locations in this specification do not necessarily refer to the same embodiment. Furthermore, certain features, structures, or characteristics of one or more embodiments of this specification may be appropriately combined.
[0092] In addition, unless expressly stated in the claims, the order of the processing elements and sequences, the use of alphanumeric characters, or the use of other names described in this specification are not intended to limit the order of the processes and methods of this specification. Although the above disclosure discusses some of the invention embodiments currently considered useful through various examples, it should be understood that such details are for illustrative purposes only, and the appended claims are not limited to the disclosed embodiments. On the contrary, the claims are intended to cover all modifications and equivalent combinations that are consistent with the spirit and scope of the embodiments of this specification. For example, although the system components described above can be implemented by hardware devices, they can also be implemented only by software solutions, such as installing the described system on an existing server or mobile device.
[0093] Similarly, it should be noted that, in order to simplify the presentation of this specification and thus facilitate understanding of one or more embodiments of the invention, the foregoing descriptions of the embodiments of this specification sometimes combine multiple features into a single embodiment, figure, or description thereof. However, this disclosure method does not imply that the subject matter of this specification requires more features than those recited in the claims. In fact, an embodiment may have fewer features than all of the features of a single disclosed embodiment.
[0094] In some embodiments, numbers describing the quantity of components and attributes are used. It should be understood that such numbers used in the description of the embodiments are modified by the modifiers "about", "approximately" or "substantially" in some examples. Unless otherwise stated, "about", "approximately" or "substantially" indicate that the numbers are allowed to vary by ±20%. Accordingly, in some embodiments, the numerical parameters used in the description and claims are approximate values, which may vary according to the required features of the individual embodiments. In some embodiments, the numerical parameters should take into account the specified significant digits and adopt the general method of retaining digits. Although the numerical domains and parameters used to confirm the breadth of their range in some embodiments of this specification are approximate values, in specific embodiments, the settings of such numerical values are as accurate as possible within the feasible range.
[0095] Each patent, patent application, patent application publication, and other materials, such as articles, books, specifications, publications, and documents, cited in this specification is hereby incorporated by reference in its entirety. This excludes any application history documents that are inconsistent with or conflicting with the content of this specification, as well as any documents (currently or subsequently appended to this specification) that limit the broadest scope of the claims of this specification. It should be noted that if the descriptions, definitions, and / or terminology used in the accompanying materials are inconsistent or conflicting with the content of this specification, the descriptions, definitions, and / or terminology used in this specification shall prevail.
[0096] Finally, it should be understood that the embodiments described in this specification are intended only to illustrate the principles of the embodiments of this specification. Other variations may also fall within the scope of this specification. Therefore, by way of example and not limitation, alternative configurations of the embodiments of this specification may be considered consistent with the teachings of this specification. Accordingly, the embodiments of this specification are not limited to the embodiments explicitly described and illustrated in this specification.
Claims
1. A method for determining suspicious traffic in encrypted traffic, characterized in that: The method comprises: Collect the encrypted traffic to be tested and extract the encrypted traffic characteristics of the encrypted traffic to be tested; wherein the encrypted traffic characteristics include a first traffic characteristic and a second traffic characteristic, the first traffic characteristic includes access characteristic information, protocol characteristic information, transmission characteristic information, and a byte distribution probability vector, and the second traffic characteristic is obtained through a relationship graph; the relationship graph includes domain name nodes, entity nodes, edges connecting the entity nodes, and edges connecting the entity nodes and the domain name nodes, and the edge attributes of the edges of the relationship graph include communication-related data, traffic type, and the byte distribution probability vector; Determining the traffic type of the encrypted traffic to be tested based on the encrypted traffic characteristics of the encrypted traffic to be tested includes: The encrypted traffic features of the encrypted traffic to be tested are processed based on a suspicious traffic identification model to determine the traffic type of the encrypted traffic to be tested; the traffic types include normal traffic and suspicious traffic; the suspicious traffic identification model is a machine learning model, the input of the suspicious traffic identification model includes the encrypted traffic features and a reference malicious value of the byte distribution probability vector, and the output includes one of the suspicious traffic and the normal traffic and a classification vector, the reference malicious value is obtained based on the relationship graph, and the classification vector includes the confidence level that the encrypted traffic to be tested belongs to different categories of suspicious traffic; In response to the traffic type of the encrypted traffic to be tested being the suspicious traffic, a decryption DPI module is used to perform subsequent decryption analysis on the encrypted traffic to be tested.
2. A system for determining suspicious traffic in encrypted traffic, characterized in that: The system comprises: A traffic feature acquisition module is configured to collect encrypted traffic to be tested and extract encrypted traffic features of the encrypted traffic to be tested; wherein the encrypted traffic features include a first traffic feature and a second traffic feature, the first traffic feature including access feature information, protocol feature information, transmission feature information, and a byte distribution probability vector, and the second traffic feature is acquired through a relationship graph; the relationship graph includes domain name nodes, entity nodes, edges connecting the entity nodes, and edges connecting the entity nodes and the domain name nodes, and the edge attributes of the edges of the relationship graph include communication-related data, traffic type, and the byte distribution probability vector; a traffic type determining module, configured to determine the traffic type of the encrypted traffic to be measured based on the encrypted traffic characteristics of the encrypted traffic to be measured; The traffic type determination module is further configured to: The encrypted traffic features of the encrypted traffic to be tested are processed based on a suspicious traffic identification model to determine the traffic type of the encrypted traffic to be tested; the traffic types include normal traffic and suspicious traffic; the suspicious traffic identification model is a machine learning model, the input of the suspicious traffic identification model includes the encrypted traffic features and a reference malicious value of the byte distribution probability vector, and the output includes one of the suspicious traffic and the normal traffic and a classification vector, the reference malicious value is obtained based on the relationship graph, and the classification vector includes the confidence level that the encrypted traffic to be tested belongs to different categories of suspicious traffic; The decryption DPI module performs subsequent decryption analysis on the encrypted traffic to be tested in response to the traffic type of the encrypted traffic to be tested being the suspicious traffic.
3. A device for determining suspicious traffic in encrypted traffic, the device comprising at least one processor and at least one memory; the at least one memory is used to store computer instructions; the at least one processor is used to execute at least part of the computer instructions to implement the method as claimed in claim 1. 4 . A computer-readable storage medium storing computer instructions, which implement the method according to claim 1 when the computer instructions are executed by a processor.
Citation Information
Patent Citations
Network monitoring method, network monitoring device and electronic equipment
CN110611651A
Malicious encrypted traffic detection method and system based on behavior analysis
CN111277587A
Sandbox-based encrypted traffic processing method, system and device, and medium
CN113923021A