Illegal harmful application detection method based on encrypted traffic analysis

By analyzing encrypted traffic, constructing heterogeneous communication graphs, and fusing multi-view features, the problem of the wide variety and short lifespan of illegal and harmful applications in detection is solved, and efficient identification and classification of illegal and harmful applications is achieved.

CN121037014APending Publication Date: 2025-11-28INSTITUTE OF INFORMATION ENGINEERING CHINESE ACADEMY OF SCIENCES
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511019251.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-23
Publication Date
2025-11-28

AI Technical Summary

Technical Problem

Existing technologies are insufficient to effectively detect illegal and harmful applications, especially due to their wide variety, short lifespan, and inadequate communication feature modeling capabilities, which leads to shortcomings in the detection performance of traditional methods.

Method used

By mirroring and inspecting encrypted traffic, extracting 5-tuple information and TLS handshake traffic metadata, constructing a heterogeneous communication graph, performing heterogeneous graph representation learning, and using multi-view feature fusion and attention mechanisms, a classifier is trained to perform embedded representation and classification of applications.

Benefits of technology

It improves the accuracy and robustness of detecting illegal and harmful applications, can identify unknown applications and resist interference, is highly adaptable, and is suitable for complex scenarios in the real world.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121037014A_ABST
    Figure CN121037014A_ABST
Patent Text Reader

Abstract

The invention discloses an illegal harmful application detection method based on encrypted traffic analysis, and belongs to the technical field of network security. In order to solve the technical problems of easy failure of fingerprints, insufficient communication feature modeling capability and the like caused by a wide variety of applications and a short life cycle, the method mainly adopts a mode of constructing a heterogeneous communication graph based on TLS meta-information and fusing multi-view representation learning to extract encrypted communication behavior features between the applications and a server; semantic information under different communication relations is aggregated through meta-paths, application node representation with distinction is obtained, and illegal and harmful application detection is achieved in combination with a machine learning classifier. The method can achieve the precise recognition of unknown illegal harmful applications, and is good in generalization capability and anti-interference capability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of network security technology, specifically relating to a method for detecting illegal and harmful applications based on encrypted traffic analysis. Background Technology

[0002] The rapid development of the internet, while bringing convenience to people's lives and work, has also created opportunities for criminals. Illegal and harmful activities that were previously hidden offline are gradually shifting or expanding online, significantly impacting network order, social security, and economic development. Therefore, how to effectively detect illegal and harmful applications has become a crucial issue that urgently needs to be addressed.

[0003] Existing research can be summarized as content-based methods, which utilize visual information (such as user interfaces), textual information (such as natural language displayed in the application), or a combination of both to analyze semantic information, thereby enabling the detection of illegal and harmful applications. These methods represent a first step in the detection of illegal and harmful applications, but they are still at the stage of content analysis of the static user interfaces of gambling applications, and suffer from practical problems such as high overhead in image and text analysis and recognition, ease of bypassing, and obfuscation.

[0004] Communication features, represented by encrypted traffic, are an important component of the dynamic features of mobile applications, but have not been thoroughly explored in the field of illicit and harmful application detection. Although existing encrypted traffic analysis methods have shown excellent performance in diverse business and scenarios, they are not well-suited for the detection of illicit and harmful applications. The main reason for this is that this scenario differs from other tasks in two significant ways: (1) Illicit and harmful applications are diverse and have short lifespans. Traditional encrypted traffic analysis methods rely on building fingerprints for specific applications and performing application-level detection in a closed world. However, due to the diversity of illicit and harmful applications, it is impractical to generate fingerprints for each application individually. At the same time, this method also limits its ability to discover unknown illicit and harmful applications to some extent. In addition, some illicit and harmful applications are involved in fraudulent activities, and their purpose is to quickly profit and evade regulation. Therefore, they have relatively short lifespans, which significantly shortens the validity period of the constructed fingerprints. (2) Insufficient characterization ability for illicit and harmful applications. Illicit and harmful applications have inherent family-like communication features that are different from ordinary legitimate applications. These features are mainly reflected in the inter-flow correlation and inter-application correlation of application communication traffic. Traditional methods mostly rely on extracting features from a single stream, but cannot capture the relationships between streams and applications, thus failing to fully characterize the specific communication features of illegal and harmful applications. Summary of the Invention

[0005] The purpose of this invention is to propose an illegal and harmful application detection method based on encrypted traffic analysis, which can solve problems such as the wide variety of applications, short application lifespan leading to short fingerprint validity period, and insufficient communication feature modeling capabilities, thereby improving the detection effect.

[0006] To achieve the above objectives, the technical solution adopted by the present invention is as follows:

[0007] A method for detecting malicious and illegal applications based on encrypted traffic analysis includes the following steps:

[0008] 1) Mirror and inspect the encrypted traffic generated by the application, extract the five-tuple information from each packet: source IP, source port, destination IP, destination port and protocol; and reassemble the data stream based on the five-tuple, further extract various metadata from the TLS handshake traffic to form the initial characteristics of each stream;

[0009] 2) Construct a heterogeneous communication graph based on the initial features, including constructing a node set, a node feature matrix, and a relationship set consisting of application nodes and server nodes, to represent the communication relationship between the application and the server;

[0010] 3) Based on the heterogeneous communication graph, the application node is embedded in the representation of the server view and the application view from the perspectives of server view and application view, respectively, through server view information aggregation, application view information aggregation, multi-view feature fusion and classification model training.

[0011] 4) Use the trained classifier to classify the embedded representations of unlabeled application nodes and output the corresponding application category results to determine whether the application is illegal or harmful.

[0012] Further, in step 1), 18 types of metadata are extracted, including TLS version, cipher suite, signature algorithm, compression method, key exchange algorithm, client random number digest, server random number digest, SNI information, ALPN protocol declaration, supported extended fields, supported elliptic curves, elliptic curve point format, Session Ticket identifier, common name in the server certificate, organization name, certificate authority, certificate validity period, and server public key digest.

[0013] Furthermore, the steps in step 2) to construct the node set include:

[0014] Build application nodes based on the unique application's package name;

[0015] The server node is constructed based on the server's IP address.

[0016] Furthermore, step 2) involves constructing the node feature matrix, which includes:

[0017] By using the word segmentation dictionary mapping method, the non-numerical features in the TLS meta-information are converted into numerical features that can be processed;

[0018] Logarithmic transformation is performed on the two features of the flow average load and the total transmission load within the time window, which follow the long-tailed distribution, to compress the numerical interval;

[0019] The features of each dimension obtained through the above processing are aggregated by feature splicing to form the final original features of the node;

[0020] Based on the final original features, a node feature matrix is constructed.

[0021] Furthermore, the steps of constructing the relationship set in step 2) include:

[0022] For all communication flows between each pair of applications and servers, communication flow blocks are divided based on a sliding time window;

[0023] According to the magnitude relationship between the traffic payload p and the flow initiation frequency τ of the communication flow block and the preset thresholds p0 and τ0 respectively, the edge relationships are classified to generate different relationship types and a relationship set is constructed.

[0024] Furthermore, four relationship types R0, R1, R2, and R3 are generated by classification, specifically including:

[0025] If p < p0 and τ < τ0, the relationship type is R0, indicating that the application and the server only have low-frequency and low-data-volume requests, and there is a probability of having some heartbeat connections;

[0026] If p < p0 and τ > τ0, the relationship type is R1, indicating that there are high-frequency and low-data-volume requests between the application and the server, and there is a probability of having a configuration distribution connection with extremely high latency requirements when the application starts;

[0027] If p > p0 and τ < τ0, the relationship type is R2, indicating that there is a probability of having a data and resource transmission connection;

[0028] If p > p0 and τ > τ0, the relationship type is R3, indicating that the application and the server have high-frequency and high-data-volume communication requests.

[0029] Furthermore, the steps of server view aggregation in step 3) include:

[0030] The original feature vectors of the server nodes and application nodes are respectively mapped to the same latent feature space through linear transformation;

[0031] Based on different communication relationship types, the multi-head attention mechanism is used to calculate the attention weights of the server neighbor nodes to the application nodes;

[0032] The potential vectors of the server's neighboring nodes are weighted and aggregated according to the calculated weights to obtain the server view embedding representation of the application node.

[0033] Furthermore, the steps in step 3) of applying view aggregation include:

[0034] Based on the different communication relationship edge types between application nodes and server nodes in the heterogeneous communication graph, a variety of meta-paths with specific semantics are defined;

[0035] Based on the meta-path, sample relevant application nodes that are associated with the target application node from the application node's neighbors;

[0036] An attention mechanism is used to calculate the attention weights of relevant application nodes to the target node. The latent vectors of relevant application nodes are aggregated according to the weights to generate the node representation of the application view.

[0037] Furthermore, step 3) of multi-view feature fusion includes:

[0038] Linear transformation and activation calculation are performed on the embedded representations of application nodes generated from server views and application views to obtain the global importance weights of view-relationship pairs;

[0039] Based on the global weights, the node embedding results of each view and each relationship are weighted and fused to generate the final application node embedding representation.

[0040] Furthermore, the steps for training the classification model in step 3) include:

[0041] The generated embedded representations of labeled application nodes are used as input to train a linear support vector machine classifier, with the focus loss function used to optimize the model parameters during training.

[0042] The beneficial effects achieved by this invention are as follows:

[0043] 1. This invention can avoid the problems of high overhead and easy bypass in the feature acquisition, modeling and recognition stages of image recognition methods, and achieve more stable identification of illegal and harmful applications by encrypting the metadata and behavioral characteristics of traffic.

[0044] 2. This invention effectively models the communication behavior characteristics between applications and servers by extracting TLS metadata, traffic load, and packet timing characteristics, combined with a sliding time window mechanism and flow block partitioning strategy, thereby improving the ability to model complex communication intentions.

[0045] 3. This invention proposes an application-server heterogeneous graph construction method, which models the node information and communication relationships in encrypted traffic as multiple types of nodes and edges, further improving the ability to express communication semantics and model node behavior.

[0046] 4. This invention employs a heterogeneous graph representation learning method based on meta-paths, learning application node embedding representations from both server view and application view perspectives, and integrating multi-relationship and multi-view information to enhance the distinguishability between illegal and harmful applications and normal applications in high-dimensional space.

[0047] 5. This invention designs a graph representation learning structure that integrates attention mechanism and multi-head mechanism, which can fully explore the influence of key neighbor nodes and communication relationships on node classification, and improve the robustness and discriminativeness of the representation.

[0048] 6. This invention combines the focus loss function to optimize the classifier training process, which can improve the detection accuracy and enhance the ability to focus on difficult-to-identify samples, especially given the scarcity of illegal and harmful applications.

[0049] 7. This invention ultimately classifies the learned node representations using a linear support vector machine, exhibiting excellent detection performance for unknown illegal and harmful applications, version update applications, and obfuscated traffic in encrypted communications, making it more adaptable.

[0050] 8. Compared with existing methods, the present invention achieves higher detection performance on real-world sample datasets, demonstrating better accuracy, generalization ability and anti-interference ability. Attached Figure Description

[0051] Figure 1 This is an architecture diagram of a method for detecting illegal and harmful applications based on encrypted traffic analysis.

[0052] Figure 2 This is a flowchart of a method for detecting illegal and harmful applications based on encrypted traffic analysis.

[0053] Figure 3 The figure shows the experimental results when there is noise flow disturbance.

[0054] Figure 4 This is a graph showing the experimental results when there are application version updates. Detailed Implementation

[0055] To make the various technical features, advantages, or effects of the present invention more apparent and understandable, detailed descriptions are provided below through embodiments and accompanying drawings.

[0056] This invention proposes a method for detecting illegal and harmful applications based on encrypted traffic analysis, such as... Figure 1-2 As shown, the specific processing steps are explained below:

[0057] S1: Traffic Acquisition and Preprocessing.

[0058] First, all application-generated traffic is mirrored and inspected, and a five-tuple of information (source IP, source port, destination IP, destination port, and protocol) is extracted from each packet. Packets with identical five-tuple information are then reassembled into a data stream. Each stream represents a communication process between the application and the server. Subsequently, a key set of metadata is extracted from the TLS handshake traffic, including TLS version, cipher suite, signature algorithm, compression method, key exchange algorithm, client random number digest, server random number digest, SNI information, ALPN protocol declaration, supported extended fields, supported elliptic curves, elliptic curve point format, Session Ticket identifier, Common Name (CN), Organization Name (O), Certificate Authority (Issuer), certificate validity period, and server public key digest from the server certificate—a total of 18 types of metadata. This provides initial features for subsequent application and server node representation learning.

[0059] S2: Heterogeneous communication graph construction.

[0060] This step constructs a heterogeneous communication graph G = (V, R, F) for the application based on the preprocessed traffic and extracted TLS metadata in S1, including the construction of three types of entities: node set V, node feature matrix F, and relation set R.

[0061] S2-1: Construction of node set V.

[0062] Heterogeneous communication graphs model the communication relationships between applications and servers. Therefore, applications and servers correspond to two different types of nodes in the heterogeneous communication graph: application nodes V. A and server node V S This step uses the package name (a unique application identifier in the Android application ecosystem) and IP address to represent the application node and server node, respectively, i.e., each application node i∈V A Each server node j∈V corresponds to a unique application package name. S This represents the IP address of a server.

[0063] S2-2: Construction of the node feature matrix F.

[0064] Aggregate the TLS meta-information extracted in step S1 with the software package name of the application node and the IP address of the server node as the primary keys to form the corresponding TLS fingerprint, which is used as the initial feature representation of the node. Since the extracted TLS meta-information is heterogeneous, containing numerical and non-numerical fields, it cannot be directly used to represent node features. Therefore, the original data needs to be uniformly encoded before feature aggregation. Specifically, for non-numerical fields in the TLS meta-information, a tokenizer is used for mapping processing to convert them into a processable numerical form; for numerical fields such as flow average load and total transmission load within a time window that follow a long-tailed distribution, logarithmic transformation is performed to compress the data scale and enhance distinguishability. Finally, the features of each dimension obtained through the above processing are aggregated by feature splicing to form the final raw features of the node, and a node feature matrix F is generated based on this.

[0065] S2-3: Construction of the relationship set R.

[0066] Model the inter-flow association between the communication intention of the application connecting to the server and the generated traffic. In different communication processes, the communication intentions of illegal and harmful applications and the server are also different. This difference is directly reflected in the effective payload of the traffic and the flow initiation frequency, and can be modeled as different communication relationships between the application and the server, thus further promoting the learning of better node representations. According to the traffic effective payload and flow initiation frequency, model the traffic interaction between the application and the server as a communication relationship between the application and the server, forming an edge between the application node and the server node in the heterogeneous graph. First, set a sliding time window threshold, denoted as T window . For all communication flows between each pair of application and server, if the time difference Δt between the request initiation times of two consecutive flows is less than T window , merge these two flows into the same flow block. This process is iterated continuously to merge the traffic into multiple flow blocks, where each flow block represents a relatively complete interactive communication intention. Subsequently, calculate the traffic effective payload p and flow initiation frequency τ in each flow block. According to the size relationship between these two metrics and the preset thresholds p0 and τ0, map these flow blocks to four specific relationship types R0, R1, R2, R3. The specific classification method is as follows:

[0067] If p < p0 and τ < τ0, the relationship type is R0, indicating that the application and the server only have low-frequency and low-data-volume requests, and there may be some heartbeat connections; [

[0068] If p < p0 and τ > τ0, the relationship type is R1, indicating that there are high-frequency and low-data-volume requests between the application and the server, and there may be a configuration download connection with extremely high latency requirements when the application starts;

[0069] If p>p0, τ<τ0, then the relation type is R2, which may involve data and resource transfer connections;

[0070] If p>p0 and τ>τ0, then the relation type is R3, indicating that the application and the server have high-frequency, high-volume communication requests.

[0071] S3: Application node representation learning.

[0072] This step is based on the heterogeneous communication graph constructed in S2 to learn the distinguishable representation of application nodes. It enhances the representation of application nodes from two different views: server and application. It includes four sub-steps: server view information aggregation, application view information aggregation, multi-view feature fusion, and classification model training.

[0073] S3-1: Server view information aggregation.

[0074] This approach facilitates representation learning for application nodes by aggregating server-side information. The characteristics of servers accessed by an application directly reflect its communication intent. For example, if an application accesses a server suspected of providing illegal or harmful services, it is more likely to be an illegal or harmful application compared to other applications that have not accessed such servers. Therefore, utilizing the characteristics of servers that have direct communication relationships with the application is crucial for determining whether an application is illegal or harmful. To address this issue, the features of server nodes that communicate with each application node are aggregated into the application feature representation to fuse server view information. Since server nodes and application nodes are of different types, they reside in different feature spaces. Therefore, for each application node, server neighbor nodes are first randomly sampled according to the sampling ratio of neighbor nodes. Then, linear transformations are performed on both application nodes and server nodes separately, projecting the feature vectors onto the same latent vector space. Formally, for V... T The type is T∈{V A V s For node i in}, we have:

[0075] h i =W T ·x i

[0076] in and Let be the original feature vector and the projected latent vector of node i, respectively. It is the parameter weight matrix of a node of type T.

[0077] Different relationships represent different communication intentions; therefore, servers with different communication relationships should contribute differently to the representation of application nodes. An attention mechanism is employed to learn the importance of each server's neighbors and to fuse information from these neighboring nodes. Given... For application node i and server node In a set of server neighbors based on relation r, application-server node pairs<i,j> Its importance can be expressed as follows:

[0078]

[0079]

[0080] in, h is an intermediate variable used for attention calculation. i and h j It is the predicted latent vector of application node i and server node j, α r Let be the attention matrix of relation r, where σ(·) represents the activation function and || represents the concatenation operation.

[0081] Finally, by adding all the server neighbor features to their corresponding importance, we can obtain the server view embedding, as follows:

[0082]

[0083] Where K represents the number of heads in the multi-head attention mechanism, This represents the concatenation operation of k attention vectors.

[0084] S3-2: Application view information aggregation.

[0085] This paper proposes a method to further enhance application node representations by leveraging inter-application associations. Two applications are defined as "related" applications if they access the same server through the same communication relationship, and they will contribute to each other's node representations. The main reasons for this association stem from three aspects: the booming development of third-party libraries, the widespread use of the Android Software Development Kit (SDK), and the family-based development code templates in gambling applications, which further lead to partially similar communication sequences and characteristics in encrypted traffic. Therefore, two applications connected to the same server S and having the same communication relationship r may have the same communication intent. Appropriately utilizing these related application node features helps generate more comprehensive application node representations. Based on the above analysis, this paper proposes a meta-path-based approach to fuse inter-application association features. In heterogeneous graphs, such related nodes can be interpreted through different semantic paths, which are called meta-paths. Since the communication relationship between an application and a server is defined as four different relation edge types R0, R1, R2, and R3, this method defines four corresponding meta-paths to reveal different inter-application communication associations between two applications, including AR0A meta-path, AR1A meta-path, AR2A meta-path, and AR3A meta-path.

[0086] AR0A metapath. The AR0A metapath connects two application nodes i∈V. A and j∈V A Both nodes are connected to the same server node k∈V through R0. S This means that the two applications only exist within a single time window T. window The system makes several accesses to the server, but the amount of data transmitted is minimal.

[0087] AR1A metapath. The AR1A metapath connects two application nodes i∈V. A and j∈V A They are connected to the same server node k∈V via R1. S This indicates that both applications are in the time window T window The system frequently attempted to access the server, but only transferred a small amount of data.

[0088] AR2A metapath. The AR2A metapath exists between two application nodes i∈V. A and j∈V A A semantic relationship has been established between the two nodes, and both nodes are connected to the same server node k∈V through R2. S This means that the two applications are in a time window T window The number of internal server accesses was infrequent, but a large amount of data was transferred.

[0089] AR3A metapath. The AR3A metapath exists between two application nodes i∈V. A and j∈V A A reachable path is established between the two nodes, both of which are linked to the same server node k∈V via R3. S This means that both applications are in the time window T window It frequently accesses the server and transfers large amounts of data.

[0090] Based on predefined semantically rich meta-paths, information fusion resulting from inter-application associations is integrated into the representation of application nodes. Specifically, for each application node, meta-path-based application neighbor nodes are sampled, and an attention mechanism is used to aggregate information from neighbor nodes from different meta-paths with different weights. Formally, given... For application node i and application node j in In the context of neighbor nodes based on metapath P, the neighbor node pair based on metapath P is...<i,j> Attention value This can be expressed as follows:

[0091]

[0092] Among them, hi and h j Let be the predicted latent vectors of nodes i and j, σ(·) denotes the activation function, and || denotes the concatenation operation.

[0093] Similar to server view information fusion, application view information fusion can also be obtained by aggregating all applications based on metapaths, as shown below:

[0094]

[0095] Where K represents the number of heads in the multi-head attention mechanism.

[0096] S3-3: Multi-view feature fusion.

[0097] After generating the corresponding representations from the server view and application view, the above information is organically integrated to generate a more comprehensive node representation. Since different relationships represent different communication characteristics, it is necessary to measure the importance of each relationship in different views. Therefore, this method computes the view T∈{V A V s Importance weights based on the embedding of relation r in} as follows:

[0098]

[0099] in and These are the parameters for the linear layer.

[0100] Based on the weights of each relation learned in each view, the embedding representation h of the application node i is calculated. i :

[0101]

[0102] S3-4: Classification model training.

[0103] The goal of training the classification model is to distinguish between illegal / harmful applications and legitimate applications based on the embedding representations of application nodes. This problem is equivalent to a binary classification task. Therefore, a linear support vector machine (LinearSVC) classifier is trained using labeled application node embedding representations as input to obtain the final classification result.

[0104] y c =W c ·h i +b c

[0105] p c =softmax(y c )

[0106] Among them W c and b c For linear layer parameters, y c For the output of the linear layer, p c This represents the predicted probability distribution after softmax smoothing.

[0107] Because of the class imbalance between illegal and harmful applications and normal applications, a design is made to minimize the focus loss. fl To achieve model optimization:

[0108] L fl =-(1-pc) γ log(p c )

[0109] Where γ≥0 represents the focusing parameter.

[0110] By minimizing focal length loss, this method enables the model to quickly focus on applications that are difficult to classify and to learn more discriminative node representations for them.

[0111] S4: Detection of illegal and harmful applications.

[0112] The trained classifier is used to predict the embedding representation of unlabeled application nodes, and the binary classification result of the application node as normal or illegal / harmful is output to complete the detection of illegal / harmful applications.

[0113] The following are specific examples:

[0114] The dataset used in this embodiment comprises two sub-datasets: AndGamb2023 and AndGamb2023-Openworld, which cover 175 mainstream gambling applications currently used in China, obtained in cooperation with an authoritative domestic institution. The types of gambling games involved include lottery, sports, card games, and online games; and 245 mainstream normal applications included in CSTNET-300(25). The two sub-datasets cover the same range of applications, differing only in dataset size and class sample balance. The former contains 19K gambling traffic entries and 72K normal traffic entries; the latter contains 193K gambling traffic entries and 875K normal traffic entries. Detailed information for each dataset is shown in Tables 1 and 2.

[0115] Table 1. Detailed information about the AndGamb2023 dataset.

[0116] category platform Application Number Sample size Number of streams Flow rate Gambling apps Android 175 767 19k 0.26G Normal application Android 245 1037 72k 24.95G

[0117] Table 2. Detailed information about the AndGamb2023-Openworld dataset.

[0118] category platform Application Number Sample size Number of streams Flow rate Gambling apps Android 175 767 193k 1.29G Normal application Android 245 12361 875k 59.04G

[0119] Processing is performed according to step S1 of the method of the present invention, using the captured AndGamb2023 and AndGamb2023-OpenWorld datasets as experimental data, and obtaining data packets from the PCAP file and reassembling them into a stream, extracting a total of 18 kinds of metadata from the TLS handshake traffic, including TLS version, cipher suite, signature algorithm, certificate fields, SNI information, etc., as well as the initial representation of the application and server nodes.

[0120] Processing is performed according to step S2 of the method of the present invention, and a heterogeneous communication graph is constructed using the parsed traffic data, wherein the feature dimension of the application node is 193, the feature dimension of the server node is 1044, the sliding time window threshold is 0.5s, the traffic effective load threshold is 8192B, the flow initiation frequency threshold is 3 times, and four edge relationship types are confirmed as shown in Table 3.

[0121] Table 3. Mapping of flow payload and flow initiation frequency for four edge relationship types.

[0122]

[0123] The method of the present invention is processed according to step S3, using a heterogeneous communication graph learning algorithm to learn a distinguishable representation of the application node, wherein the number of heads of the attention mechanism is 8, the hidden layer dimension of each attention head is set to 50, the learning rate is set to 0.005, the focus parameter of the focus loss is set to 2, the node encoder update process is realized through gradient descent, and the model parameters are saved for generating node representations, and the generated node representations are used to train a linear support vector machine classifier.

[0124] The method of the present invention is processed according to step S4, and the trained classifier is used to represent the unlabeled application node to obtain the normal / harmful binary classification result corresponding to the application node, thereby realizing the detection of illegal and harmful applications.

[0125] Experimental test:

[0126] Representative methods in the field of encrypted traffic analysis, including AppScanner, k-FP, DF, FS-Net, DLWF, App-net, GAP-WF, and GraphDApp, were selected as comparison methods to be tested against the method of this invention (also known as the communication association graph-based method, HeCGamb). First, the effectiveness of detecting illegal and harmful applications in a closed world was verified on the AndGamb2023 dataset. The experimental results are shown in Table 4. The experimental results show that the method achieves an F1 score of 0.994 with low computational time overhead, outperforming the other comparison methods.

[0127] Table 4 Comparative Experiment Results of Illegal and Harmful Application Detection in a Closed World

[0128]

[0129] The performance of the method was then verified in the scenario of identifying unknown illegal and harmful applications. The experiment randomly divided AndGamb2023 into AndGamb2023-Known and AndGamb2023-Unknown. The former contained 140 illegal and harmful applications and 195 normal applications, while the latter involved 35 illegal and harmful applications and 50 normal applications. The experimental results are shown in Table 5. The results show that the method achieved the best AUC value of 0.992, confirming that the method can generalize the features of known illegal and harmful applications to the detection task of unknown illegal and harmful applications.

[0130] Table 5. Comparison of experimental results for detecting unknown illegal and harmful applications.

[0131]

[0132] To further validate the effectiveness of the method in an open world scenario, the experiment was conducted using AndGamb2023-Openworld. The experimental setup followed the real-world scenario, which may involve unknown illegal or harmful applications, and the natural class imbalance between normal and illegal / harmful applications. The experimental results are shown in Table 6. The results show that the method still achieved an AUC value of 0.968 in this scenario, which is better than other comparative methods, and can be used for the detection of illegal and harmful applications in the real world.

[0133] Table 6. Comparative Experiment Results of Illegal and Harmful Application Detection in Open World

[0134]

[0135] Meanwhile, experiments evaluated the method's performance under real-world noise flows and application version update interference, with results as follows: Figure 3-4 As shown in the figure, experimental results demonstrate that the method maintains stable performance under noisy traffic interference scenarios at various noise ratios, achieving an AUC of 93.6% even under high perturbation conditions with a noise ratio of 15%. In application version update scenarios, the method's AUC exceeds 96%, showing only a 0.7 percentage point decrease compared to historical versions. These experimental results confirm that the method maintains stable and robust performance in complex scenarios involving noisy traffic and application version updates, particularly for illegal and harmful detection tasks.

[0136] Although the present invention has been disclosed above with reference to embodiments, it is not intended to limit the present invention. Appropriate modifications or equivalent substitutions made by those skilled in the art to the technical solutions of the present invention should be covered within the protection scope of the present invention, which is defined by the claims.

Claims

1. An illegal and harmful application detection method based on encrypted traffic analysis, comprising the following steps: 1) Mirror and check the encrypted traffic generated by the application, extract five-tuple information from each data packet: source IP, source port, destination IP, destination port, and protocol; and recombine the data stream based on the five-tuple, and further extract various meta-information from the TLS handshake traffic to form the initial features of each flow. 2) Construct a heterogeneous communication graph based on the initial features, including constructing a node set, a node feature matrix, and a relationship set composed of application nodes and server nodes to represent the communication relationship between the application and the server. 3) Based on the heterogeneous communication graph, through server view information aggregation, application view information aggregation, multi-view feature fusion, and classification model training, perform embedded representation learning on the application nodes from two perspectives of the server view and the application view respectively. 4) Use the trained classifier to classify the embedded representation of the unlabeled application nodes, and output the corresponding application category results to determine whether the application is an illegal and harmful application.

2. The method as described in claim 1, characterized in that, In step 1), 18 types of meta-information are extracted, including TLS version, cipher suite, signature algorithm, compression method, key exchange algorithm, client random number digest, server random number digest, SNI information, ALPN protocol declaration, supported extension fields, supported elliptic curves, elliptic curve point formats, Session Ticket identifier, common name in the server certificate, organization name, certificate authority, certificate validity period, server public key digest.

3. The method as described in claim 1, characterized in that, The steps for constructing the node set in step 2) include: Construct an application node according to the package name of the unique application. Construct a server node according to the IP address of the server.

4. The method as described in claim 1, characterized in that, The steps for constructing the node feature matrix in step 2) include: Adopt a word segmentation dictionary mapping method to convert non-numeric features in the TLS meta-information into processable numerical features. Perform logarithmic transformation on two features, the average flow load and the total transmission load within the time window, which follow a long-tailed distribution, to compress the numerical interval. Aggregate the features of each dimension obtained through the above processing by feature splicing to form the final original features of the node. Construct a node feature matrix based on the final original features.

5. The method as described in claim 1, characterized in that, The steps for constructing the relationship set in step 2) include: For all communication flows between each pair of applications and servers, divide the communication flow into blocks based on a sliding time window. Classify the edge relationships according to the size relationships between the traffic payload p and the flow initiation frequency τ of the communication flow block and the preset thresholds p0 and τ0 respectively, generate different relationship types, and construct the relationship set.

6. The method as described in claim 5, characterized in that, Four relationship types R0, R1, R2, and R3 are generated by classification, specifically including: If p < p0 and τ < τ0, the relationship type is R0, indicating that the application and the server have only low-frequency and low-data-volume requests, and there is a probability of partial heartbeat connections. If p < p0 and τ > τ0, the relationship type is R1, indicating that there are high-frequency and low-data-volume requests between the application and the server, and there is a probability of a configuration download connection with extremely high latency requirements when the application starts. If p>p0, τ<τ0, then the relation type is R2, indicating that there is a probability that a data and resource transfer connection exists; If p>p0 and τ>τ0, then the relation type is R3, indicating that the application and the server have high-frequency, high-volume communication requests.

7. The method as described in claim 1, characterized in that, Step 3) of server view aggregation includes the following steps: The original feature vectors of the server node and the application node are mapped to the same latent feature space through linear transformations. Based on different communication relationship types, a multi-head attention mechanism is used to calculate the attention weights of the server's neighbor nodes to the application node; The potential vectors of the server's neighboring nodes are weighted and aggregated according to the calculated weights to obtain the server view embedding representation of the application node.

8. The method as described in claim 1, characterized in that, Step 3) involves applying view aggregation, including: Based on the different communication relationship edge types between application nodes and server nodes in the heterogeneous communication graph, a variety of meta-paths with specific semantics are defined; Based on the meta-path, sample relevant application nodes that are associated with the target application node from the application node's neighbors; An attention mechanism is used to calculate the attention weights of relevant application nodes to the target node. The latent vectors of relevant application nodes are aggregated according to the weights to generate the node representation of the application view.

9. The method as described in claim 1, 7, or 8, characterized in that, Step 3) of multi-view feature fusion includes the following steps: Linear transformation and activation calculation are performed on the embedded representations of application nodes generated from server views and application views to obtain the global importance weights of view-relationship pairs; Based on the global weights, the node embedding results of each view and each relationship are weighted and fused to generate the final application node embedding representation.

10. The method as described in claim 1, characterized in that, Step 3) includes the following steps for training the classification model: The generated embedded representations of labeled application nodes are used as input to train a linear support vector machine classifier, with the focus loss function used to optimize the model parameters during training.