Network traffic identification method and system for unknown cryptographic protocol

By preprocessing and feature extraction of dynamic traffic data, building and dimensionality reduction knowledge graphs, combining cluster analysis and evaluation strategies, the detection problem of unknown cryptographic protocols is solved, and accurate identification and network security are improved.

CN120110796AActive Publication Date: 2025-06-06LIZHUANG INFORMATION TECH (SUZHOU) CO LTD

Patent Information

Application Number
CN202510580699.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-07
Publication Date
2025-06-06
Estimated Expiration
2045-05-07

AI Technical Summary

Technical Problem

The prior art lacks effective detection methods for unknown password protocols, resulting in the inability to identify encrypted traffic or hidden traffic types, which in turn leads to network security vulnerabilities.

Method used

By preprocessing the captured dynamic traffic data, reading the predetermined feature dimensions, collecting a multi-dimensional feature parameter set, building an initial knowledge graph, and introducing a graph dimensionality reduction mechanism to obtain the target knowledge graph. Taking the target knowledge graph as the clustering benchmark, cluster analysis is carried out on the network traffic knowledge graph database, and combined with the predetermined clustering degree evaluation strategy, the target class cluster is identified and its protocol type is determined.

Benefits of technology

It realizes accurate identification of unknown password protocols, improves the accuracy of network traffic analysis, enhances network security protection capabilities, and can effectively identify unknown password protocol types.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120110796A_ABST
    Figure CN120110796A_ABST
Patent Text Reader

Abstract

The invention provides a network traffic identification method and system for an unknown cryptographic protocol, and relates to the technical field of network traffic analysis, and the method comprises the steps: carrying out the preprocessing of captured dynamic traffic data, and obtaining target traffic data; acquiring a multi-dimensional feature parameter set based on predetermined feature dimension collection; constructing an initial knowledge graph, and introducing a graph dimension reduction mechanism to carry out dimension reduction processing to obtain a target knowledge graph; performing clustering analysis on the network flow knowledge graph database to obtain a target class cluster; reading a predetermined clustering degree evaluation strategy, and performing clustering degree evaluation analysis to obtain a target clustering degree; and if the preset clustering degree limit value is reached, marking the target protocol type corresponding to the target class cluster as the protocol type of the target traffic data. The technical problem that in the prior art, a static feature recognition method is generally adopted, effective detection on an unknown password protocol is lacked, the requirement for continuous evolution of the protocol cannot be met, and then network security vulnerabilities are caused is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of network traffic analysis, and in particular to a method and system for identifying network traffic of unknown cryptographic protocols. Background Art

[0002] Unknown cryptographic protocols usually refer to network protocols that hide their actual content through encryption or other methods. These protocols can evade traditional network traffic analysis and intrusion detection systems, thus causing certain security risks. At present, network traffic analysis technology mainly relies on traffic identification methods based on protocol features. These methods usually rely on the identification of features of known protocols, but are incapable of detecting unknown cryptographic protocols. Existing technologies lack effective detection methods for unknown cryptographic protocols, and therefore cannot adapt to the evolving needs of protocols, which leads to the inability to identify encrypted traffic or hidden traffic types, and in turn leads to network security vulnerabilities that attackers can exploit to carry out malicious activities, such as data theft, illegal access, etc. Summary of the invention

[0003] This application provides a method and system for identifying network traffic of unknown cryptographic protocols, aiming to solve the technical problem that the existing technology usually adopts a static feature recognition method, lacks effective detection of unknown cryptographic protocols, and is unable to adapt to the needs of the ever-evolving protocols, which in turn leads to network security vulnerabilities.

[0004] The first aspect disclosed in the present application provides a method for identifying network traffic of an unknown cryptographic protocol, the method comprising: preprocessing captured dynamic traffic data to obtain target traffic data; reading predetermined feature dimensions, and collecting a multidimensional feature parameter set of the target traffic data based on the predetermined feature dimensions; constructing an initial knowledge graph in combination with the multidimensional feature parameter set, and introducing a graph dimensionality reduction mechanism to perform dimensionality reduction processing on the initial knowledge graph to obtain a target knowledge graph; using the target knowledge graph as a clustering benchmark, performing clustering analysis on a network traffic knowledge graph database to obtain a target cluster corresponding to the target knowledge graph; reading a predetermined clustering degree evaluation strategy, and performing clustering degree evaluation analysis on the target cluster and the target traffic data according to the predetermined clustering degree evaluation strategy to obtain a target clustering degree; if the target clustering degree reaches a predetermined clustering degree limit, the target protocol type corresponding to the target cluster is recorded as the protocol type of the target traffic data.

[0005] The second aspect disclosed in the present application provides a network traffic identification system for an unknown cryptographic protocol, the system is used for the above-mentioned network traffic identification method for an unknown cryptographic protocol, the system includes: a data preprocessing module, used to preprocess the captured dynamic traffic data to obtain target traffic data; a feature parameter acquisition module, used to read a predetermined feature dimension, and collect a multidimensional feature parameter set of the target traffic data based on the predetermined feature dimension; a dimensionality reduction processing module, used to construct an initial knowledge graph in combination with the multidimensional feature parameter set, and introduce a graph dimensionality reduction mechanism to reduce the dimension of the initial knowledge graph Processing to obtain a target knowledge graph; a clustering analysis module, used to perform clustering analysis on the network traffic knowledge graph database with the target knowledge graph as the clustering benchmark, to obtain a target cluster corresponding to the target knowledge graph; an evaluation and analysis module, used to read a predetermined clustering degree evaluation strategy, and perform clustering degree evaluation analysis on the target cluster and the target traffic data according to the predetermined clustering degree evaluation strategy to obtain a target clustering degree; a protocol type determination module, used to record the target protocol type corresponding to the target cluster as the protocol type of the target traffic data if the target cluster reaches a predetermined clustering degree limit.

[0006] One or more technical solutions provided in this application have at least the following beneficial effects: By preprocessing the captured dynamic traffic data, redundant data, noise and invalid information can be effectively removed, making subsequent analysis and identification work more accurate and efficient. The target traffic data obtained after preprocessing is purer, which improves the accuracy of subsequent feature extraction and protocol identification. By reading the predetermined feature dimensions and extracting the multidimensional feature parameter set of the target traffic data based on these dimensions, the multidimensional information of the network traffic can be captured comprehensively and finely. Different feature dimensions provide a more comprehensive basis for protocol identification and enhance the richness and accuracy of feature expression. By combining the multidimensional feature parameter set to construct the initial knowledge graph and introducing the graph dimensionality reduction mechanism, high-dimensional data can be converted into a graph representation in a low-dimensional space. After dimensionality reduction, the target knowledge graph is convenient for further processing and analysis, while effectively reducing the computational complexity. The graph after dimensionality reduction retains important structural information and improves The efficiency and accuracy of subsequent analysis are improved; by taking the target knowledge graph as the clustering benchmark and performing cluster analysis on the network traffic knowledge graph database, the clustering pattern of similar traffic data can be identified. The target clusters after clustering can clearly identify different protocols or communication modes, thereby providing a clear classification basis for protocol identification; by using a predetermined clustering degree evaluation strategy, the clustering degree of the target clusters and the target traffic data is evaluated, which can quantify the degree of clustering. The evaluation results help confirm whether the clustering effect meets the predetermined standards, further improving the reliability and accuracy of the clustering analysis; if the clustering degree of the target cluster reaches the predetermined limit, the protocol type corresponding to the cluster can be accurately used as the protocol type of the target traffic data. This method ensures the automatic identification of protocol types, improves the accuracy in network traffic analysis, and can effectively identify unknown cryptographic protocol types.

[0007] The above description is only an overview of the technical solution of the present application. In order to more clearly understand the technical means of the present application, it can be implemented in accordance with the contents of the specification. In order to make the above and other purposes, features and advantages of the present application more obvious and easy to understand, the specific implementation methods of the present application are listed below. BRIEF DESCRIPTION OF THE DRAWINGS

[0008] Figure 1 A flowchart of a method for identifying network traffic of an unknown cryptographic protocol provided in an embodiment of the present application.

[0009] Figure 2 A schematic diagram of the structure of a network traffic identification system for an unknown cryptographic protocol provided in an embodiment of the present application.

[0010] Explanation of the reference numerals: data preprocessing module 10 , characteristic parameter acquisition module 20 , dimensionality reduction processing module 30 , clustering analysis module 40 , evaluation and analysis module 50 , protocol type determination module 60 . DETAILED DESCRIPTION

[0011] The embodiments of the present application provide a method and system for identifying network traffic of unknown cryptographic protocols, thereby solving the technical problem that the prior art generally adopts a static feature identification method, lacks effective detection of unknown cryptographic protocols, and is unable to adapt to the needs of the ever-evolving protocols, thereby causing network security vulnerabilities.

[0012] After introducing the basic principles of the present application, various non-limiting implementation methods of the present application will be specifically introduced below in conjunction with the accompanying drawings of the specification. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0013] Embodiment 1, as Figure 1 As shown, an embodiment of the present application provides a method for identifying network traffic of an unknown cryptographic protocol, the method comprising: The captured dynamic traffic data is preprocessed to obtain the target traffic data.

[0014] Capture dynamic traffic data from the network in real time. Dynamic traffic data includes various characteristics of data packets, such as source IP address, destination IP address, port number, packet length, protocol type, etc. Since the captured data may contain noise or irrelevant information, it needs to be preprocessed. The preprocessing process includes removing missing or incomplete data packets; removing outliers, such as wrong data packets or data packets that do not conform to the expected format; and deduplicating data to ensure that each traffic record is unique. Aggregate multiple data packets from the same session or the same data stream into a single target traffic data for further processing.

[0015] The predetermined characteristic dimensions are read, and a multi-dimensional characteristic parameter set of the target flow data is collected based on the predetermined characteristic dimensions.

[0016] Read the predetermined feature dimensions, which at least include port features, IP address features, statistical features, and content features. After reading the predetermined feature dimensions, extract specific values ​​from the target traffic data based on these dimensions and summarize them into a multi-dimensional feature parameter set. Each dimension corresponds to a feature of the traffic data, such as port, IP address, statistical value, etc. The combination of these feature parameters helps to fully describe the behavior pattern of the traffic.

[0017] An initial knowledge graph is constructed in combination with the multi-dimensional feature parameter set, and a graph dimensionality reduction mechanism is introduced to reduce the dimensionality of the initial knowledge graph to obtain a target knowledge graph.

[0018] By visualizing the relationship between multi-dimensional feature parameter sets, a graph structure is constructed to form an initial knowledge graph. Specifically, when constructing the initial knowledge graph, each network traffic feature vector is taken as a node, and each node corresponds to a multi-dimensional feature parameter set of network traffic data. The relationship (edge) between nodes is determined based on similarity measurement. For example, if two traffic data have high similarity in multiple features, such as port number, IP address, etc., there will be an edge connecting them. This edge can be assigned a weight to represent the similarity of the two nodes. Common similarity measurement methods include Euclidean distance, cosine similarity, etc. By connecting traffic data with similar features, an initial knowledge graph is constructed. In this graph, nodes represent traffic features and edges represent similarity relationships between features.

[0019] Since the initial knowledge graph constructed has high-dimensional characteristics, especially when the network traffic data contains multiple features, the feature space is very large. In order to better analyze it, the graph is reduced in dimension. The high-dimensional data is mapped to two-dimensional or three-dimensional space by adopting the graph dimensionality reduction mechanism while maintaining the local structure of the data. After dimensionality reduction, the target knowledge graph is obtained. This graph contains the node positions in the low-dimensional space and can more intuitively display the relationship and clustering structure between traffic data.

[0020] Taking the target knowledge graph as a clustering benchmark, cluster analysis is performed on the network traffic knowledge graph database to obtain the target cluster corresponding to the target knowledge graph.

[0021] The similarity structure of the nodes in the target knowledge graph will be used as the basis for clustering. Through the clustering method, different clusters will be divided based on the similarity between the nodes in the graph. For example, any graph is extracted from the network traffic knowledge graph database as a comparison object, and the similarity coefficient between the target knowledge graph and other graphs is calculated. According to the similarity coefficient, similar knowledge graphs are classified into one category. These clusters represent the patterns or protocol types of traffic data in a specific feature space. After clustering analysis, the target clusters are finally obtained. These clusters represent a set of traffic data with similar characteristics in network traffic.

[0022] A predetermined clustering degree evaluation strategy is read, and clustering degree evaluation analysis is performed on the target cluster and the target traffic data according to the predetermined clustering degree evaluation strategy to obtain a target clustering degree.

[0023] Read the predetermined clustering evaluation strategy, which is used to evaluate the quality and consistency of the target cluster. Specifically, the clustering evaluation strategy mainly defines how to quantify the similarity of data points within a cluster and the difference between clusters. The evaluation indicators include intra-cluster compactness and inter-cluster separation. The intra-cluster compactness is measured by calculating the distance between all data points in the cluster. For example, the average distance or minimum distance between data points in the cluster can be used as a measure of compactness; the inter-cluster separation is measured by calculating the distance between the cluster and other clusters. For example, the minimum distance or average distance between the target cluster and other clusters can be used as a measure of separation. By performing the above evaluation on the target cluster and the target traffic data, the target clustering value is finally obtained, which is used to judge the clustering quality of the cluster.

[0024] If the target clustering degree reaches a predetermined clustering degree limit, the target protocol type corresponding to the target cluster is recorded as the protocol type of the target traffic data.

[0025] The predetermined clustering degree limit is a pre-set threshold used to judge the quality of clustering. For example, the clustering degree limit is set to 0.7, which means that if the clustering degree value is greater than or equal to 0.7, the clustering quality is considered satisfactory; if the clustering degree value is less than 0.7, it is necessary to further adjust the clustering strategy or abandon the cluster.

[0026] If the target clustering degree is greater than or equal to the predetermined clustering degree limit, that is, the clustering quality is good, then the target cluster is considered to be a valid cluster. Each cluster represents a specific network protocol type or communication mode. Usually, the protocol type is determined based on the source, destination, port number, protocol characteristics, etc. of the traffic data. For example, a cluster may correspond to the HTTP protocol, and another cluster may correspond to the FTP protocol. When the target clustering degree meets the predetermined limit, the protocol type corresponding to the target cluster is marked as the protocol type of the target traffic data. This means that based on clustering analysis and evaluation, the network traffic data is successfully identified as a specific protocol type.

[0027] On the contrary, if the target clustering degree does not reach the predetermined limit, that is, the clustering quality is poor, the target cluster needs to be adjusted, such as re-clustering, adjusting clustering algorithm parameters, etc., or abandoning the cluster and performing other analysis methods.

[0028] Furthermore, the predetermined feature dimensions include at least port features, IP address features, statistical features and content features.

[0029] Ports are identifiers that distinguish different services or applications in computer networks. Port features are usually used to identify specific applications or protocols involved in traffic data. By analyzing the port number, the protocol used in the network can be identified. For example, port 80 is usually used for HTTP protocol, port 443 is used for HTTPS, and port 21 is used for FTP. IP address is a unique address used to identify a computer or network device, including source IP address and destination IP address. By analyzing IP addresses, the transmission path of network traffic and the devices involved can be understood, and potential attacks can be discovered by detecting abnormal IP address patterns.

[0030] Statistical features involve analyzing network traffic data and extracting some digital features that reflect traffic behavior, including traffic size, number of packets, duration, flow rate, etc. These statistical features can help analyze traffic behavior patterns, capacity, and frequency. For example, concentrated traffic during peak hours or excessive traffic fluctuations may be abnormal. Content features involve analyzing the payload content of network packets. In many protocols, the content of the packet carries protocol-specific information. Analyzing these contents helps identify protocol types. For example, HTTP requests usually contain specific header fields, and DNS requests contain query fields.

[0031] Furthermore, an initial knowledge graph is constructed in combination with the multidimensional feature parameter set, and a graph dimension reduction mechanism is introduced to reduce the dimension of the initial knowledge graph to obtain a target knowledge graph, including: The graph similarity probability distribution is set based on the graph dimensionality reduction mechanism; the initial knowledge graph is reduced in dimension according to the graph similarity probability distribution to obtain the target knowledge graph; wherein the expression of the graph similarity probability distribution is: ; in, It is used to characterize the difference between the initial knowledge graph and the target knowledge graph. The initial knowledge graph is subjected to dimensionality reduction processing with the constraint of tending to 0. Characterize the divergence of the similarity between the initial knowledge graph and the target knowledge graph, Representing the nodes in the initial knowledge graph With Node The similarity probability of Representing the nodes in the target knowledge graph With Node The similarity probability.

[0032] The purpose of graph dimensionality reduction is to convert the high-dimensional features in the initial knowledge graph into a low-dimensional representation, while trying to maintain the similarity and connection relationship between the graph nodes. By reducing the dimensionality, the graph can be made more concise and easier to analyze in the low-dimensional space. In order to achieve this goal, a similarity probability distribution is pre-set to help decide which nodes or data points should be mapped to similar positions during the dimensionality reduction process.

[0033] According to the set graph similarity probability distribution, the initial knowledge graph is mapped from the high-dimensional space to the low-dimensional space using a dimensionality reduction algorithm to obtain the target knowledge graph. Specifically, dimensionality reduction is achieved by minimizing the difference between the similarity distributions in the high-dimensional space and the low-dimensional space. Specifically, the similarity probability distributions between nodes in the high-dimensional space are calculated, and these similarity probability distributions are mapped to the low-dimensional space, and similar nodes are kept as close as possible in the low-dimensional space, while dissimilar nodes are kept as far away. The target knowledge graph after dimensionality reduction contains the node representations in the low-dimensional space. At this time, the positions of the nodes in the graph reflect the similarity relationship between them, and the structural information in the original graph is retained as much as possible.

[0034] Specifically, the expression of graph similarity probability distribution is: ; Representing nodes in the initial knowledge graph With Node The similarity probability of Representing nodes in the target knowledge graph With Node For each pair of nodes, the formula calculates the KL divergence between them to measure the difference in the probability distribution of the two nodes. Specifically calculate how the initial spectrum deviates from the target spectrum. The difference between the two probability distributions is quantified, and the goal is to minimize the KL divergence, indicating that the target knowledge graph Q should be as similar to the initial knowledge graph P after transformation as possible. In general, this formula expresses that by calculating the KL divergence between each pair of nodes, the overall difference between the initial graph and the target graph is measured. By minimizing this difference, it can ensure that the target knowledge graph is as close to the initial knowledge graph as possible in structure and relationship.

[0035] Furthermore, taking the target knowledge graph as a clustering benchmark, cluster analysis is performed on the network traffic knowledge graph database to obtain a target cluster corresponding to the target knowledge graph, including: Extract any knowledge graph in the network traffic knowledge graph database; calculate any similarity coefficient between the target knowledge graph and the arbitrary knowledge graph; when the arbitrary similarity coefficient is within a predetermined coefficient threshold, form the target cluster based on the target knowledge graph and the arbitrary knowledge graph.

[0036] The network traffic knowledge graph database contains multiple analyzed and processed network traffic graphs. Each graph represents different network traffic characteristics, protocol types or communication modes. Each graph consists of multiple nodes (network traffic data) and edges connecting these nodes (similarity relationships between nodes). Any knowledge graph refers to a graph randomly selected from the network traffic knowledge graph database.

[0037] Calculate any similarity coefficient between the target knowledge graph and any knowledge graph. The similarity coefficient is used to quantify the degree of similarity between two knowledge graphs. The similarity coefficient calculation methods include cosine similarity and Euclidean distance. Exemplarily, the similarity is measured by calculating the cosine value between the vector representations of nodes and edges in the graph. The value range of cosine similarity is from -1 to +1, where -1 means completely different and +1 means completely different.

[0038] The predetermined coefficient threshold is a pre-set similarity standard used to determine whether the similarity between the target knowledge graph and any knowledge graph is high enough to classify them into the same cluster. This threshold is usually adjusted in experiments according to specific problems. For example, if the similarity between the target graph and any graph is higher than 0.8, they are considered to belong to the same cluster.

[0039] If the similarity between the target graph and any graph is higher than a predetermined threshold, the two graphs can be classified into the same cluster to form a target cluster. Furthermore, whenever a new graph is found to be similar to the target graph, the graph will also be included to form a larger target cluster. The target cluster represents a set of network traffic data with similar characteristics or protocols.

[0040] Further, a predetermined clustering degree evaluation strategy is read, and clustering degree evaluation analysis is performed on the target cluster and the target traffic data according to the predetermined clustering degree evaluation strategy to obtain a target clustering degree, including: Obtain a first knowledge graph corresponding to the first traffic data in the target cluster; perform eigendecomposition on a first Laplace matrix of the first knowledge graph to obtain a first decomposition result; construct a cluster eigenvalue scatter plot based on a first correspondence between a first eigenvalue in the first decomposition result and the first traffic data; perform spline curve processing on the cluster eigenvalue scatter plot to obtain a cluster spline curve; obtain a target eigenvalue of the target knowledge graph, and calculate the distance from the target eigenvalue to the cluster spline curve to obtain a target distance; normalize the target distance to obtain the target clustering degree.

[0041] The first traffic data refers to any traffic data in the target cluster except the target traffic data. This traffic data is other data samples in the cluster. Each traffic data can be constructed into a corresponding knowledge graph through a predetermined feature dimension, where nodes represent traffic features and edges represent similarities between features. The first knowledge graph corresponding to the first traffic data is extracted for subsequent analysis.

[0042] The Laplace matrix is ​​a basic matrix in graph theory, which reflects the connection relationship between nodes in the graph. In graph analysis, the Laplace matrix is ​​often used for eigenvalue decomposition of the graph to extract the main features of the graph structure. The construction process of the Laplace matrix is ​​as follows: Suppose there is an undirected graph G, whose adjacency matrix is ​​A and degree matrix is ​​D. The Laplace matrix L is defined as L=DA, where D is a diagonal matrix whose diagonal elements represent the degree of the node, and A is the adjacency matrix, which represents the connection relationship between the nodes. For a knowledge graph, the nodes represent the traffic characteristics and the edges represent the similarity between the nodes. The Laplace matrix can reflect the connection strength between these nodes and the overall structure of the graph.

[0043] By performing eigendecomposition on the first Laplacian matrix of the first knowledge graph, the eigenvalues ​​and eigenvectors of the graph can be obtained as the first decomposition results. These eigenvalues ​​and eigenvectors are used to reflect the relationship between the nodes in the first knowledge graph and the structural characteristics of the graph.

[0044] The first eigenvalue obtained from the feature decomposition represents some important properties of the graph structure, such as the connection strength between nodes, the compactness of the graph, etc. By analyzing the relationship between these eigenvalues ​​and the first traffic data, we can further understand the distribution of these traffic data in the graph. The cluster eigenvalue scatter plot visualizes the eigenvalue of each traffic data as the coordinate in the graph. The horizontal axis of the scatter plot can be the first eigenvalue, and the vertical axis is other eigenvalues ​​related to it or other attributes of the traffic data such as traffic size, duration, etc., which is used to intuitively display the characteristic distribution of nodes in the graph and the structure of the cluster.

[0045] Spline curve is a method of fitting a smooth curve through a set of data points. It usually uses piecewise polynomials to fit the relationship between data points. Common spline curves include cubic splines, which can ensure smooth transitions between data points. Specifically, each data point in the cluster feature value scatter plot is used as an input data point, and a spline curve fitting method, such as cubic spline interpolation, is used to connect these data points into a smooth curve to obtain a cluster spline curve. One feature of the spline curve is that it can not only smoothly fit the relationship between data points, but also maintain a low risk of overfitting.

[0046] The target Laplacian matrix of the target knowledge graph is subjected to eigendecomposition using the same method as the first eigenvalue to obtain a target decomposition result, and the target eigenvalue of the target knowledge graph is extracted from the target decomposition result. The distance from the target eigenvalue to the cluster spline curve is calculated, for example, the Euclidean distance method is used to calculate the straight-line distance between the target eigenvalue and the cluster spline curve to obtain the target distance, which indicates the degree of difference between the target eigenvalue and the cluster.

[0047] In order to make the target distance value suitable for further analysis and be able to be compared with other clustering results, the target distance needs to be normalized. The purpose of normalization is to limit the range of the target distance to a specific interval, usually between 0 and 1. The common normalization method is to standardize the target distance according to its maximum value, or use the minimum and maximum normalization method. After normalization, the target clustering degree is obtained, which is used to measure the similarity between the target knowledge graph and the target cluster. The closer the target clustering degree value is to 1, the higher the similarity between the target graph and the cluster, and the better the clustering quality; the closer the value is to 0, the poorer the clustering effect.

[0048] Furthermore, the first Laplacian matrix of the first knowledge graph is subjected to eigendecomposition to obtain a first decomposition result, including: Convert the first knowledge graph into a first undirected graph; obtain a first degree matrix and a first adjacency matrix of the first undirected graph respectively; and take the difference between the first degree matrix and the first adjacency matrix as the first Laplacian matrix.

[0049] The knowledge graph is a graph structure composed of nodes and edges. The purpose of converting the first knowledge graph into the first undirected graph is to analyze it using the methods in graph theory. An undirected graph is a graph composed of a set of nodes and undirected edges connecting these nodes. Unlike a directed graph, the edges in an undirected graph have no directionality, that is, the edges represent the symmetrical relationship between nodes. Specifically, each node in the knowledge graph represents an entity or a feature. In an undirected graph, these nodes remain unchanged as vertices in the graph. In a knowledge graph, edges represent the relationship between nodes and are usually directional, for example, from one node to another. In the process of converting directed edges into undirected edges, the directionality of the edges is removed, indicating that the relationship between nodes is symmetrical.

[0050] Get the first degree matrix and the first adjacency matrix of the first undirected graph respectively, where the degree matrix is ​​a diagonal matrix, which indicates the degree of each node in the undirected graph. The degree of a node refers to the number of edges connected to the node. For an undirected graph, the degree matrix is ​​an N×N diagonal matrix, where N is the total number of nodes. The elements on the diagonal indicate the degree of the node, that is, the number of edges connected to the node, and the elements in other positions are 0. The adjacency matrix is ​​an N×N matrix, which indicates the connection relationship between nodes in an undirected graph. For an undirected graph, the adjacency matrix is ​​a symmetric matrix, where each element indicates whether there is an edge between two nodes, and each element is 0 or 1. If there is an edge connecting the two nodes, the corresponding element is 1, otherwise it is 0.

[0051] The Laplace matrix is ​​a basic matrix in graph theory. It reflects the relationship between the nodes of a graph and the structure of the graph. The first Laplace matrix is ​​defined by the difference between the first degree matrix and the first adjacency matrix. Through the first Laplace matrix, we can further analyze the structural characteristics of the graph, such as the connectivity of the graph, the balance of the graph, and the similarity between nodes.

[0052] Furthermore, the cluster feature value scatter plot is processed into a spline curve to obtain a cluster spline curve, including: A first scatter point group is formed based on the cluster feature value scatter plot; the first scatter point group is processed into a spline curve to obtain a first spline curve; the cluster feature value scatter plot is compared with the first scatter point group to obtain a first verification group; a first verification result of the first verification group on the first spline curve is obtained; when the first verification result meets a predetermined curvilinear constraint, the first spline curve is recorded as the cluster spline curve.

[0053] The first scatter point group is a subset extracted from the cluster feature value scatter plot. By selecting and filtering the scatter points in the cluster feature value scatter plot, a part of the scatter points is selected as the first scatter point group. These scatter points are used for further analysis. When selecting scatter points, they are screened based on their distribution in the scatter plot, density or similarity with other scatter points in the cluster to ensure that the selected scatter points are representative.

[0054] Spline curve is a technique for fitting a smooth curve through a series of control points. Through spline processing, a smooth fitting curve can be provided for the discrete data points in the first scattered point group, helping to reveal the overall trend of the data. Specifically, using techniques such as cubic spline interpolation, the discrete data points in the first scattered point group are fitted into a smooth curve. Cubic spline interpolation ensures a smooth transition of the curve between the data points by constructing a polynomial between each adjacent scattered point. The first spline curve obtained through spline processing can more clearly show the distribution trend of the characteristics in the cluster, simplify data analysis, and make subsequent comparison and verification more intuitive.

[0055] Compare the cluster feature value scatter plot with the first scatter point group to verify whether the characteristics of the cluster are consistent with the spline curve fitted by the scatter point group, including evaluating whether the first scatter point group can fit the cluster feature value scatter plot well, or whether they show similar distribution trends; evaluate the accuracy of the fit by calculating the distance between the scatter points and the spline curve, such as the least squares error. By comparison, a first verification group is obtained, which contains a group of data points that conform to the distribution of the cluster feature values ​​and are similar to the pattern of the cluster feature value scatter plot in the feature space.

[0056] Verify whether the first spline curve can effectively fit the data points in the first verification group, which means that it is hoped that the distance between these data points and the spline curve is as small as possible, and the curve can reflect the trend of the data points. Exemplarily, the sum of squared errors between each data point in the first verification group and the spline curve is calculated to obtain an error index. If the error is small, it indicates that the spline curve fits well. The first verification result is used to determine whether the spline curve can reasonably fit the data features in the cluster.

[0057] The predetermined curvilinear constraint is an expected condition set for the fitting result of the spline curve, for example, a predetermined error range is set. When the errors between the data points of all the validation groups and the spline curve are within the range, the fitting is considered to be valid. If the first validation result meets the predetermined curvilinear constraint, that is, the error of the spline curve is less than the predetermined error range, and meets the constraints such as the smoothness requirement and the error limit, it means that the spline curve successfully describes the characteristics of the cluster. At this time, the first spline curve is confirmed as a cluster spline curve and is used as the representative feature of the cluster.

[0058] Furthermore, if the target clustering degree reaches a predetermined clustering degree limit, the target protocol type corresponding to the target cluster is recorded as the protocol type of the target traffic data, including: Obtain the first protocol type corresponding to the first traffic data, and count the first total number of the first protocol type in the target cluster; descend the first protocol type based on the first total number to obtain a descending list of cluster protocol types; take the first protocol type in the descending list of cluster protocol types as the target protocol type.

[0059] In the network traffic data, each traffic data corresponds to a protocol type, such as HTTP, FTP, DNS, etc. The corresponding first protocol type is obtained from the first traffic data to indicate the protocol type to which this traffic data belongs. The first total number of the first protocol type is counted in the target cluster, that is, by traversing all traffic data in the target cluster, the number of traffic data with the same protocol type as the first traffic data is counted. For example, if the protocol type of the first traffic data is HTTP, the number of all traffic data with the protocol type of HTTP in the target cluster needs to be counted.

[0060] Taking the first total number as the descending basis, sort the number of occurrences of each protocol type in the target cluster to obtain a descending list of cluster protocol types. Common sorting algorithms, such as quick sort, merge sort, etc., can be used to sort the number of occurrences of protocol types in descending order to ensure that the most frequently occurring protocol type is at the front.

[0061] According to the descending list of cluster protocol types obtained by sorting, the protocol type ranked first is selected as the target protocol type. This is the type with the largest number of occurrences, indicating that the protocol type is the most representative in the target cluster.

[0062] Further, reading a predetermined feature dimension, and collecting and obtaining a multi-dimensional feature parameter set of the target flow data based on the predetermined feature dimension, includes: Obtaining a predetermined weight distribution of the predetermined feature dimension; and adjusting the multi-dimensional feature parameter set based on the predetermined weight distribution.

[0063] The predetermined feature dimensions include at least port features, IP address features, statistical features and content features. Each feature dimension represents an attribute in the network traffic and is used to describe the nature and behavior of the traffic. The predetermined weight allocation refers to allocating a weight to each feature dimension in the predetermined feature dimension. The weight indicates the importance of the feature in the overall analysis. For example, feature selection algorithms such as information gain, chi-square test, etc. are used to evaluate the contribution of each feature dimension and then determine the weight of each feature.

[0064] The predetermined weight distribution is applied to the multidimensional feature parameter set for adjustment. The adjustment method is usually to multiply each eigenvalue by its corresponding weight so that the contribution of each feature in the feature vector matches its importance in the analysis. By adjusting the multidimensional feature parameter set, it can be effectively ensured that the importance of the feature is correctly reflected in the data, thereby producing more accurate results in subsequent analysis processes such as clustering, classification or protocol identification.

[0065] In summary, the network traffic identification method of an unknown cryptographic protocol provided by the embodiment of the present application has the following technical effects: By preprocessing the captured dynamic traffic data, redundant data, noise and invalid information can be effectively removed, making subsequent analysis and identification work more accurate and efficient. The target traffic data obtained after preprocessing is purer, which improves the accuracy of subsequent feature extraction and protocol identification. By reading the predetermined feature dimensions and extracting the multidimensional feature parameter set of the target traffic data based on these dimensions, the multidimensional information of the network traffic can be captured comprehensively and finely. Different feature dimensions provide a more comprehensive basis for protocol identification and enhance the richness and accuracy of feature expression. By combining the multidimensional feature parameter set to construct the initial knowledge graph and introducing the graph dimensionality reduction mechanism, high-dimensional data can be converted into a graph representation in a low-dimensional space. After dimensionality reduction, the target knowledge graph is convenient for further processing and analysis, while effectively reducing the computational complexity. The graph after dimensionality reduction retains important structural information and improves The efficiency and accuracy of subsequent analysis are improved; by taking the target knowledge graph as the clustering benchmark and performing cluster analysis on the network traffic knowledge graph database, the clustering pattern of similar traffic data can be identified. The target clusters after clustering can clearly identify different protocols or communication modes, thereby providing a clear classification basis for protocol identification; by using a predetermined clustering degree evaluation strategy, the clustering degree of the target clusters and the target traffic data is evaluated, which can quantify the degree of clustering. The evaluation results help confirm whether the clustering effect meets the predetermined standards, further improving the reliability and accuracy of the clustering analysis; if the clustering degree of the target cluster reaches the predetermined limit, the protocol type corresponding to the cluster can be accurately used as the protocol type of the target traffic data. This method ensures the automatic identification of protocol types, improves the accuracy in network traffic analysis, and can effectively identify unknown cryptographic protocol types.

[0066] Embodiment 2 is based on the same inventive concept as the method for identifying network traffic of an unknown cryptographic protocol in the above embodiment. Figure 2 As shown, an embodiment of the present application provides a network traffic identification system for an unknown cryptographic protocol, the system comprising: The data preprocessing module 10 is used to preprocess the captured dynamic flow data to obtain target flow data.

[0067] The characteristic parameter acquisition module 20 is used to read the predetermined characteristic dimension and collect and obtain a multi-dimensional characteristic parameter set of the target flow data based on the predetermined characteristic dimension.

[0068] The dimensionality reduction processing module 30 is used to construct an initial knowledge graph in combination with the multi-dimensional feature parameter set, and introduce a graph dimensionality reduction mechanism to perform dimensionality reduction processing on the initial knowledge graph to obtain a target knowledge graph.

[0069] The clustering analysis module 40 is used to perform clustering analysis on the network traffic knowledge graph database based on the target knowledge graph as a clustering benchmark to obtain a target cluster corresponding to the target knowledge graph.

[0070] The evaluation and analysis module 50 is used to read a predetermined clustering degree evaluation strategy, and perform clustering degree evaluation and analysis on the target cluster and the target traffic data according to the predetermined clustering degree evaluation strategy to obtain a target clustering degree.

[0071] The protocol type determination module 60 is configured to record the target protocol type corresponding to the target cluster as the protocol type of the target traffic data if the target clustering degree reaches a predetermined clustering degree limit.

[0072] Furthermore, the predetermined feature dimensions include at least port features, IP address features, statistical features and content features.

[0073] Furthermore, the dimension reduction processing module 30 is used to perform the following operation steps: The graph similarity probability distribution is set based on the graph dimensionality reduction mechanism; the initial knowledge graph is reduced in dimension according to the graph similarity probability distribution to obtain the target knowledge graph; wherein the expression of the graph similarity probability distribution is: ; in, It is used to characterize the difference between the initial knowledge graph and the target knowledge graph. The initial knowledge graph is subjected to dimensionality reduction processing with the constraint of tending to 0. Characterize the divergence of the similarity between the initial knowledge graph and the target knowledge graph, Representing the nodes in the initial knowledge graph With Node The similarity probability of Representing the nodes in the target knowledge graph With Node The similarity probability.

[0074] Furthermore, the cluster analysis module 40 is used to perform the following operation steps: Extract any knowledge graph in the network traffic knowledge graph database; calculate any similarity coefficient between the target knowledge graph and the arbitrary knowledge graph; when the arbitrary similarity coefficient is within a predetermined coefficient threshold, form the target cluster based on the target knowledge graph and the arbitrary knowledge graph.

[0075] Furthermore, the evaluation and analysis module 50 is used to perform the following operation steps: Obtain a first knowledge graph corresponding to the first traffic data in the target cluster; perform eigendecomposition on a first Laplace matrix of the first knowledge graph to obtain a first decomposition result; construct a cluster eigenvalue scatter plot based on a first correspondence between a first eigenvalue in the first decomposition result and the first traffic data; perform spline curve processing on the cluster eigenvalue scatter plot to obtain a cluster spline curve; obtain a target eigenvalue of the target knowledge graph, and calculate the distance from the target eigenvalue to the cluster spline curve to obtain a target distance; normalize the target distance to obtain the target clustering degree.

[0076] Furthermore, the evaluation and analysis module 50 is used to perform the following operation steps: Convert the first knowledge graph into a first undirected graph; obtain a first degree matrix and a first adjacency matrix of the first undirected graph respectively; and take the difference between the first degree matrix and the first adjacency matrix as the first Laplacian matrix.

[0077] Furthermore, the evaluation and analysis module 50 is used to perform the following operation steps: A first scatter point group is formed based on the cluster feature value scatter plot; the first scatter point group is processed into a spline curve to obtain a first spline curve; the cluster feature value scatter plot is compared with the first scatter point group to obtain a first verification group; a first verification result of the first verification group on the first spline curve is obtained; when the first verification result meets a predetermined curvilinear constraint, the first spline curve is recorded as the cluster spline curve.

[0078] Furthermore, the protocol type determination module 60 is used to perform the following operation steps: Obtain the first protocol type corresponding to the first traffic data, and count the first total number of the first protocol type in the target cluster; descend the first protocol type based on the first total number to obtain a descending list of cluster protocol types; take the first protocol type in the descending list of cluster protocol types as the target protocol type.

[0079] Furthermore, the characteristic parameter acquisition module 20 is used to perform the following operation steps: Obtaining a predetermined weight distribution of the predetermined feature dimension; and adjusting the multi-dimensional feature parameter set based on the predetermined weight distribution.

[0080] Through the above-mentioned detailed description of the network traffic identification method of an unknown cryptographic protocol in this specification, technical personnel in this field can clearly understand the network traffic identification system of an unknown cryptographic protocol in this embodiment. Since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part description.

[0081] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to the embodiments shown herein, but will conform to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for identifying network traffic of an unknown cryptographic protocol, characterized in that: The method comprises: Preprocess the captured dynamic traffic data to obtain target traffic data; Reading a predetermined characteristic dimension, and collecting and obtaining a multi-dimensional characteristic parameter set of the target flow data based on the predetermined characteristic dimension; An initial knowledge graph is constructed in combination with the multidimensional feature parameter set, and a graph dimension reduction mechanism is introduced to reduce the dimension of the initial knowledge graph to obtain a target knowledge graph; Taking the target knowledge graph as a clustering benchmark, cluster analysis is performed on the network traffic knowledge graph database to obtain a target cluster corresponding to the target knowledge graph; Reading a predetermined clustering degree evaluation strategy, and performing clustering degree evaluation analysis on the target cluster and the target traffic data according to the predetermined clustering degree evaluation strategy to obtain a target clustering degree; If the target clustering degree reaches a predetermined clustering degree limit, the target protocol type corresponding to the target cluster is recorded as the protocol type of the target traffic data.

2. According to the method for identifying network traffic of an unknown cryptographic protocol according to claim 1, it is characterized in that: The predetermined feature dimensions include at least port features, IP address features, statistical features and content features.

3. According to the method for identifying network traffic of unknown cryptographic protocol in claim 1, it is characterized in that: An initial knowledge graph is constructed in combination with the multidimensional feature parameter set, and a graph dimension reduction mechanism is introduced to reduce the dimension of the initial knowledge graph to obtain a target knowledge graph, including: Setting a graph similarity probability distribution based on the graph dimension reduction mechanism; According to the graph similarity probability distribution, the initial knowledge graph is reduced in dimension to obtain the target knowledge graph; Among them, the expression of the probability distribution of the graph similarity is: ; in, It is used to characterize the difference between the initial knowledge graph and the target knowledge graph. The initial knowledge graph is subjected to dimensionality reduction processing with the constraint of tending to 0. Characterize the divergence of the similarity between the initial knowledge graph and the target knowledge graph, Representing the nodes in the initial knowledge graph With Node The similarity probability of Representing the nodes in the target knowledge graph With Node The similarity probability.

4. According to the method for identifying network traffic of unknown cryptographic protocol in claim 1, it is characterized in that: Taking the target knowledge graph as the clustering benchmark, cluster analysis is performed on the network traffic knowledge graph database to obtain the target cluster corresponding to the target knowledge graph, including: Extract any knowledge graph in the network traffic knowledge graph database; Calculate any similarity coefficient between the target knowledge graph and any knowledge graph; When the arbitrary similarity coefficient is within a predetermined coefficient threshold, the target cluster is formed based on the target knowledge graph and the arbitrary knowledge graph.

5. According to the method for identifying network traffic of unknown cryptographic protocol as described in claim 4, it is characterized in that: Reading a predetermined clustering degree evaluation strategy, and performing clustering degree evaluation analysis on the target cluster and the target traffic data according to the predetermined clustering degree evaluation strategy to obtain a target clustering degree, including: Obtaining a first knowledge graph corresponding to the first traffic data in the target cluster; Performing eigendecomposition on a first Laplacian matrix of the first knowledge graph to obtain a first decomposition result; Constructing a cluster eigenvalue scatter plot according to a first corresponding relationship between a first eigenvalue in the first decomposition result and the first traffic data; Performing spline curve processing on the cluster feature value scatter plot to obtain a cluster spline curve; Obtaining a target feature value of the target knowledge graph, and calculating a distance from the target feature value to the cluster spline curve to obtain a target distance; The target distance is normalized to obtain the target clustering degree.

6. According to the method for identifying network traffic of unknown cryptographic protocol according to claim 5, it is characterized in that: Performing eigendecomposition on the first Laplacian matrix of the first knowledge graph to obtain a first decomposition result, including: Converting the first knowledge graph into a first undirected graph; Respectively obtain a first degree matrix and a first adjacency matrix of the first undirected graph; The difference between the first degree matrix and the first adjacency matrix is ​​taken as the first Laplacian matrix.

7. According to the method for identifying network traffic of unknown cryptographic protocol in claim 5, it is characterized in that: The cluster characteristic value scatter plot is processed into a spline curve to obtain a cluster spline curve, including: Forming a first scatter point group based on the cluster feature value scatter plot; Performing spline curve processing on the first scattered point group to obtain a first spline curve; Comparing the cluster feature value scatter plot with the first scatter point group to obtain a first verification group; Obtaining a first verification result of the first spline curve by the first verification group; When the first verification result meets the predetermined curvilinear constraint, the first spline curve is recorded as the cluster spline curve.

8. According to the method of claim 5, the method is characterized in that: If the target clustering degree reaches a predetermined clustering degree limit, the target protocol type corresponding to the target cluster is recorded as the protocol type of the target traffic data, including: Acquire a first protocol type corresponding to the first traffic data, and count a first total number of the first protocol type in the target cluster; The first protocol types are sorted in descending order based on the first total number to obtain a cluster protocol type descending list; The first protocol type in the descending list of the cluster protocol types is taken as the target protocol type.

9. The method for identifying network traffic of an unknown cryptographic protocol according to claim 1, characterized in that: Reading a predetermined feature dimension, and collecting and obtaining a multi-dimensional feature parameter set of the target flow data based on the predetermined feature dimension, including: Obtaining a predetermined weight distribution of the predetermined feature dimension; The multi-dimensional feature parameter set is adjusted based on the predetermined weight distribution.

10. A network traffic identification system for unknown cryptographic protocols, characterized in that: A method for identifying network traffic of an unknown cryptographic protocol according to any one of claims 1 to 9, the system comprising: The data preprocessing module is used to preprocess the captured dynamic flow data to obtain the target flow data; A characteristic parameter acquisition module, used for reading a predetermined characteristic dimension, and collecting and obtaining a multi-dimensional characteristic parameter set of the target flow data based on the predetermined characteristic dimension; A dimensionality reduction processing module is used to construct an initial knowledge graph in combination with the multi-dimensional feature parameter set, and introduce a graph dimensionality reduction mechanism to perform dimensionality reduction processing on the initial knowledge graph to obtain a target knowledge graph; A clustering analysis module is used to perform clustering analysis on the network traffic knowledge graph database using the target knowledge graph as a clustering benchmark to obtain a target cluster corresponding to the target knowledge graph; An evaluation and analysis module, used for reading a predetermined clustering degree evaluation strategy, and performing clustering degree evaluation and analysis on the target cluster and the target traffic data according to the predetermined clustering degree evaluation strategy to obtain a target clustering degree; The protocol type determination module is used to record the target protocol type corresponding to the target cluster as the protocol type of the target traffic data if the target clustering degree reaches a predetermined clustering degree limit.

Citation Information

Patent Citations

  • Traffic recognition method, traffic determination method and knowledge graph establishment method

    CN112822121A

  • Knowledge graph-based domestic and overseas electricity market research hotspot tracking method

    CN113609303A

  • Clustering result processing method and device, electronic equipment and storage medium

    CN116776171A

  • Wideband satellite network abnormal behavior detection method based on knowledge graph

    CN117113994A

  • Knowledge graph expansion method, system and equipment based on knowledge disambiguation and medium

    CN118333152A

Cited By

  • Encryption protocol reasoning method and system based on graph driving

    CN120856473A

  • Encrypted anonymous network traffic analysis and identification method based on traffic reconstruction

    CN120935287A