Message Type Recognition Method and Device Based on Data Mining
Through a data mining-based method, the continuous sequence pattern and factor graph model are used to identify keyword fields in the binary network protocol, which solves the problems of low alignment quality and high complexity in the prior art, and achieves efficient and accurate message type recognition.
Patent Information
- Application Number
- CN202111674303.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-31
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2041-12-31
AI Technical Summary
The prior art has problems of lower alignment quality and high complexity of sequence alignment algorithms in reverse analysis of binary network protocols, and the failure to effectively consider semantic constraints, resulting in insufficient efficiency and accuracy of identifying message types in large-scale data.
Using a data mining method, frequent continuous subsequences are generated through a continuous sequence mode algorithm, and the position-related candidate keyword fields are generated using the key continuous sequence mode algorithm, and the probability that the candidate keyword field becomes a keyword is calculated through the factor graph model. Finally, the candidate keyword field with the greatest probability is selected to determine the message type.
Quickly identify keyword fields with dynamic length and position under the condition of misalignment of messages, improving the accuracy and efficiency of message type recognition, and reducing time and space complexity.
Smart Images

Figure CN114417857B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of message type recognition, and particularly relates to a method and device for message type recognition based on data mining. Background Art
[0002] Network protocols can be briefly divided into two categories: text protocols and binary protocols. Text protocols often contain delimiters and keywords recognizable by natural language processing, while most binary network protocols lack these structural features required for natural language processing. Therefore, current protocol reverse research based on network traffic mainly focuses on binary protocols.
[0003] The similarity-based method is the most commonly used technique for binary protocol reverse analysis, and most of these methods are derived from bioinformatics algorithms. When analyzing network protocols using the similarity-based method, messages are usually represented as sequences of bytes or other basic units. First, a sequence comparison algorithm is used to align the messages and calculate similarity scores. Then, the messages are clustered based on the similarity scores. Finally, the message type is determined by analyzing the commonalities of the messages within the cluster and the message format is summarized. However, this method has the following two disadvantages: (1) The diversity of message content reduces the alignment quality. On the one hand, it is impossible to ensure the diversity of the collected samples. On the other hand, the same fields of messages of the same type may have very different values, and different types of messages may also have some common fields with the same values; (2) The applicability of the sequence alignment algorithm cannot be guaranteed. Precise sequence alignment cannot be used for large-scale data due to its exponential complexity, and widely used recursive clustering algorithms such as UPGMA are derived from bioinformatics and cannot be fully applied to protocol analysis. For example, in bioinformatics, UPGMA uses a phylogenetic tree to reflect the evolution of genomic sequences, but there is no such evolutionary relationship in protocol messages.
[0004] In fact, after receiving a message, the server and the client only determine the message type through keywords. Therefore, if the fields representing the keywords can be inferred, then the message type is clear. However, it is not easy to find keywords from dense binary data, and many methods can only indirectly determine the message type through keyword-related sequences. The only method that attempts to determine the message type through a keyword field is NETPLIER.
[0005] NETPLIER first aligns the messages from the client and server sides using the multiple sequence alignment (MSA) algorithm, then divides the messages into field sequences and differentiates fixed, dynamic, and variable-length fields. A random variable is introduced for each dynamic field to represent the probability that the field becomes a keyword. Assuming a dynamic field is a keyword, the messages can be divided into different clusters according to the values of this field, and these clusters will satisfy some constraint relationships. NETPLIER uses Message Similarity Constraints, Remote Coupling Constraints, Structure Coherence Constraints, and Dimension Constraints to construct probability constraint relationships. Finally, probability inference is performed to derive the posterior probability of the random variable, and the one with the highest probability among all candidate fields is selected as the keyword to identify the message. The experimental results of NETPLIER are better than other similar technologies. However, it still has the following deficiencies: (1) The practice of using sequence alignment to generate candidate keywords brings huge time and space complexity; (2) The constraint relationships do not consider semantic constraints; (3) The inferred probability setting tends to generate fewer clusters and is not applicable in large-scale messages. Summary of the Invention
[0006] In view of the defects existing in the prior art, the present invention proposes a method and device for identifying message types based on data mining, which uses data mining to quickly determine candidate keyword fields and improves the probability constraint relationship, can identify message types in a short time with high accuracy, and solves the problem of extracting keyword fields with dynamic length and position without aligning messages.
[0007] To achieve the above object, the present invention adopts the following technical solutions:
[0008] The present invention provides a method for identifying message types based on data mining, comprising the following steps:
[0009] Generate frequent continuous subsequences for the message sequence using the continuous sequence pattern algorithm;
[0010] Generate position-related candidate keyword fields on the selected frequent continuous subsequences through the key continuous sequence pattern algorithm;
[0011] Calculate the probability that the candidate keyword fields become keywords based on the factor graph model;
[0012] Select the candidate keyword field with the highest probability as the keyword to determine the message type.
[0013] Further, the continuous sequence pattern algorithm specifically includes the following steps:
[0014] First, extract subsequences of length 1 basic length from all sequences and store them in the subsequence set;
[0015] Then, calculate frequent continuous subsequences that reach the minimum support in subsequences of length l - 1 basic length, where the support is defined as the number of sequences containing the target subsequence, and use the frequent continuous subsequences of length l - 1 to generate continuous subsequences of length l according to the Apriori strategy, and iterate this step until no new continuous subsequences can be extracted;
[0016] Finally, take the union of all frequent continuous subsequences, delete the subsequences included in other sequences in the set, and return them after sorting in descending order of support.
[0017] Further, the expression for the set of candidate keyword fields composed of multiple candidate keyword fields is as follows:
[0018] KeySeqSet = {K 1 , K 2 ,..., K n}
[0019] where n represents the number of candidate keyword fields;
[0020] Then, the structural expression of the candidate keyword field K i is as follows:
[0021] K i = [MinOffset, MaxDepth, ModeOffset, ModeLen, kw]
[0022] where MinOffset represents the minimum starting position of the current frequent continuous subsequence fs, MaxDepth represents the maximum ending position of the current frequent continuous subsequence fs, ModeOffset represents the minimum mode of the starting positions of all frequent continuous subsequences within kw, ModeLen represents the minimum mode of the lengths of all frequent continuous subsequences within kw, and kw is the set of all frequent continuous subsequences within the range from MinOffset to MaxDepth in the current frequent continuous subsequence fs and the message sequence set after removing the messages containing fs. Therefore, by setting this structure, it is ensured that the key continuous sequence pattern algorithm can identify keyword fields with unfixed positions.
[0023] Furthermore, when performing message clustering subsequently, first search for frequent consecutive subsequences in kw in descending order of support within the range from MinOffset to MaxDepth. If any are found, identify the category of the message as that frequent consecutive subsequence; otherwise, intercept the sequence within the range from MinOffset to MaxDepth according to ModeOffset and ModeLen for category identification.
[0024] Furthermore, generating position-related candidate keyword fields on the selected frequent consecutive subsequences through the key consecutive sequence pattern algorithm specifically includes:
[0025] First, select subsequences in the frequent consecutive subsequence set that meet the following three conditions: (1) The standard deviation of the subsequence position is less than a preset value, indicating that the position variation range of the subsequence is not large; (2) The support of the subsequence is not 1, indicating that the subsequence does not appear in all messages; (3) The subsequence does not exist in an existing candidate keyword field to prevent duplicate calculations;
[0026] Then create a set of message sequences that do not contain subsequences meeting the above conditions and truncate these message sequences according to the minimum start position and maximum end position of the current frequent consecutive subsequence;
[0027] Finally, run the consecutive sequence pattern algorithm on the newly created message sequences. If new frequent consecutive subsequences are obtained, combine the results obtained by the consecutive sequence pattern algorithm with the current frequent consecutive subsequence to form a candidate keyword field set and save the information according to the candidate keyword field structure.
[0028] Furthermore, the factor graph model models the uncertainty in keyword recognition as the joint distribution of observed values and a set of random variables. Each random variable represents the probability that a candidate keyword field is a message keyword, and the observed value is the relevant probability of the four constraint relationships; the four constraint relationships include: (1) Message similarity constraint M(f, c), defined as messages within the same cluster are as similar as possible, and messages between different clusters are as separated as possible; (2) Semantic consistency constraint L(f, c), defined as messages within the same cluster have similar semantic information; (3) Remote coupling constraint R(f, c), defined as the request or response messages corresponding to messages within the same cluster belong to the same cluster; (4) Dimension constraint D(f), defined as the number of clusters divided according to the values of keyword fields is not very large and each cluster contains a sufficient number of messages; each constraint relationship includes an observation constraint and an inference constraint. The observation constraint refers to whether the constraint relationship itself holds, and the inference constraint refers to the inference relationship between the constraint relationship and the random variable.
[0029] Further, the random variable is denoted as K(f), where the field f is a protocol keyword field;
[0030] The expression for the probability of the message similarity constraint M(f, c) holding is as follows:
[0031]
[0032] Let IntraDis denote the average distance between a certain sample and other samples within its cluster, InterDis denote the average distance between a certain sample and samples in other clusters, and n be the number of samples in the cluster. Then the silhouette coefficient sc for a certain sample and the silhouette coefficient SC for a certain cluster are respectively defined as:
[0033]
[0034]
[0035] The expression for the probability of the semantic consistency constraint L(f, c) holding is as follows:
[0036]
[0037] Among them, UnionNum represents the number of different frequent consecutive subsequences among all kws in the set of candidate keyword fields owned by a certain cluster, and InterNum represents the number of frequent consecutive subsequences jointly owned by all messages in the cluster;
[0038] The expression for the probability of the remote coupling constraint R(f, c) holding is as follows:
[0039]
[0040] For a cluster with MsgNum messages, calculate the maximum number of request or response messages belonging to the same cluster in the messages of this cluster, denoted by MaxNum;
[0041] The expression for the probability of the dimension constraint D(f) holding is as follows:
[0042]
[0043] Among them, CluNum represents the number of clusters divided according to the values of keyword fields, MsgsNum represents the number of all messages, and SingleNum represents the number of clusters containing only one message;
[0044] The inference probabilities of the four constraint relationships and the random variables are set according to prior knowledge.
[0045] Further, the process of using the factor graph model for probability calculation includes:
[0046] First, use probability functions to describe each constraint relationship. Let k represent K(f), and x i represents a specific cluster or a certain constraint relationship of the entire clustering result;
[0047] Then, let f j represent a certain probability function. The joint probability function of variables k and the observed values is expressed as the product of the current values of s probability functions divided by the sum of the products of all possible values of the s probability functions. n and s represent the number of constraint relationships and probability functions respectively. The formula for the joint probability function is as follows:
[0048]
[0049] Finally, the probability that the candidate keyword field f becomes a message keyword is expressed as the marginal probability of k, as shown in the following formula:
[0050]
[0051] The present invention also provides a message type recognition device based on data mining, including:
[0052] A frequent continuous subsequence generation module, which is used to generate frequent continuous subsequences for the message sequence using the continuous sequence pattern algorithm;
[0053] A candidate keyword field selection module, which is used to generate position-related candidate keyword fields on the selected frequent continuous subsequences through the key continuous sequence pattern algorithm;
[0054] A probability calculation module, which is used to calculate the probability that the candidate keyword field becomes a keyword based on the factor graph model;
[0055] A message type recognition module, which is used to select the candidate keyword field with the maximum probability as the keyword to determine the message type.
[0056] Compared with the prior art, the present invention has the following advantages:
[0057] The method for identifying message types based on data mining of the present invention does not require message alignment. First, the continuous sequence pattern algorithm is used to generate n-gram frequent continuous subsequences, and then the position-related candidate keyword fields are mined based on the frequent continuous subsequences, solving the problem of extracting keyword fields with dynamic lengths and positions without message alignment. In terms of keyword field judgment, aiming at the deficiency that NETPLIER does not consider semantic constraints, the setting of message similarity constraints is improved, semantic consistency constraints are introduced to calculate probabilities, and finally the factor graph model is used to deduce the probability of each candidate keyword field becoming the final keyword, and the one with the highest probability is selected to determine the message type. Compared with the sequence alignment method, the method for identifying message types based on data mining of the present invention can accurately identify keywords and then determine the message type in a shorter time-consuming situation. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can also obtain other drawings based on these drawings without creative efforts.
[0059] Figure 1 is a flowchart of the method for identifying message types based on data mining according to an embodiment of the present invention;
[0060] Figure 2 is a schematic diagram of the factor graph model according to an embodiment of the present invention;
[0061] Figure 3 is a graph showing the change of time-consuming of four methods in the present invention under different protocol scales;
[0062] Figure 4 is a comparison graph of the running time-consuming of the method of the present invention and NETPLIER on different protocols according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0063] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of the present invention.
[0064] Aiming at the deficiencies of NETPLIER, this embodiment proposes a method for identifying message types based on data mining, as Figure 1 shown, including the following steps:
[0065] Step S11: Use the continuous sequence pattern algorithm to generate frequent continuous subsequences for the message sequence.
[0066] Step S12: Generate position-related candidate keyword fields on the selected frequent continuous subsequences through the key continuous sequence pattern algorithm.
[0067] Step S13: Calculate the probability that the candidate keyword field becomes a keyword based on the factor graph model.
[0068] Step S14: Select the candidate keyword field with the highest probability as the keyword to determine the message type.
[0069] Data mining refers to obtaining knowledge from data. The knowledge that can be obtained includes association rules, classification rules, sequence patterns, clustering rules, generalization rules, and similarity search, etc. Obtaining association rules is the most common application of data mining. The most influential association rule mining algorithm is the Apriori algorithm. The Apriori algorithm uses an iterative method of layer-by-layer search to find the relationships of item sets in the database to form rules. Its process consists of joining and pruning. In this algorithm, the concept of an item set is the set of items. The frequency of an item set refers to the number of transactions containing the item set. If the frequency of an item set meets the minimum support, it is called a frequent item set. The Apriori algorithm is simple and easy to implement, and has spawned many algorithms. These algorithms all follow the Apriori principle: if a set is a frequent item set, then all its subsets are frequent item sets; if a set is not a frequent item set, then all its supersets are not frequent item sets.
[0070] The process of sequence pattern mining is similar to that of association rule mining. The main difference between them is that the purpose of association rule mining is to extract frequently occurring concurrent item sets, while the purpose of sequence pattern mining is to extract frequently occurring time series patterns. Sequence pattern mining algorithms based on the Apriori principle include AprioriAll, AprioriSome, and Generalized Sequential Pattern (GSP). Goo et al. proposed an algorithm suitable for extracting protocol specifications, called the Continuous Sequential Pattern (CSP) algorithm, in order to extract field formats, message formats, and session formats through sequence pattern mining techniques and then extract protocol syntax. The goal of the CSP algorithm is to find frequent continuous common subsequences. Since protocol messages are continuous sequences of fields, this algorithm does not allow gaps in the pattern. To eliminate gaps, this algorithm improved the GSP algorithm with gap constraints so that both its minimum and maximum gaps are zero to avoid gaps in the mined patterns.
[0071] The pseudocode of the CSP algorithm is as follows:
[0072]
[0073]
[0074] In step S11, the continuous sequence pattern algorithm specifically includes the following steps:
[0075] Step S111: Extract subsequences of the basic length (e.g., 1 byte) from all sequences and store them in the subsequence set F 1 (lines 1 - 5).
[0076] Step S112: Calculate the frequent continuous subsequences that reach the minimum support in subsequences of the length of l - 1 basic lengths, and iteratively extract all candidate subsequences of the length of l basic lengths until no new subsequences can be extracted (lines 6 - 16). This iterative process includes two parts. The first part is to calculate the support of the current subsequence and exclude candidate subsequences that do not meet the minimum support threshold (lines 8 - 13). The support is defined as the number of sequences containing the target subsequence, and frequent continuous subsequences are obtained by extracting subsequences with a support higher than the predefined minimum support. The second part is to use the frequent continuous subsequences of the length of l - 1 to generate candidate subsequences of the length of l for the next iteration according to the Apriori strategy (line 14) to accelerate the subsequence generation speed.
[0077] Step S113: Take the union of all frequent continuous subsequences, delete the subsequences included in other sequences within the set, and return them sorted in descending order of support (lines 17 - 19).
[0078] The pseudocode of the Key Continuous Sequence Pattern (KCSP) algorithm is as follows:
[0079]
[0080] Before introducing the key continuous sequence pattern algorithm, it is necessary to first explain the structure of its output candidate keyword field set KeySeqSet. Formula (1) represents a set of candidate keyword fields:
[0081] KeySeqSet = {K 1 , K 2 ,..., K n} (1)
[0082] Among them, n represents the number of candidate keyword fields.
[0083] Formula (2) represents the structure of the candidate keyword field K i :
[0084] K i= [MinOffset, MaxDepth, ModeOffset, ModeLen, kw] (2)
[0085] Among them, MinOffset represents the minimum starting position of the current frequent consecutive subsequence fs, MaxDepth represents the maximum ending position of the current frequent consecutive subsequence fs, ModeOffset represents the minimum mode of the starting positions of all frequent consecutive subsequences within kw, ModeLen represents the minimum mode of the lengths of all frequent consecutive subsequences within kw, and kw is the set of all frequent consecutive subsequences within the range from MinOffset to MaxDepth in the current frequent consecutive subsequence fs and the message sequence set temp after removing the messages containing fs. Since the messages are not aligned in advance, the positions of the keyword fields in the messages may not be the same. Therefore, the present invention sets this structure to ensure that the KCSP algorithm can identify the keyword fields with unfixed positions. When performing message clustering subsequently, first, check whether there is a frequent consecutive subsequence in kw within the range from MinOffset to MaxDepth in descending order of support. If there is, identify the category of the message as that frequent consecutive subsequence; otherwise, intercept the sequence within the range from MinOffset to MaxDepth according to ModeOffset and ModeLen for category identification.
[0086] The key consecutive sequence pattern algorithm in step S12 specifically includes the following steps:
[0087] Step S121, select subsequences (lines 1 - 3) that meet the following three conditions from the set of frequent consecutive subsequences: (1) The standard deviation of the subsequence position is less than the position standard deviation threshold t, indicating that the position variation range of the subsequence is not large; (2) The support of the subsequence is not 1, indicating that the subsequence does not appear in all messages; (3) The subsequence does not exist in an existing candidate keyword field to prevent duplicate calculation.
[0088] Step S122, create a set of message sequences that do not contain the subsequences meeting the above conditions and truncate these message sequences according to the minimum starting position and the maximum ending position of the current frequent consecutive subsequence (lines 4 - 7).
[0089] Step S123, run the consecutive sequence pattern algorithm on the newly created message sequences. If new frequent consecutive subsequences are obtained, combine the results obtained by the consecutive sequence pattern algorithm with the current frequent consecutive subsequence to form a candidate keyword field set and save the information into KeySeqSet according to the candidate keyword field structure (lines 8 - 16).
[0090] The factor graph model in step S13 models the uncertainty in keyword recognition as the joint distribution of observations and a set of random variables. Each random variable represents the probability that a candidate keyword field is a message keyword, and the observations are the probabilities involved in the four constraint relationships. According to the set of candidate keyword fields obtained in step S123, specific constraint relationships are constructed to obtain observations for inferring which field is most likely to be the keyword.
[0091] To determine the keyword fields, NETPLIER differentiates between server-side and client messages and sets the following constraint relationships for the clusters obtained by clustering based on the values of keyword fields according to the following 4 observations: (1) Messages within the same cluster should be more similar than messages in different clusters; (2) There should be a corresponding relationship between the clusters to which client and server messages belong; (3) Messages in the same cluster follow the same field structure; (4) The number of clusters should not be too large, and there should be a sufficient number of messages in each cluster. For observation (1), the present invention modifies it based on the silhouette coefficient, that is, messages within the same cluster should be as similar as possible, and messages between different clusters should be as separated as possible. The specific implementation of observation (3) requires the use of alignment padding inserted during the message alignment process. Since the present invention does not perform message alignment, the structural similarity between messages cannot be calculated. However, the KSCP algorithm can filter out frequent subsequences during operation, and these sequences can be used as semantic information to calculate the semantic similarity between messages. Therefore, the present invention proposes a semantic-based constraint setting: Messages within the same cluster should have similar semantic information.
[0092] The above-mentioned observation settings have uncertainties. The true clustering results may not fully follow such settings, and the clustering that satisfies these settings may not necessarily be the correct result. Therefore, a random variable is introduced to quantitatively represent whether a candidate keyword field can become a true keyword. The random variable can form a joint probability distribution with the actual observations, thereby transforming the keyword field recognition into a probability inference problem. By calculating the marginal posterior probability of the given keyword field random variable and comparing them, the field most likely to be the keyword can be obtained.
[0093] Table 1 Concepts related to constraint relationships
[0094]
[0095] Table 1 shows several concepts involved in the constraint relationships proposed by the present invention, where the probability p of K(f) holding k is the calculation target. By calculating and comparing the p of all candidate fields k , the final keyword can be determined. M(f, c), L(f, c), R(f, c), and D(f) refer to the message similarity constraint, semantic consistency constraint, remote coupling constraint, and dimension constraint respectively, and their observed constraint probabilities pm 、p l 、p r and p d Calculated from experimental data, the inferred constraint probability between K(f) and the constraint relationship needs to be assigned based on prior knowledge. Existing probability reasoning literature generally uses prior knowledge to preset prior probabilities when encountering prior probabilities. Since the probability inference algorithm often undergoes multiple rounds of iterations, the inference results are not sensitive to these prior probability values. The role of the prior probability value is mainly to characterize the reliability of different constraint relationships. When assigning the observed constraint probability, the present invention takes 0.9 and 0.1 to represent possible and impossible, respectively. The first type of inferred constraint probability (p → ) should be greater than the second type of inference constraint probability (p ← ) is large, because when a certain field is determined to be a message keyword, the characteristics of this position will most likely meet the constraints designed for the message keywords, but conversely, some fields that meet these constraints may not necessarily be keywords. Since the message similarity is only calculated by the proportion of the same byte values between messages, without considering information such as position, its reliability is low. Therefore, the present invention sets the first-class inference constraint probability of M(f,c) to 0.8, and the remaining first-class is set to 0.9. NETPLIER sets the second-class inference constraint probability between [0.6,0.8] according to the number of samples divided into each cluster. The specific calculation method is shown in formula (3), where csize represents the number of samples contained in the current cluster, and sum represents the number of samples contained in all clusters. This assignment method is beneficial to fields that cluster messages into relatively few clusters with a relatively balanced number of samples in each cluster. It works better when the number of message samples is low, but as the number of samples increases, the number of message categories contained in the data increases. At this time, continuing to use this method will make the real keyword field difficult to be discovered. Therefore, when calculating data with more than 100 samples, the present invention replaces sum with maxsize, that is, the maximum number of samples contained in all clusters.
[0096]
[0097] The message similarity constraint is defined as messages within the same cluster should be as similar as possible, and messages between different clusters should be as separated as possible. The silhouette coefficient is used here to quantify the value. The silhouette coefficient ranges from [-1,1] and can be used as an evaluation indicator to measure the clustering effect. The closer the value is to 1, the better the clustering effect is, and the closer it is to -1, the worse the clustering effect is. Let IntraDis represent the average distance between a sample and other samples in its cluster, and InterDis represent the average distance between a sample and samples in other clusters. The silhouette coefficient sc of a sample and the silhouette coefficient SC of a cluster can be defined as (n is the number of samples in the cluster):
[0098]
[0099]
[0100] For p of this cluster m , its assignment method is as follows:
[0101]
[0102] Semantic consistency constraint is defined as that the messages within the same cluster have similar semantic information. The KCSP algorithm can obtain a set of candidate keyword field sets KeySeqSet. Let UnionNum denote the number of different frequent consecutive subsequences among all kws in KeySeqSet within a certain cluster, and InterNum denote the number of frequent consecutive subsequences commonly owned by all messages in this cluster. Then for p of this cluster l , its assignment method is as follows:
[0103]
[0104] Remote coupling constraint is defined as that the request or response messages corresponding to the messages within the same cluster should all belong to the same cluster. Through preprocessing, the traffic can be segmented into two-way sessions based on the five-tuple, and the candidate keywords are respectively used to cluster the client and server messages. The messages in the session can be marked with the clusters they belong to. For a cluster with MsgNum messages, calculate the maximum number of request or response messages belonging to the same cluster in the messages of this cluster, denoted by MaxNum. Then for p of this cluster r , its assignment method is as follows:
[0105]
[0106] Dimension constraint is defined as that the number of clusters divided according to the keyword field values is not very large and each cluster contains a sufficient number of messages. Let CluNum denote the number of clusters divided according to the keyword field values, MsgsNum denote the number of all messages, and SingleNum denote the number of clusters containing only one message. Then for p of the current keyword field d , its assignment method is as follows:
[0107]
[0108] The process of using the factor graph model for probability calculation includes:
[0109] Step S131, use probability functions to describe each constraint relationship. For simplicity of expression, let k represent K(f), and x i represent a specific constraint relationship of a specific cluster or the entire clustering result; assume x iIf it represents the message similarity constraint M(f, c) of a certain cluster c, then the message similarity observation constraint and the two inference constraints of this cluster can be represented by Formula (10), Formula (11) and Formula (12) respectively.
[0110]
[0111]
[0112]
[0113] Step S132, let f j represent a certain probability function. The joint probability function of variable k and the observed value is expressed as the product of the current values of s probability functions divided by the sum of the products of all possible values of s probability functions. n and s represent the number of constraint relationships and probability functions respectively. The joint probability function is as shown in Formula (13):
[0114]
[0115] In step S133, the probability that the candidate keyword field f becomes the message keyword is represented by the marginal probability of k, as shown in Formula (14):
[0116]
[0117] The present invention uses a factor graph to represent all probability functions and perform probability calculations. As a probabilistic graphical model, a factor graph is a bipartite graph with factor nodes and variable nodes. Factor nodes represent probability functions, variable nodes represent variables used in probability functions, and the connections indicate the correlation between variable nodes and factor nodes. The keyword discrimination factor graph model used in the present invention is as Figure 2 shown. The factor graph can efficiently solve the marginal distribution of each variable through the belief propagation algorithm and is widely used in fields such as signal processing, systems biology, and systems dynamics. The specific implementation can refer to the research of Ankan et al. After obtaining the probabilities of all candidate keyword fields, the field with the highest probability is selected as the message keyword.
[0118] A specific experiment is given below to verify the effect of the present invention.
[0119] (1) Experimental data
[0120] The two datasets used in the experiments of the present invention are the Basic Application Protocol Dataset and the NETPLIER Public Dataset respectively. NETPLIER is one of the most cutting-edge works in the field of packet type recognition at present, and its source code and part of the experimental data have been open-sourced. Since the scale of each protocol in the open-sourced experimental data is about 100 packets, in order to verify the generality of the method, it is necessary to construct other datasets with larger scales. For the convenience of comparison, the present invention selects the Basic Application Protocol Dataset whose protocol types are all included in the NETPLIER dataset.
[0121] The traffic data are all processed in units of sessions. First, the present invention randomly selects data of different scales from the Basic Application Protocol Dataset in units of packets. Considering the experimental environment, the data scale of each protocol is set to three types: 100, 500, and 1000 packets.
[0122] The Basic Application Protocol Dataset filters the traffic of three common application protocols from the public traffic dataset, including DHCP, SMB, and NTP. DHCP (Dynamic Host Configuration Protocol) is a local area network protocol that communicates based on the UDP protocol and mainly functions to centrally manage and allocate IP addresses. SMB (Server Message Block) is a network file system protocol that uses the application programming interface of NetBIOS, and Microsoft's file sharing function uses this protocol for data transmission. NTP (Network Time Protocol) is a time synchronization protocol defined by RFC1305, which is used to synchronize time between distributed time servers and clients. This protocol transmits packets based on UDP and uses port number 123. Table 2 shows the number of packet types and the number of packets of various protocols in this dataset.
[0123] Table 2 Basic Application Protocol Dataset
[0124]
[0125]
[0126] The NETPLIER dataset consists of 10 common protocol packets filtered from public traffic datasets. Except for the TFTP protocol, at least 1000 packets of other protocols are filtered. The publicly available NETPLIER dataset only has 9 binary protocols, and the scale of the protocols is small. The specific information is shown in Table 3. Among the 9 protocols included, DHCP has a complex field structure, so the similarity of its packets is low; ICMP and NTP have simple structures, but the packet types include broadcast packets, which may cause the failure of remote coupling constraints; SMB and SMB2 are the previous and later versions of the same protocol, but their field structures are different and both have many packet types; TFTP is a file transfer protocol, and the lengths of different packets may vary greatly; ZeroAccess is used for the communication of P2P botnets and is a typical C&C (command and control) protocol; DNP3 and Modbus are two commonly used industrial control protocols. It can be seen that this dataset has a wide coverage and strong category diversity. The disadvantage is that the publicly available data scale is too small and most of the data contains too few protocol packet types.
[0127] Table 3 Public NETPLIER Dataset
[0128]
[0129] (2) Evaluation Metrics
[0130] The present invention uses V-measure as the metric for packet type recognition. V-measure is the harmonic mean of homogeneity and completeness, and its value range is [0, 1]. The closer it is to 1, the better the recognition effect. Homogeneity h and completeness c respectively refer to that each cluster only contains members of a single class and all samples of a given class are assigned to the same cluster, and their definitions are shown in Formulas (15) and (16) respectively.
[0131]
[0132]
[0133] Let n represent the total number of samples, n t and n c respectively represent the number of samples belonging to type T and cluster C, while n t,c represents the number of samples divided from type T to cluster C. Then the definitions of H(T|C) and H(T) are shown in Formulas (17) and (18) respectively, and H(C|T) and H(C) can be obtained similarly.
[0134]
[0135]
[0136] The specific calculation method of V-measure is as follows:
[0137]
[0138] (3) Experimental results
[0139] Since the message keywords of the protocol used in the present invention are all fixed-length fields and their positions do not change, the position standard deviation threshold t is set to 0 in the experiment. By analyzing a large number of protocol formats, it is found that the field lengths of most protocols are below four bytes and the basic unit is a byte. Therefore, the basic unit of frequent item mining in the present invention is a byte, and the highest n-gram is set to 4 during the recursive mining of frequent items. The two minimum supports Min_Supp 1 and Min_Supp 2 are both set to 10% of the total number of samples.
[0140] The first experiment compares the method of the present invention with NETPLIER, NEMETYL, and Netzob on the basic application protocol dataset, where the experimental parameters of NETPLIER are all taken as default values and the similarity threshold of Netzob is taken as 50%. Table 4 shows the experimental results obtained by the four methods on the basic application protocols with the scales of 100, 500, and 1000 messages.
[0141] As can be seen from Table 4, the method of the present invention performs well on the NTP and SMB protocols. For these two protocols, the V-measure of the method of the present invention has the highest value regardless of the data scale, and the keywords of the SMB protocol are accurately identified. On the DHCP protocol, the method of the present invention performs better than the other three methods on the data scale of 100 messages. However, its recognition effect decreases as the data scale increases, and the V-measure obtained on the other two data scales lags behind NETPLIER, with a gap of about 8% in both cases. The reason for this phenomenon is that the message length of the DHCP protocol is large, and as the data scale increases, the number of frequent subsequences increases, interfering with the recognition of keywords. In fact, only by relaxing the judgment range of the keyword field, the method of the present invention can also obtain accurate results on the DHCP protocol. In the comparison with the other two methods, the method of the present invention has an overall advantage, and the V-measure obtained is higher than that of NEMETYL and Netzob in all cases.
[0142] Table 4 Comparison of message type recognition effects I
[0143]
[0144] Figure 3It shows the time-consuming changes of four methods under different protocol scales. In the same experimental environment, the running time of the method of the present invention is the most stable, basically remaining at the same order of magnitude under three data scales, all below 100 seconds. NETPLIER takes the longest running time in most cases, and its time-consuming increases sharply with the increase of data scale, especially when facing protocols with large message lengths such as DHCP and SMB. The operation time-consuming of NEMETYL and Netzob is between the method of the present invention and NETPLIER. Since Netzob also aligns the messages, when facing long-message protocols such as DHCP, the trend of its time-consuming changing with the increase of data scale is close to that of NETPLIER.
[0145] In order to further compare with NETPLIER, the present invention uses the publicly available dataset of NETPLIER to conduct experiments on the method of the present invention and NETPLIER. The experimental results are shown in Table 5. Both methods accurately identified the keywords in protocols such as SMB, TFTP, ZeroAccess, and Modbus. In addition, the method of the present invention separately inferred the keywords of the DNP3 protocol, while NETPLIER separately inferred the keywords of the DHCP protocol and the ICMP protocol. In the remaining two protocols, the performance of the method of the present invention is slightly better than that of NETPLIER. Generally speaking, except for the DHCP protocol and the ICMP protocol, the performance of the method of the present invention is better than that of NETPLIER.
[0146] Table 5 Comparison of Message Type Recognition Effects II
[0147]
[0148] Figure 4 It gives the running time-consuming of the method of the present invention and NETPLIER on different protocols. In the same experimental environment, the running time of the method of the present invention is within 50 seconds, and the time-consuming of NETPLIER on all protocols exceeds that of the method of the present invention. Except for the DNP3 protocol and the Modbus protocol, the time-consuming gap of the remaining protocols is very large.
[0149] Corresponding to the above message type recognition method based on data mining, this embodiment also proposes a message type recognition device based on data mining, including:
[0150] A frequent continuous subsequence generation module, which is used to generate frequent continuous subsequences for the message sequence using the continuous sequence pattern algorithm;
[0151] A candidate keyword field selection module, which is used to generate position-related candidate keyword fields on the selected frequent continuous subsequences through the key continuous sequence pattern algorithm;
[0152] A probability calculation module, configured to calculate the probability of a candidate keyword field becoming a keyword based on a factor graph model;
[0153] A message type recognition module, configured to select the candidate keyword field with the highest probability as the keyword to determine the message type.
[0154] Those of ordinary skill in the art can understand that all or part of the steps to implement the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps including the above method embodiments; and the foregoing storage medium includes: various media such as ROM, RAM, magnetic disk, or optical disc that can store program codes.
[0155] Finally, it should be noted that the above are only the preferred embodiments of the present invention, which are only used to illustrate the technical solutions of the present invention and are not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention are included in the protection scope of the present invention.
Claims
1. A method for identifying message types based on data mining, characterized in that, it includes the following steps: Step 1: Use the continuous sequence pattern algorithm to generate frequent continuous subsequences for the message sequence; the continuous sequence pattern algorithm specifically includes the following steps: First, extract subsequences with a basic length of 1 from all sequences and store them in the subsequence set; Then, calculate the frequent continuous subsequences that reach the minimum support in the subsequences with a length of l - 1 basic lengths, where the support is defined as the number of sequences containing the target subsequence, and use the frequent continuous subsequences with a length of l - 1 to generate continuous subsequences with a length of l according to the Apriori strategy, and iterate this step until no new continuous subsequences can be extracted; Finally, take the union of all frequent continuous subsequences, delete the subsequences included in other sequences in the set, sort them in descending order of support and return; Step 2: Generate position-related candidate keyword fields on the selected frequent continuous subsequences through the key continuous sequence pattern algorithm, and multiple candidate keyword fields form a candidate keyword field set; The generating of position-related candidate keyword fields on the selected frequent continuous subsequences through the key continuous sequence pattern algorithm specifically includes: First, select subsequences in the frequent continuous subsequence set that meet the following three conditions: (1) The standard deviation of the subsequence position is less than the preset value, indicating that the position change range of the subsequence is not large; (2) The support of the subsequence is not 1, indicating that the subsequence does not appear in all messages; (3) The subsequence does not exist in an existing candidate keyword field to prevent repeated calculation; Then, create a message sequence set that does not include the subsequences that meet the above conditions and truncate these message sequences according to the minimum start position and maximum end position of the current frequent continuous subsequence; Finally, run the continuous sequence pattern algorithm on the newly created message sequences. If new frequent continuous subsequences are obtained, combine the results obtained by the continuous sequence pattern algorithm with the current frequent continuous subsequences to form a candidate keyword field set and save the information according to the candidate keyword field structure; Step 3: Calculate the probability that the candidate keyword fields in the candidate keyword field set become keywords based on the factor graph model; Step 4: Select the candidate keyword field with the highest probability as the keyword to determine the message type.
2. The method for identifying message types based on data mining according to claim 1, characterized in that, The expression for multiple candidate keyword fields to form a candidate keyword field set is as follows: KeySeqSet = {K 1 , K 2 ,..., K n} where n represents the number of candidate keyword fields; Then, the candidate keyword field K i has the following structural expression: K i = [MinOffset, MaxDepth, ModeOffset, ModeLen, kw] Among them, MinOffset represents the minimum starting position of the current frequent consecutive subsequence fs, MaxDepth represents the maximum ending position of the current frequent consecutive subsequence fs, ModeOffset represents the minimum mode of the starting positions of all frequent consecutive subsequences within kw, ModeLen represents the minimum mode of the lengths of all frequent consecutive subsequences within kw, and kw is the set of all frequent consecutive subsequences within the range from MinOffset to MaxDepth in the current frequent consecutive subsequence fs and the message sequence set after removing the messages containing fs. Therefore, by setting this structure, it is ensured that the key consecutive sequence pattern algorithm can identify keyword fields with unfixed positions.
3. The message type recognition method based on data mining according to claim 2, characterized in that, when performing message clustering subsequently, first search in the range from MinOffset to MaxDepth in descending order of support to find whether there is a frequent consecutive subsequence in kw. If there is, identify the category of the message as the frequent consecutive subsequence. Otherwise, intercept the sequence in the range from MinOffset to MaxDepth according to ModeOffset and ModeLen for category identification.
4. The message type recognition method based on data mining according to claim 1, characterized in that, the factor graph model models the uncertainty in keyword recognition as the joint distribution of observed values and a set of random variables. Each random variable represents the probability that a candidate keyword field is a message keyword, and the observed value is the relevant probability of the four constraint relationships; the four constraint relationships include: (1) message similarity constraint M(f, c), defined as messages within the same cluster are as similar as possible, and messages between different clusters are as separated as possible; (2) semantic consistency constraint L(f, c), defined as messages within the same cluster have similar semantic information; (3) remote coupling constraint R(f, c), defined as the request or response messages corresponding to messages within the same cluster belong to the same cluster; (4) dimension constraint D(f), defined as the number of clusters divided according to the values of keyword fields is not very large and each cluster contains a sufficient number of messages; each constraint relationship includes an observation constraint and an inference constraint. The observation constraint refers to whether the constraint relationship itself holds, and the inference constraint refers to the inference relationship between the constraint relationship and the random variables.
5. The message type recognition method based on data mining according to claim 4, characterized in that, the random variable is represented as K(f), and the field f is a protocol keyword field; the expression for the probability of the message similarity constraint M(f, c) holding is as follows: Let IntraDis represent the average distance between a certain sample and other samples within its cluster, InterDis represent the average distance between a certain sample and samples in other clusters, and n be the number of samples in the cluster. Then the silhouette coefficient sc for a certain sample and the silhouette coefficient SC for a certain cluster are respectively defined as: the expression for the probability of the semantic consistency constraint L(f, c) holding is as follows: Among them, UnionNum represents the number of different frequent consecutive subsequences among all kws in the candidate keyword field set within a certain cluster, and InterNum represents the number of frequent consecutive subsequences jointly owned by all messages in this cluster; The expression for the probability of the remote coupling constraint R(f, c) to hold is as follows: For a cluster with MsgNum messages, calculate the maximum number of request or response messages belonging to the same cluster in the messages of this cluster, denoted by MaxNum; The expression for the probability of the dimension constraint D(f) to hold is as follows: Among them, CluNum represents the number of clusters divided according to the values of keyword fields, MsgsNum represents the number of all messages, and SingleNum represents the number of clusters containing only one message; The inference probabilities of the four constraint relationships and random variables are set according to prior knowledge.
6. The method for identifying message types based on data mining according to claim 5, characterized in that The process of using the factor graph model for probability calculation includes: First, use probability functions to describe each constraint relationship. Let k represent K(f), and x i represent a specific cluster or a certain constraint relationship of the entire clustering result; Then, let f j represent a certain probability function. The joint probability function of variables k and the observed values is expressed as the product of the current values of s probability functions divided by the sum of the products of all possible values of the s probability functions. n and s represent the number of constraint relationships and probability functions respectively. The formula for the joint probability function is as follows: Finally, the probability that the candidate keyword field f becomes the keyword of the message is expressed as the marginal probability of k, as shown in the following formula:
7. A device for identifying message types based on data mining, characterized in that It is used to implement the method for identifying message types based on data mining according to any one of claims 1-6. The device includes: A frequent consecutive subsequence generation module, which is used to generate frequent consecutive subsequences for the message sequence using the continuous sequence pattern algorithm; A candidate keyword field selection module, which is used to generate position-related candidate keyword fields on the selected frequent consecutive subsequences through the key continuous sequence pattern algorithm; A probability calculation module, which is used to calculate the probability that the candidate keyword field becomes the keyword based on the factor graph model; A message type identification module, which is used to select the candidate keyword field with the highest probability as the keyword to determine the message type.
Citation Information
Patent Citations
Protocol format inference method based on closed sequential pattern mining
CN108667839A
Data packet desensitization method and device
CN111935081A