A Method for Dynamically Identifying Attack Organizations Based on Cyberspace Detection Behaviors

By establishing a multi-classification model and an attack organization identification mechanism based on cyberspace detection behavior, the problem of missing threat intelligence information is solved, and accurate identification and dynamic tracking of attack organizations is achieved.

CN116405275BActive Publication Date: 2025-06-13SHENYANG INST OF AUTOMATION - CHINESE ACAD OF SCI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202310322636.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-29
Publication Date
2025-06-13
Estimated Expiration
2043-03-29

AI Technical Summary

Technical Problem

The prior art is difficult to deal with the problem of lack of threat intelligence information when identifying attack organizations, especially when attacking organizations use multiple attack tools, which are difficult to accurately identify.

Method used

By scanning industrial control devices exposed to the Internet using open source cyberspace detection tools, extracting the feature vectors of network attack traffic, establishing a multi-classification model based on the One-Vs-One algorithm and the unparalleled supervised binary classification algorithm, combining the attack pattern categories output by the C-type network segment and the multi-classification model, an attack organization identification mechanism is established.

Benefits of technology

Improves the accurate identification capabilities of attacking organizations and can automatically identify new cyberspace detection tools without human participation to adapt to the diversity and scale of complex attacking organizations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116405275B_ABST
    Figure CN116405275B_ABST
Patent Text Reader

Abstract

The present invention proposes a method for dynamically identifying attack organizations based on cyber space detection behavior. An open-source cyber space detection tool is used to scan industrial control devices exposed on the Internet to form a reference network attack traffic set; the network attack traffic is segmented with traffic sessions as the basic unit, and an attack pattern containing feature vectors of all industrial control traffic is extracted; a multi-classification model for the attack pattern is established, and a new category determination mechanism and a self-update determination mechanism are established to enable the multi-classification model to have self-expansion and self-update functions; an attack organization identification mechanism is established to enable it to determine the organization to which the attacker belongs according to the attacker's IP and the classification and determination results of the multi-classification model. This method can accurately and dynamically identify existing and unknown attack organizations in the absence of threat intelligence information, enabling network defenders to master attack organization information in advance before an attack event occurs.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of industrial control system network security, and specifically to a method for dynamically identifying attack organizations based on cyber space detection behaviors. Background Art

[0002] With the introduction of concepts such as industrial Internet, the degree of integration between industrial control systems and information systems has been continuously deepened, promoting the gradual transformation of traditional industries towards networking, digitization, and intelligence. As a result, the security issues of industrial control systems have received unprecedented attention. How to achieve in-depth excavation of attack organizations, especially identification at the early stage of attack implementation, has become a difficult problem that attracts the attention of the academic and industrial communities.

[0003] Currently, the identification of attack organizations is still in a stage of being aware of things after they have happened. That is, usually after an industrial enterprise has suffered an attack event, network security companies, based on network attack models (such as STRIDE, CyberKillChain, ATT&CK, CAPEC, etc.), summarize and review the logs and alarm information reported by various security products to dig out latent attack organizations. However, according to the seven stages of the kill chain model, namely reconnaissance, weaponization, delivery, exploitation, installation, command and control, and actions on objectives, attack organizations usually start to perform reconnaissance tasks on the attack targets on the network in the first stage of the attack, which is also the best stage to discover the action traces of attack organizations in the external cyber space. Distributed industrial control honeynets have unique advantages in this regard. By deploying distributed industrial control honeynets on the public network, a large number of detection behaviors targeting exposed industrial control devices can be captured. Such detection behaviors can better reflect the attack purposes of attackers and the attack patterns of attack tools, thereby providing a basis for attack organization identification. With the gradual deployment of distributed industrial control honeynets and the continuous accumulation of attack traffic, how to accurately identify attack organizations based on cyber space detection behaviors has become an urgent problem to be solved.

[0004] At the present stage, the methods for identifying attack organizations based on cyber space detection behaviors mainly include two categories. One category is to use machine learning methods to establish a classification model based on the relevant attack traffic of detection behaviors to achieve classification of attack patterns, and divide attack organizations according to attack patterns. The main problem of this type of method is that it is difficult to identify attack organizations that contain multiple attack tools (usually with different attack patterns). The other category is the method for identifying attack organizations based on threat intelligence. This type of method uses network crawlers to obtain relevant information based on the attacker IPs captured by the honeynet, and then determines which IPs belong to the same attack organization. Although this method has relatively high accuracy, the main problem it currently faces is the lack of threat intelligence, especially the threat intelligence about attackers is even more difficult to obtain. Summary of the Invention

[0005] In view of the deficiencies of the prior art, the present invention proposes a dynamic recognition method for attack organizations based on cyber space detection behavior, aiming to solve the problem of accurately identifying attack organizations from external networks in the absence of threat intelligence information, and improving the accuracy compared with the current recognition method based on attack patterns.

[0006] The technical solution adopted by the present invention to achieve the above object is:

[0007] A dynamic recognition method for attack organizations based on cyber space detection behavior, comprising the following steps:

[0008] Use open-source cyber space detection tools to scan industrial control devices exposed on the Internet to obtain a network attack traffic set, and use the open-source cyber space detection tools as class labels;

[0009] Segment the network attack traffic set with traffic sessions as the basic unit, extract the attack pattern AP from the traffic sessions, and form an attack pattern set;

[0010] Use the One-Vs-One algorithm and a parameter-free supervised binary classification algorithm to establish a multi-classification model based on attack patterns, and train it with the attack pattern set;

[0011] Establish a new category determination mechanism for the multi-classification model;

[0012] Establish a self-update determination mechanism for the multi-classification model;

[0013] Use the Class C network segment and the attack pattern categories output by the multi-classification model as the basis for identifying attack organizations, and establish an attack organization identification mechanism to identify attack organizations in the cyber space.

[0014] The attack pattern set is a set composed of attack patterns AP. The attack pattern AP is composed of the feature vectors of all industrial control traffic in the traffic session, and is expressed as:

[0015] AP = {A 1 , A 2 , …, A m}

[0016] Among them, A i represents the feature vector of industrial control traffic, i = 1, 2, …, m, and m represents the number of feature vectors. The industrial control traffic is a data packet containing industrial control protocol request data, and A i is expressed as:

[0017] A i = {a 1 , a 2 , …, a q}

[0018] Among them, ai Represents specific feature attributes, and q represents the number of feature attributes.

[0019] The multi-classification model MC adopts a voting mechanism, and the class with the most votes is determined as the final output class, expressed as:

[0020] MC = {M 1 , M 2 , …, M p}

[0021] Among them, M i represents a parameter-free supervised binary classification model, i = 1, 2, …, p, and p = n(n - 1) / 2 represents the total number of binary classification models, where n is the total number of classes.

[0022] The new class determination mechanism includes two determination steps: threshold determination and model determination. When the new attack pattern passes both determinations, it is considered that the new attack pattern represents a new class, and a corresponding new binary classification model is created using this attack pattern to expand the multi-classification model.

[0023] The threshold determination includes two sub-determination conditions: attack pattern fingerprint determination and average feature vector determination, and the two are in an "or" relationship, that is, if either of them is established, the threshold determination passes.

[0024] The attack pattern fingerprint determination is expressed as:

[0025]

[0026] Among them, APF represents the fingerprint of the test attack pattern, and the attack pattern is initially classified into class C i by the multi-classification model, APF ij represents the j-th attack pattern fingerprint in class C i , dist() represents the Euclidean distance between vectors, d min () represents the minimum distance between APF and all attack pattern fingerprints in class C i , D fin represents the fingerprint threshold, and APF is expressed as:

[0027] APF = {V 1 , V 2 , …, V n}

[0028] Among them, V i represents the average number of votes for the attack pattern belonging to class C i , i = 1, 2, …, n, expressed as:

[0029]

[0030] Among them, v ji represents the feature vector A j belonging to the category C i of votes, expressed as:

[0031]

[0032] Among them, m k represents the classification result of a binary classification model, and p k represents the probability corresponding to the classification result m k I i is an indicator function. When m k is i, the value of I i is 1, otherwise the value is 0.

[0033] The average feature vector determination is expressed as:

[0034]

[0035] Among them, D vec represents the vector threshold, AFV represents the average feature vector of the test attack pattern, and AFV ij represents the average feature vector of the j-th attack pattern in the category C i , and AFV is expressed as:

[0036] AFV = {AF 1 , AF 2 , …, AF q}

[0037] Among them, AF i represents the i-th average attribute of AFV, where i = 1, 2, …, q, and is expressed as:

[0038]

[0039] Among them, a ji represents the i-th attribute of the j-th feature vector of the attack pattern.

[0040] The model determination is specifically as follows:

[0041] Take the attack pattern as a new category to generate a new binary classification model and add it to the existing multi-classification model; if the extended multi-classification model can correctly classify the attack pattern, the model determination passes.

[0042] The self-update determination is as follows:

[0043]

[0044] Among them, D up represents the update threshold, and |C i | represents the category Ci The number of average feature vectors owned, U represents an update coefficient for controlling the update ratio of the attack pattern, and μ is an update coefficient function.

[0045] The attack organization recognition mechanism is specifically as follows:

[0046] 1) Establish an attack organization set AO:

[0047] AO = {O 1 , O 2 , …, O l}

[0048] Among them, O j represents an attack organization, j = 1, 2 …, l, where l is the number of attack organizations, and O j is expressed as:

[0049] O j = <C j , S j , D j >

[0050] Among them, C j represents the label category to which O j belongs, S j represents the set of C-class network segments included in O j , and is expressed as:

[0051] S j = {s 1 , s 2 , …, s x}

[0052] Among them, x represents the number of C-class network segments, and D j represents the set of complete IP addresses included in O j , and is expressed as:

[0053] D j = {d 1 , d 2 , …, d y}

[0054] Among them, y represents the number of complete IP addresses;

[0055] 2) For an input attack pattern AP i , extract its C-class network segment s i , complete IP address d i and the classification result c i of the multi-classification model;

[0056] 3) Determine whether s i belongs to an attack organization O j, if it belongs to an attack organization O j , then determine c i whether it is consistent with O j in category, and execute step 4); otherwise, execute step 8);

[0057] 4) If c i is consistent with O j in category, then determine whether the category C j of O j is the category to which the open-source cyberspace detection tool belongs, and execute step 5); otherwise, execute step 7);

[0058] 5) If C j is the category to which the open-source cyberspace detection tool belongs, then determine whether AP i passes the new category determination, and execute step 6); otherwise, determine that d i belongs to O j ;

[0059] 6) If AP i passes the new category determination, then remove the attack organization O j from the category of the open-source cyberspace detection tool to which it belongs, merge it with AP i into a new category, and create a new binary classification model and attack organization for it; otherwise, determine that d i belongs to O j ;

[0060] 7) Determine whether c i is the category to which the open-source cyberspace detection tool belongs. If c i is the category to which the open-source cyberspace detection tool belongs, then determine that d i belongs to O j ; otherwise, merge the attack organization O i corresponding to c i with O j , and at the same time merge the binary classification models corresponding to O i and O j ;

[0061] 8) Determine whether AP i passes the new category determination. If it passes the new category determination, then create a new binary classification model and attack organization for AP i ; otherwise, determine whether c i is the category to which the open-source cyberspace detection tool belongs;

[0062] 9) If c i is the category to which the open-source cyberspace detection tool belongs, then create a new attack organization for it based on the C-class network segment s i ; otherwise, determine that d i belongs to ci The corresponding attacking organization O i .

[0063] The present invention has the following beneficial effects and advantages:

[0064] 1. Based on the in-depth analysis of network space detection behaviors, an attack pattern is established with traffic sessions as the basic unit and feature vectors containing all industrial control traffic, which can more deeply and comprehensively depict the behavioral characteristics of network space detection behaviors and is more conducive to distinguishing the network space detection tools used by different attacking organizations;

[0065] 2. The established multi-classification model of attack patterns with self-expansion and self-update functions can automatically identify new network space detection tools other than open-source network space detection tools without manual participation, and can continuously enrich the existing binary classification model with subsequent attack patterns and prevent the rapid expansion of the multi-classification model;

[0066] 3. The proposed attacking organization recognition mechanism makes use of the characteristic that attackers belonging to the same Class C network segment usually come from the same attacking organization, effectively making up for the deficiencies of the previous method that only uses attack patterns or threat intelligence as the basis for evaluating attacking organizations, and can discover complex attacking organizations with diverse attack patterns and a large number of members. BRIEF DESCRIPTION OF THE DRAWINGS

[0067] Figure 1 is a schematic flowchart of the method of the present invention;

[0068] Figure 2 is the traffic session generated by the open-source detection tool Plcscan in the embodiment of the present invention;

[0069] Figure 3 is a schematic diagram of the multi-classification model in the embodiment of the present invention;

[0070] Figure 4 is a schematic flowchart of the new category determination process in the embodiment of the present invention;

[0071] Figure 5 is an expanded schematic diagram of the multi-classification model in the embodiment of the present invention;

[0072] Figure 6 is a schematic flowchart of the attacking organization recognition process in the embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0073] The following further elaborates the present invention in detail in conjunction with the drawings and specific implementation schemes, but does not limit the technical solution of the present invention.

[0074] The overall process of an attack organization dynamic recognition method based on network space detection behaviors provided by the present invention is as Figure 1 shown, including the following steps:

[0075] 1) Use open-source cyberspace detection tools to scan industrial control devices exposed on the Internet, form a reference network attack traffic set, and use the corresponding open-source cyberspace detection tools as category labels;

[0076] 2) Segment the network attack traffic with traffic sessions as the basic unit, extract attack patterns from the traffic sessions to form an attack pattern set, where the extracted attack patterns are composed of the feature vectors of all industrial control traffic in the traffic sessions;

[0077] 3) Use the One-Vs-One (OVO) algorithm and parameter-free supervised binary classification algorithms (such as CART algorithm, random forest algorithm, SVM algorithm, etc.) to establish an attack pattern multi-classification model with self-expansion and self-update functions, and use the attack pattern set generated in step 2) as the training data set for training;

[0078] 4) Establish a new category determination mechanism for the multi-classification model generated in step 3). Its main purpose is to discover new types of attack tools outside the multi-classification model, specifically including two determination steps: threshold determination and model determination. When the new attack pattern passes both determinations, it is considered that the new attack pattern represents a new category and is used to create a corresponding new binary classification model to expand the multi-classification model;

[0079] 5) Establish a self-update determination mechanism for the multi-classification model generated in step 3). Its main purpose is to continuously enrich the existing binary classification models and prevent the rapid expansion of the multi-classification model. When the new attack pattern fails to pass the new category determination and is determined to be an existing category, a self-update determination is performed on it. If it passes the determination, it is used as new training data to update the binary classification model corresponding to the category to which it belongs;

[0080] 6) Establish an attack organization identification mechanism based on the C-class network segment (i.e., in the 4-segment number of the IP address, the network number composed of the first 3 segment numbers) and the attack pattern category as the identification basis for the attack organization;

[0081] The traffic session in step 2) is defined as: for a pair of given hosts, all traffic connections with a time interval within T between them will be aggregated into the same traffic session. When the transport layer protocol is the TCP protocol, the traffic connection refers to the TCP connection. When the transport layer protocol is the UDP protocol, the traffic connection refers to a single network data packet. aggreg The attack pattern (AP) in step 2) is expressed as:

[0082] AP = {A

[0083] , A 1 , A 2 , …, A m}

[0084] Among them, A i represents the feature vector of industrial control traffic, m represents the number of feature vectors, where industrial control traffic refers to data packets containing industrial protocol request data, and A i is expressed as:

[0085] A i ={a 1 , a 2 , …, a q}

[0086] Among them, a i represents the specific feature attribute, and q represents the number of feature attributes.

[0087] The multi-classification model (MC) in step 3) is expressed as:

[0088] MC={M 1 , M 2 , …, M p}

[0089] Among them, M i represents a parameter-free supervised binary classification model, p = n(n - 1) / 2 represents the total number of binary classification models, and n is the total number of categories. The multi-classification model adopts a voting mechanism, and the category with the most votes is determined as the final output category.

[0090] The threshold determination in step 4) includes two sub-determination conditions: attack mode fingerprint determination and average feature vector determination, and the two are in an "or" relationship, that is, if either of them is established, the threshold determination passes.

[0091] The attack mode fingerprint determination is expressed as:

[0092]

[0093] Among them, D fin represents the fingerprint threshold, APF represents the fingerprint of the attack mode, and the attack mode is initially classified into C i categories by the multi-classification model, dist() represents the Euclidean distance between vectors, and APF is expressed as:

[0094] APF={V 1 , V 2 , …, V n}

[0095] Among them, V i represents the average number of votes for the attack mode belonging to category C i , and is expressed as:

[0096]

[0097] Among them, v ji represents the feature vector A i belonging to the category C i of votes, expressed as:

[0098]

[0099] Among them, among them, m k represents the classification result of a binary classification model, p k represents the probability corresponding to the classification result m k and I i is the indicator function.

[0100] The determination of the average feature vector is expressed as:

[0101]

[0102] Among them, D vec represents the vector threshold, AFV represents the average feature vector of the attack mode, expressed as:

[0103] AFV = {AF 1 , AF 2 , …, AF q}

[0104] Among them, AF i represents the i-th average attribute of AFV, expressed as:

[0105]

[0106] Among them, a ji represents the i-th attribute of the j-th feature vector of the attack mode.

[0107] The model determination in step 4) means using the attack mode as a new category to generate a new binary classification model and adding it to the existing multi-classification model; if the extended multi-classification model can correctly classify the attack mode, then it passes the model determination;

[0108] The self-update determination in step 5) is expressed as:

[0109]

[0110] Among them, D up represents the update threshold, |C i | represents the number of average feature vectors owned by the category C i and U represents the update coefficient, which is used to control the update ratio of the attack mode.

[0111] The attack organization recognition mechanism in step 6) includes the following steps:

[0112] 6.1) Establish the attack organization set (AO):

[0113] AO = {O 1 , O 2 , …, O l}

[0114] where O j represents an attack organization, expressed as:

[0115] O j = <C j , S j , D j >

[0116] where C j represents the label category to which O j belongs, S j represents the set of C-class network segments included in O j , expressed as:

[0117] S j = {s 1 , s 2 , …, s l}

[0118] D j represents the set of complete IP addresses included in O j , expressed as:

[0119] D j = {d 1 , d 2 , …, d l}

[0120] 6.2) For an input attack pattern AP i , extract its C-class network segment s i , complete IP address d i and the classification result c i of the multi-classification model;

[0121] 6.3) Determine whether s i belongs to an attack organization in the attack organization set;

[0122] 6.4) If it belongs to an attack organization O j , then determine whether c i is consistent with the category of O j ;

[0123] 6.5) If c i is consistent with the category of O j , then determine the category C j of O jWhether it is the category of open-source cyberspace detection tools;

[0124] 6.6) If C j is the category of open-source cyberspace detection tools, then determine whether AP i passes the new category determination;

[0125] 6.7) If AP i passes the new category determination, then remove the attack organization O j from the category of the open-source cyberspace detection tool it belongs to, merge it with AP i to form a new category, and create a new binary classification model and attack organization for it; otherwise, determine that d i belongs to O j ;

[0126] 6.8) If C j is not the category of open-source cyberspace detection tools, then determine that d i belongs to O j ;

[0127] 6.9) If c i is inconsistent with the category of O j then determine whether c i is the category of open-source cyberspace detection tools;

[0128] 6.10) If c i is the category of open-source cyberspace detection tools, then determine that d i belongs to O j ; otherwise, merge the attack organization O i corresponding to c i with O j and at the same time merge the binary classification models corresponding to O i and O j ;

[0129] 6.11) If s i does not belong to the set of attack organizations, then determine whether AP i passes the new category determination;

[0130] 6.12) If it passes the new category determination, then create a new binary classification model and attack organization for AP i ; otherwise, determine whether c i is the category of open-source cyberspace detection tools;

[0131] 6.13) If c i is the category of open-source cyberspace detection tools, then create a new attack organization for it based on the C-class network segment s i ; otherwise, determine that d i belongs to c iThe corresponding attack organization O i 。

[0132] Embodiment

[0133] Taking the Modbus industrial control protocol as an example, this embodiment specifically includes the following steps:

[0134] Step 1: Use an open-source cyberspace detection tool to scan industrial control devices exposed on the Internet, form a reference network attack traffic set, and use the corresponding open-source cyberspace detection tool as a category label. The information of the obtained data set is as follows:

[0135] Table 1 Cyberspace Detection Data Set

[0136]

[0137] Step 2: Segment the network attack traffic with the traffic session as the basic unit, extract the attack patterns from the traffic session to form an attack pattern set. For the network attack traffic in Table 1, segment and extract features with T aggreg = 100 seconds, and the following attack pattern representation is obtained:

[0138]

[0139] Among them, Seq i represents the sequence number of the industrial control protocol request data packet within a traffic connection, Res i indicates whether the data packet has received a response, Pro i represents the type of transport layer protocol, including TCP and UDP, Len i represents the length of the data packet, Dep i represents the total number of industrial control protocol request data packets in a traffic connection, CNo i represents the sequence number of the current traffic connection in a traffic session, CTo i represents the total number of traffic connections in a traffic session, Tra i , Uni i and Fun i correspond to the "Transaction Identifier", "Unit Identifier", and "Function Code" fields in the Modbus protocol respectively. Pre i refers to the Uni i-1 value of the previous industrial control protocol request data packet. When i ≤ 1, its default value is -1. Taking a traffic session generated by the detection tool Plcscan as an example, as Figure 2 shown, its corresponding attack pattern is as follows:

[0140]

[0141] Step 3: Use the One-Vs-One (OVO) algorithm and the parameter-free supervised binary classification algorithm (taking the CART algorithm as an example) to establish a multi-classification model for attack patterns with self-expansion and self-update functions. Use the attack pattern set generated in step 2) as the training data set for training, as Figure 3 shown, and its specific representation form is as follows:

[0142] MC = {M 1 , M 2 , …, M 21}

[0143] Among them, M 1 represents a CART decision tree composed of class 1 and class 2, M 2 represents a CART decision tree composed of class 1 and class 3, M 21 represents a CART decision tree composed of class 6 and class 7, and so on. From the attack patterns, the input of a CART decision tree is a feature vector containing 11-dimensional elements, and the output is the determined class represented in the form of probability. For example, the output <1, 0.9> of M 1 means that the determined class is 1 and the determination probability is 0.9. Finally, the multi-classification model adopts a voting mechanism, and the class with the most votes is determined as the final output class.

[0144] Step 4: Establish a new class determination mechanism for the multi-classification model generated in step 3), and the specific determination process is as Figure 4 shown. When the total number of classes is 7, the following attack pattern fingerprint representation form is obtained:

[0145] APF = {V 1 , V 2 , V 3 , V 4 , V 5 , V 6 , V 7}

[0146] Through the analysis of the experimental data, when the fingerprint threshold D fin takes values in the range of [1.6, 2.8], better classification effects will be obtained. Taking D fin equal to 1.9 as an example, when the attack pattern fingerprint set corresponding to class 7 is as follows:

[0147]

[0148] If the fingerprint corresponding to the attack pattern to be classified is:

[0149] APF 1 = {4, 2, 2, 0, 0, 6, 7}

[0150] Since d min (APF 1 , C 7 ) = 1.41 < 1.9, so it fails the attack pattern fingerprint determination. If the fingerprint corresponding to the classified attack pattern is:

[0151] APF 2 = {4, 1, 3, 0, 0, 6, 7}

[0152] Since d min (APF 2 , C 7 ) = 2.83 > 1.9, so it passes the attack pattern fingerprint determination. Since the average feature vector determination process is similar to the attack pattern fingerprint determination process, it will not be elaborated here. In addition, when the vector threshold D vec takes values in the range of [0.1, 0.9], better classification results will be obtained.

[0153] The extension process of the multi-classification model is as Figure 5 shown. When the classified attack pattern is re-input into the new multi-classification model and is determined to be a new category, it passes the model determination. In addition, since the total number of categories of the extended multi-classification model increases, it is necessary to use the newly generated binary classification model to classify the training data of the existing multi-classification model, and use the classification results to update the attack pattern fingerprints of the training data.

[0154] Step 5: Establish a self-update determination mechanism for the multi-classification model generated in step 3). When the classified attack pattern fails the new category determination and is determined to be an existing category, perform a self-update determination on it. Through the analysis of experimental data, when the update coefficient U takes values in the range of [10, 30], better classification results will be obtained. Taking U equal to 20 and D vec equal to 0.9 as an example, when the average feature vector set corresponding to category 7 is as follows:

[0155]

[0156] If the average feature vector corresponding to the classified attack pattern is:

[0157] AFV 1 = {1, 0, 1, 65, 1, 2, 2, 0, 0, -1, 43}

[0158] Since d min (AFV 1 , C 7 ) = 0 < D vec = 0.135, so it fails the self-update determination. If the average feature vector corresponding to the classified attack pattern is:

[0159] AFV 2 ={1, 0, 1, 64, 1, 2, 2, 0, 0, -1, 43}

[0160] Since d min (AFV 2 , C 7 ) = 1 > D vec = 0.135, so through self-update determination, it will be used to update the binary classification model corresponding to category 7. Accordingly, the average feature vector set corresponding to category 7 becomes:

[0161]

[0162] Step 6: Use the C-class network segment and attack mode category as the identification basis for the attack organization, and establish an attack organization identification mechanism. The specific identification process is as Figure 6 shown.

[0163] Step 6.1: Establish an attack organization set (AO). Since there is no initial input of attack organizations, the attack organization set is empty:

[0164] For various situations that may occur during the identification process, three of the most typical cases are used for illustration, and the reasons for relevant operations are briefly explained. The remaining situations can be analogously deduced from the above three cases and will not be elaborated.

[0165] Case 1: For AP 1 Create a new binary classification model and attack organization

[0166] Step 6.2: For the first input attack mode AP 1 , its relevant information can be represented as a triple, that is, <s 1 , d 1 , c 1 >;

[0167] Step 6.3: Determine whether s 1 belongs to an attack organization in the attack organization set. Obviously not, go to step 6.11;

[0168] Step 6.11: Determine whether AP 1 passes the new category determination. Assume it passes the determination, then go to step 6.12;

[0169] Step 6.12: Create a new binary classification model and attack organization for AP 1 . Since no attack organization can be matched with AP 1 through the C-class network segment, and AP 1Represents a new type of attack pattern, so a new binary classification model (class 8) and an attack organization are created, and the attack organization set is updated:

[0170] AO = {O 1},

[0171] O 1 = <C 1 , S 1 , D 1 , C 1 = 8, S 1 = {s 1}, D 1 = {d 1}

[0172] Case 2: Based on the C-class network segment, a new attack organization is created for d 2

[0173] Step 6.2: For the second input attack pattern AP 2 , its relevant information can be represented as a triple, i.e., <s 2 , d 2 , c 2 >;

[0174] Step 6.3: Determine whether s 2 belongs to an attack organization in the attack organization set. Assume it does not belong, i.e., s 2 ≠ s 1 , and go to Step 6.11;

[0175] Step 6.11: Determine whether AP 2 passes the new category determination. Assume it fails the determination, then go to Step 6.13;

[0176] Step 6.13: Assume c 2 is the category 2 of the open-source cyberspace detection tool. Then, based on the C-class network segment s 2 , a new attack organization is created for it. Since no attack organization can match AP 2 through the C-class network segment, and AP 2 is determined to be the category of the open-source cyberspace detection tool, and attack organizations belonging to different C-class network segments can all use the open-source tool for detection, so the open-source tool category cannot be used as the attack organization identification criterion. Therefore, for the attack patterns classified as the open-source tool category, the attack organization is further divided based on the C-class network segment, that is, a new attack organization is created for d 2 , and the attack organization set is updated:

[0177] AO = {O 1 , O 2},​

[0178] O 1 = <C 1 , S 1 , D 1 】> C 1 = 8, S 1 = {s 1} D 1 = {d 1}

[0179] O 2 = <C 2 , S 2 , D 2 】> C 2 = 2, S 2 = {s 2} D 2 = {d 2}

[0180] Case 3: For AP 3 and O 2 Create a new binary classification model and an attack organization

[0181] Step 6.2: For the third input attack pattern AP 3 , its relevant information can be represented as a triple, i.e., <s 2 , d 3 , c 3 】>. Obviously, AP 3 and AP 2 belong to the same C-class network segment s 2 ;

[0182] Step 6.3: Determine whether s 2 belongs to an attack organization in the attack organization set. Obviously, it belongs to the attack organization O 2 , and proceed to Step 6.4;

[0183] Step 6.4: Determine whether c 3 is consistent with O 2 in category. Assume they are consistent, i.e., both belong to category 2, and proceed to Step 6.5;

[0184] Step 6.5: Determine whether the category C 2 of O 2 is the category of open-source cyberspace detection tools. Obviously, it is, and proceed to Step 6.6;

[0185] Step 6.6: Determine whether AP 3 passes the new category determination. Assume it passes, and proceed to Step 6.7;

[0186] Step 6.7: For AP 3 and O2 Create a new binary classification model and attack organization. Although, AP 2 is initially determined to belong to category 2, but the APs within the same attack organization 3 are determined to be a new category. The possible reasons include that the attack organization has modified existing open-source detection tools or used new detection tools, thus showing a new attack pattern. Therefore, for AP 3 and O 2 create a new binary classification model (category 9) and attack organization, and at the same time update the set of attack organizations:

[0187] AO = {O 1 , O 3},

[0188] O 1 = <C 1 , S 1 , D 1 >, C 1 = 8, S 1 = {s 1}, D 1 = {d 1}

[0189] O 3 = <C 3 , S 3 , D 3 >, C 3 = 9, S 3 = {s 2}, D 3 = {d 2 , d 3}

[0190] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application and are not intended to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A method for dynamically identifying attack organizations based on cyber space detection behavior, characterized in that, it includes the following steps: Use open-source cyber space detection tools to scan industrial control devices exposed on the Internet to obtain a network attack traffic set, and use the open-source cyber space detection tools as class labels; Segment the network attack traffic set with traffic sessions as the basic unit, extract the attack pattern AP from the traffic sessions, and form an attack pattern set; Use the One-Vs-One algorithm and a parameter-free supervised binary classification algorithm to establish a multi-classification model based on attack patterns, and use the attack pattern set to train it; Establish a new category determination mechanism for the multi-classification model; Establish a self-update determination mechanism for the multi-classification model; Use the C-class network segment and the attack pattern categories output by the multi-classification model as the identification basis for attack organizations, and establish an attack organization identification mechanism to identify attack organizations in the cyber space; The new category determination mechanism includes two determination steps: threshold determination and model determination. When the new attack pattern passes both determinations at the same time, it is considered that the new attack pattern represents a new category, and a corresponding new binary classification model is created using this attack pattern to expand the multi-classification model; The self-update determination is: Among them, D up represents the update threshold, |C i | represents the number of average feature vectors owned by class C i , U represents the update coefficient for controlling the update ratio of the attack pattern, μ is the update coefficient function, D vec represents the vector threshold, and AFV represents the average feature vector of the test attack pattern.

2. The method for dynamically identifying attack organizations based on cyber space detection behavior according to claim 1, characterized in that, The attack pattern set is a set composed of the attack pattern AP, and the attack pattern AP is composed of the feature vectors of all industrial control traffic in the traffic session, expressed as: AP = {A 1 , A 2 , …, A m} Among them, A i represents the feature vector of industrial control traffic. i = 1, 2, …, m, where m represents the number of feature vectors. The industrial control traffic is a data packet containing industrial control protocol request data. A i is expressed as: A i = {a 1 , a 2 , …, a q} Among them, a i represents a specific characteristic attribute, and q represents the number of characteristic attributes.

3. The method for dynamically identifying attack organizations based on cyber space detection behavior according to claim 1, characterized in that, The multi-classification model MC adopts a voting mechanism, and the category with the most votes is determined as the final output category, expressed as: MC = {M 1 , M 2 , …, M p} Among them, M i represents a parameterless supervised binary classification model, where \(i = 1, 2,\cdots, p\), and \(p=\frac{n(n - 1)}{2}\) represents the total number of binary classification models, and \(n\) is the total number of categories.

4. The method for dynamically identifying attack organizations based on cyber space detection behavior according to claim 1, characterized in that, The threshold determination includes two sub-determination conditions: attack pattern fingerprint determination and average feature vector determination, and the two are in an "or" relationship, that is, if either of them is established, the threshold determination passes.

5. The method for dynamically identifying attack organizations based on cyber space detection behavior according to claim 4, characterized in that, The attack pattern fingerprint determination is expressed as: Among them, APF represents the fingerprint of the test attack mode, and the attack mode is initially classified into C by the multi-classification model i categories, and APF ij represents the j-th attack mode fingerprint in category C i , dist() represents the Euclidean distance between vectors, and d min () represents the minimum distance between APF and all attack mode fingerprints in category C i , D fin represents the fingerprint threshold, and APF is expressed as: APF = {V 1 , V 2 , …, V n} Among them, V i represents the average number of votes for the attack mode belonging to category C i where i = 1, 2, …, n, is expressed as: Among them, v ji represents the feature vector A j belongs to the category C i the number of votes, expressed as: Among them, m k represents the classification result of a binary classification model, and p k represents the probability corresponding to the classification result m k ; I i is an indicator function. When m k is i, the value of I i is 1, otherwise the value is 0.

6. The method for dynamically identifying attack organizations based on cyber space detection behavior according to claim 4, characterized in that, The average feature vector determination is expressed as: Among them, D vec represents the vector threshold, AFV represents the average feature vector of the test attack mode, AFV ij represents the j-th attack mode in category C i The average feature vector, and AFV is expressed as: AFV = {AF 1 , AF 2 , …, AF q} Among them, AF i represents the i-th average attribute of AFV, where i = 1, 2, …, q, and is expressed as: Among them, a ji represents the i-th attribute of the j-th eigenvector of the attack mode.

7. The method for dynamically identifying attack organizations based on cyber space detection behavior according to claim 1, characterized in that, The model determination is specifically: Use the attack pattern as a new category to generate a new binary classification model and add it to the existing multi-classification model; if the expanded multi-classification model can correctly classify this attack pattern, the model determination passes.

8. The method for dynamically identifying attack organizations based on cyber space detection behavior according to claim 1, characterized in that, The attack organization identification mechanism is specifically: 1) Establish an attack organization set AO: AO = {O 1 , O 2 , …, O l} Among them, O j represents an attack organization, j = 1, 2…, l, where l is the number of attack organizations, and O j is expressed as: O j = <C j , S j , D j > Among them, C j represents the label category to which O j belongs, and S j represents the set of C-class network segments included in O j and is expressed as: S j = {s 1 , s 2 , …, s x} Among them, x represents the number of Class C network segments, and D j represents O j the complete set of IP addresses included, expressed as: D j = {d 1 , d 2 , …, d y} where y represents the number of complete IP addresses; 2) For an input attack pattern AP i , extract its Class C network segment s i , complete IP address d i and the classification result c of the multi-classification model i ; 3) Determine s i whether it belongs to an attack organization O in the attack organization set j , if it belongs to an attack organization O j , then determine c i whether it is consistent with O j in terms of category, and execute step 4), otherwise, execute step 8); 4) If c i is consistent with O j in terms of category, then determine whether the category C j of O j is the category to which the open-source cyberspace detection tool belongs, and execute step 5); otherwise, execute step 7). 5) If C j is the category to which the open-source cyberspace detection tool belongs, then determine whether AP i passes the new category determination and execute step 6), otherwise, determine that d i belongs to O j ; 6) If AP i is determined by the new category, the attacking organization O j will be removed from the category of the open-source cyberspace detection tool to which it belongs and merged with AP i to form a new category, and a new binary classification model and attacking organization will be created for it; otherwise, it is determined that d i belongs to O j ; 7) Determine c i whether it is the category to which the open-source cyberspace detection tool belongs. If c i is the category to which the open-source cyberspace detection tool belongs, then determine that d i belongs to O j ; otherwise, merge the attack organization O i corresponding to c i with O j , and at the same time merge the binary classification models corresponding to O i and O j ; 8) Determine AP i Whether it passes the new category determination. If it passes the new category determination, it is AP i Create a new binary classification model and an attack organization; otherwise, judge c i Whether it is the category to which the open-source cyberspace detection tool belongs; 9) If c i is the category to which the open-source cyberspace detection tool belongs, then a new attack organization is created for it based on the C-class network segment s i ; otherwise, it is determined that d i belongs to the attack organization O i corresponding to c i .