Network system intrusion detection method, device, equipment and medium
Through the dynamic ε value loop density clustering algorithm and hash tree storage technology, the whitelist access rules are processed in a layered manner, solving the problems of high rules redundancy and uncontrollable matching space in the whitelist mechanism, and achieving efficient and accurate network intrusion detection.
Patent Information
- Application Number
- CN202510826743.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-19
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-06-19
AI Technical Summary
The existing whitelist mechanism has problems such as high rules redundancy, uncontrollable matching space, and high risk of missed detection in network intrusion detection, making it difficult to reduce system overhead while ensuring detection accuracy and efficiency.
The dynamic ε value loop density clustering algorithm is used to preprocess the whitelist access rules, and the efficient, structured storage and matching of whitelist rules are achieved through hash tree storage and hierarchical matching technology.
It effectively reduces the redundancy of whitelist rules, improves matching efficiency, limits the probability of illegal access missed detection, and provides efficient and accurate intrusion detection.
Smart Images

Figure CN120358083A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of network security technology, and in particular, to a network system intrusion detection method, device, equipment and medium. Background Art
[0002] With the increasing complexity of network attack means, intrusion detection technology (Intrusion Detection System, IDS) has become one of the important means to ensure system security. The core goal of intrusion detection is to monitor system behavior in real time, identify potential malicious activities, and take defensive measures in a timely manner. In a large-scale network service environment, due to the huge business scale and extremely high access traffic, traditional signature-based detection or behavior analysis methods often face problems such as high consumption of computing resources and high response latency, making it difficult to meet actual needs. Therefore, how to reduce system overhead and improve detection efficiency while ensuring detection accuracy has become a key challenge for intrusion detection technology.
[0003] The whitelist mechanism is an efficient intrusion detection strategy. Its core idea is to pre-define legal system behaviors (such as process execution, file access, etc.) and intercept all abnormal operations that do not conform to the rules. Compared with the blacklist mechanism (which only intercepts known malicious behaviors), the whitelist mechanism can effectively defend against unknown attacks, especially suitable for scenarios with high security requirements.
[0004] However, the commonly used whitelist matching methods in the industry still have defects such as too large matching space and high risk of missed detection. Therefore, there is an urgent need for a low-cost solution that can significantly reduce the rule scale and improve the matching efficiency while ensuring detection accuracy. Summary of the Invention
[0005] In view of the defects existing in the prior art, the present invention provides a network system intrusion detection method, device, equipment and medium.
[0006] To achieve the above object, the technical solution adopted by the present invention is as follows: On the one hand, the present invention provides a network system intrusion detection method, including the following steps: S1. Extract the original data, which consists of whitelist access rules, and preprocess the original data to obtain preprocessed data; S2. Process the preprocessed data by using a cyclic density clustering algorithm with a dynamic ε value to obtain a clustering result composed of clustering sets with different densities; S3. Split the strings of the whitelist access rules in the clustering set into program path strings and file path strings according to the tab character, and split the program path strings and file path strings with / to obtain strings at different levels; S4. In the string of each level, identify all the common substrings in sequence and retain them in order, replace the positions of the non-common substrings in the string with wildcards, merge the wildcard replacement results of all levels, and obtain the mask rules of the cluster set; S5. A hash tree is used to store mask rules of all cluster sets, and each layer of the hash tree corresponds to a level of mask rules; S6. extracting the mask rules of the same layer of the hash tree, splitting the mask rules by wildcards to obtain multiple detection groups, wherein the detection group consists of a pair of wildcards and non-wildcards; splitting the string of the access rule to be detected by / to obtain strings of different levels as detection strings; S7. Match the detection groups and detection strings layer by layer according to the hash tree structure. When all detection groups and detection strings in the same layer are matched, the current layer is determined to be matched successfully. When all layers are matched successfully, the access rule to be detected is determined to have passed the detection. Otherwise, the access rule to be detected is considered to be an illegal access rule.
[0007] Furthermore, the original data is preprocessed to obtain preprocessed data, including: S11. For the whitelist access rules with exactly the same content in the original data, only one copy is retained, thereby deduplicating the original data; S12. Unify the path format of the whitelist access rules in the deduplicated data to obtain preprocessed data.
[0008] Furthermore, the preprocessed data is processed using a cyclic density clustering algorithm with a dynamic ε value to obtain a clustering result consisting of cluster sets with different densities, including: S21, determining an initial ε value based on the distance between the whitelist access rules in the preprocessed data; S22, using the current ε value to perform density clustering on the preprocessed data to obtain a cluster set and a discrete outlier set that meet the requirements; S23, increasing the ε value by a preset step value, and returning to S22, until there is no discrete outlier set, and obtaining a clustering result consisting of cluster sets with different densities.
[0009] Furthermore, a hash tree is used to store the mask rules of all cluster sets, including: S51, each layer of the hash tree corresponds to a level of the mask rule, each node stores the string of the level, and the string of the next level is used as a child node of the current node; S52, using tabs as special level nodes; S53. Connect the parent node and the child node through the key-value pair to complete the storage.
[0010] Further, in the strings of each level, all sequential common substrings are identified and retained in order, and the positions of the non-common substrings in the strings are replaced with wildcards. The wildcard replacement results of all levels are merged to obtain the mask rules of the clustering set, and it further includes: For the non-common substrings at the end of the string, retain the trailing characters of the longest substring, and replace the remaining parts with wildcards; For strings that cannot be replaced with wildcards at all, the original strings are retained in a hard-coded manner.
[0011] Further, match the detection group and the detection string layer by layer according to the hash tree structure, including: The detection group preferentially matches the non-wildcard content in the detection string. If the non-wildcard content exists at the corresponding position in the detection string, it is regarded as the current detection group and the detection string passing the match, and the next detection group and the detection string are matched until all detection groups and detection strings are matched.
[0012] Further, match the detection group and the detection string layer by layer according to the hash tree structure, and it further includes: For the last detection group, if the length of the detection string is less than the wildcard length, it is determined that the match passes. If the length of the detection string is greater than the wildcard length, the sliding matching method is used for matching.
[0013] On the other hand, the present invention provides a network system intrusion detection device, including: The first module is used to extract the original data, which consists of whitelist access rules, and preprocess the original data to obtain preprocessed data; The second module is used to process the preprocessed data by using the cyclic density clustering algorithm with a dynamic ε value to obtain a clustering result composed of clustering sets with different densities; The third module is used to split the strings of the whitelist access rules in the clustering set into program path strings and file path strings according to tab characters, and split the program path strings and file path strings with / to obtain strings at different levels; The fourth module is used to identify all sequential common substrings in the strings of each level and retain them in order, replace the positions of the non-common substrings in the strings with wildcards, and merge the wildcard replacement results of all levels to obtain the mask rules of the clustering set; The fifth module is used to store all the mask rules of the clustering set by using a hash tree, and each layer of the hash tree corresponds to a level of the mask rules; The sixth module is used to extract the same-level mask rules of the hash tree, split the mask rules by wildcards to obtain multiple detection groups, where each detection group consists of a pair of wildcards and non-wildcards; split the string of the access rule to be detected by " / ", and obtain different-level strings as detection strings. The seventh module is used to match the detection groups and detection strings layer by layer according to the hash tree structure. When all the detection groups and detection strings in the same level match successfully, it is determined that the current level matches successfully; when all levels match successfully, it is determined that the access rule to be detected passes the detection; otherwise, the access rule to be detected is considered an illegal access rule.
[0014] On the other hand, the present invention provides a computer device, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the network system intrusion detection method are implemented.
[0015] On the other hand, the present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the network system intrusion detection method are implemented.
[0016] Compared with the prior art, the beneficial technical effects of the present invention are as follows: The network system intrusion detection method, device, equipment and medium provided by the present invention process the preprocessed data with a dynamic ε-value cyclic density clustering algorithm to obtain a clustering result composed of clustering sets with different densities, overcoming the defects of the fixed ε-value density clustering algorithm. While ensuring a high similarity within the clustering result, it eliminates the outliers that cannot be clustered, and can better adapt to the actual distribution characteristics of the whitelist rule data; by splitting the strings of the whitelist access rules in the clustering set, after obtaining the path strings and then splitting them to achieve hierarchical division, so as to realize the hierarchical processing of a class of whitelist access rules represented by the clustering set, and only perform string comparison and analysis within the same level, avoiding the complexity and errors brought by cross-level comparison; on the basis of hierarchical division, replace the non-common substrings in the string with wildcards to obtain the mask rules of the clustering set, thereby further compressing the representation space of the same class of whitelist access rules while ensuring the rule accuracy, improving the simplicity and generalization ability of the whitelist rules; then use a hash tree to store and achieve efficient and structured storage of the mask rules of all clustering sets; then process the mask rules and the access rules to be detected to obtain detection groups and detection strings; finally, match the detection groups and detection strings layer by layer according to the hash tree structure to complete the determination of the string to be detected and determine whether the network system is invaded.
[0017] The present invention can effectively solve the problems of high rule redundancy and uncontrollable matching space in the whitelist mechanism, and effectively limit the probability of undetected illegal access; it can be widely applied to scenarios such as enterprise security protection, cloud computing platforms, and Internet of Things devices, providing efficient and accurate intrusion detection for system security. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on the structures shown in these drawings.
[0019] Figure 1 Schematic flow chart of a network system intrusion detection method provided for an embodiment; Figure 2 Schematic flow chart of a cyclic density clustering algorithm with a dynamic ε value provided for an embodiment; Figure 3 Process diagram of obtaining strings at different levels provided for an embodiment, where Figure 3 (a) is a process diagram of splitting the string of the whitelist access rule according to the tab character, Figure 3 (b) is a process diagram of splitting the path string with / ; Figure 4 Schematic diagram of a wildcard replacement process provided for an embodiment; Figure 5 Schematic diagram of sequential common substrings provided for an embodiment; Figure 6 Schematic diagram of the merging process of wildcard replacement results at all levels provided for an embodiment; Figure 7 Schematic diagram of storing the mask rule hash tree of a clustering set provided for an embodiment; Figure 8 Diagram of generating a detection group and a detection string provided for an embodiment; Figure 9 Schematic diagram of the matching process of a detection group and a detection string provided for an embodiment; Figure 10 Schematic diagram of a sliding alignment process provided for an embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0020] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0021] Reference Figure 1 , an embodiment provides a network system intrusion detection method, comprising the following steps: S1. extracting original data, wherein the original data is composed of whitelist access rules, and preprocessing the original data to obtain preprocessed data; S2, using a cyclic density clustering algorithm with a dynamic ε value to process the preprocessed data, and obtaining a clustering result consisting of cluster sets with different densities; S3, splitting the string of the whitelist access rule in the cluster set into a program path string and a file path string according to the tab character, and splitting the program path string and the file path string with / to obtain strings of different levels; S4. In the string of each level, identify all the common substrings in sequence and retain them in order, replace the positions of the non-common substrings in the string with wildcards, merge the wildcard replacement results of all levels, and obtain the mask rules of the cluster set; S5. A hash tree is used to store mask rules of all cluster sets, and each layer of the hash tree corresponds to a level of mask rules; S6. extracting the mask rules of the same layer of the hash tree, splitting the mask rules by wildcards to obtain multiple detection groups, wherein the detection group consists of a pair of wildcards and non-wildcards; splitting the string of the access rule to be detected by / to obtain strings of different levels as detection strings; S7. Match the detection groups and detection strings layer by layer according to the hash tree structure. When all detection groups and detection strings in the same layer are matched, the current layer is determined to be matched successfully. When all layers are matched successfully, the access rule to be detected is determined to have passed the detection. Otherwise, the access rule to be detected is considered to be an illegal access rule.
[0022] The processed data after preprocessing is processed by a cyclic density clustering algorithm with a dynamic ε value to obtain a clustering result composed of clustering sets with different densities, which overcomes the defects of the density clustering algorithm with a fixed ε value. While ensuring a high similarity within the clustering result, it eliminates the outliers that cannot be clustered and can better adapt to the actual distribution characteristics of the whitelist rule data. By splitting the strings of the whitelist access rules in the clustering set, the path strings are obtained and then segmented to achieve hierarchical division, so as to achieve hierarchical processing of a class of whitelist access rules represented by the clustering set, and only compare and analyze the strings within the same level, avoiding the complexity and errors brought by cross-level comparison. On the basis of hierarchical division, the non-common substrings in the strings are replaced with wildcards to obtain the mask rules of the clustering set, so as to further compress the representation space of the same class of whitelist access rules on the basis of ensuring the rule accuracy, and improve the simplicity and generalization ability of the whitelist rules. Then, a hash tree is used to store the mask rules of all clustering sets efficiently and structurally. Then, the mask rules and the access rules to be detected are processed to obtain the detection group and the detection string. Finally, the detection group and the detection string are matched layer by layer according to the hash tree structure to complete the determination of the string to be detected and determine whether the network system has been invaded.
[0023] In a preferred embodiment, the original data is preprocessed to obtain preprocessed data, including: S11. For the whitelist access rules with exactly the same content in the original data, only one copy is retained, so as to remove duplicates from the original data; S12. The path formats of the whitelist access rules in the deduplicated data are unified to obtain preprocessed data.
[0024] By removing duplicates from the original data, the storage space occupied by duplicate data is greatly reduced, saving storage resources. Moreover, through the deduplication process, the computational complexity and running time of the subsequent clustering algorithm can be simplified, enabling more efficient detection. By unifying the path formats of the whitelist access rules in the data, the path formats of the whitelist access rules in the data are made consistent, ensuring the correctness and effectiveness of subsequent calculations.
[0025] In an embodiment, the deduplicated data contains access rules with absolute path and relative path records. By supplementing the relative path record access rules in the deduplicated data with " / " or other prefixes, the relative path record access rules are converted into absolute path record access rules.
[0026] Refer to Figure 2 In a preferred embodiment, the preprocessed data is processed by a cyclic density clustering algorithm with a dynamic ε value to obtain a clustering result composed of clustering sets with different densities, including: S21. Determine an initial ε value based on the distances between the whitelist access rules in the preprocessed data; S22. Perform density clustering on the preprocessed data using the current ε value to obtain a clustering set that meets the requirements and a discrete outlier set; S23. Increase the ε value by a preset step value, return to S22 until there is no discrete outlier set, and obtain a clustering result composed of clustering sets with different densities.
[0027] By determining the initial ε value based on the distances between the whitelist access rules in the preprocessed data, a more appropriate ε value can be obtained to perform density clustering on the preprocessed data, which can effectively balance the similarity of clustering and the number of discrete outliers; then, by increasing the ε value by a preset step value to continue density clustering, through multiple rounds of cyclic clustering until there is no discrete outlier set, all data can be divided into a clustering set, thus better adapting to the actual distribution characteristics of the whitelist rule data and ensuring the accuracy of subsequent mask rule generation.
[0028] Refer to Figure 3 , in one embodiment, a process diagram for obtaining strings at different levels is provided. Among them, Figure 3 (a) is a process diagram for splitting the string of the whitelist access rule by tab characters. As shown in Figure 3 (a), / usr / bin / id is a program path string, and / lib64 / libselinux.so.1, etc. are file path strings; Figure 3 (b) is a process diagram for splitting the path string by / . As shown in Figure 3 (b), / lib64 / libselinux.so.1 is split into two levels: "lib64" and "libselinux.so.1". By restricting the comparison and analysis range of the strings, not only can the operation efficiency be effectively improved, but also the generated mask rules can be more accurate and reasonable in structure.
[0029] Refer to Figure 4 , Figure 5 , Figure 6 , one embodiment provides the specific process of the mask rule for the clustering set. As shown in Figure 4 , use as a wildcard, and it is stipulated that each can match at most one character; in the string at each level, identify all sequential common substrings and retain them in order, and use wildcards to replace the positions of non-common substrings in the string, thereby achieving the maximum compression of the string at this level. As shown in Figure 6 , merge the wildcard replacement results of all levels to obtain the unified mask representation of this type of whitelist access rule, that is, obtain the mask rule of the clustering set.
[0030] The sequential common substring is as Figure 5 shown, and to ensure the accuracy of the mask rule, only the longest common substring that appears sequentially is retained. For example, Figure 5 shown, for libselinux.so.1 and.so.clib6, the sequential common substring is.so., rather than lib and.so.; for non-sequential common substrings, on the premise of ensuring the order, the longer substring is preferably retained.
[0031] In a preferred embodiment, in the strings of each level, all sequential common substrings are identified and retained in order, and wildcards are used to replace the positions of non-common substrings in the strings. The wildcard replacement results of all levels are combined to obtain the mask rule of the clustering set, and it further includes: For the non-common substring at the end of the string, retain the tail character of the longest substring, and use wildcards to replace the rest; thus, more original information can be retained as much as possible, and the character space can be reduced.
[0032] For strings that cannot be replaced by wildcards at all, the original strings are retained in a hard-coded manner. Thereby, the accuracy and simplicity of the rule are improved.
[0033] Referring to Figure 7 , in a preferred embodiment, a hash tree is used to store the mask rules of all clustering sets, including: S51. Each layer of the hash tree corresponds to a level of the mask rule, and each node stores the string of this level, and the strings of the next layer are used as the child nodes of the current node; S52. Use the tab character as a special level node; S53. Connect the parent node and the child node through key-value pairs to complete the storage.
[0034] Using a hash tree to store the mask rules of all clustering sets, since the hash tree supports hierarchical access at the 0(1) level, it can quickly locate the target rule and greatly shorten the matching time; furthermore, hierarchical storage makes the rule organization more orderly, facilitating maintenance and expansion; on the other hand, rules with the same prefix can share nodes, thereby reducing redundant storage; at the same time, the hash tree supports dynamic insertion, deletion, and modification of rules, and can adapt to the real-time update requirements of the whitelist rules.
[0035] Referring to Figure 8 , in an embodiment, the mask rules of the same layer of the hash tree are extracted, and the mask rules are split by wildcards to obtain multiple detection groups. The detection group consists of a pair of wildcards and non-wildcards, and the lengths of the wildcards and non-wildcards can be 0; the string of the access rule to be detected is split by / to obtain strings of different levels as the detection strings.
[0036] The mask rule is disassembled into multiple detection groups, so as to achieve efficient and segmented matching of the access rule to be detected, and improve the flexibility and accuracy of detection.
[0037] In a preferred embodiment, the detection groups and the detection strings are matched layer by layer according to the hash tree structure, including: The detection group preferentially matches the non-wildcard content in the detection string. If the non-wildcard content exists at the corresponding position in the detection string, it is regarded that the current detection group and the detection string pass the match, and the next detection group and the detection string are matched until all detection groups and detection strings are matched. If the length of the subsequent string is 0, it means that there is no content to be matched, and it is determined to pass the match.
[0038] Referring to Figure 9 , the mask rule is , and the string to be detected is libexif.so.1. In the first group match, the detection group is [, lib ], that is, find lib within the first 3 characters of the string to be detected (i.e., the detection string) libexif.so.1. The result exists and the match passes. In the second group match, the detection group is , that is, find within the first 11 characters of the string to be detected exif.so.1. The result exists and the match passes. In the third group match, the detection group is , that is, find any character within the first 1 character of the string to be detected 1. The result exists and the match passes. Through the above matching, the string to be detected successfully matches the given mask rule.
[0039] In a preferred embodiment, the detection groups and the detection strings are matched layer by layer according to the hash tree structure, and it also includes: For the last detection group, if the length of the detection string is less than the wildcard length, it is determined to pass the match. If the length of the detection string is greater than the wildcard length, the sliding matching method is used for matching. Through the sliding matching method, it is ensured that accurate matching can still be completed when there are changes at the end of the string.
[0040] Referring to Figure 10 , the last detection group is , and the string to be detected is 012. At this time, 123 cannot be found in the first 6 characters of 012, and the length of 012 also exceeds the length represented by the wildcard. In this case, the algorithm allows to replace the first several characters in the detection string to achieve alignment. Slide the detection string (012) from left to right. When the end of the detection string can be aligned with the front part of the non-wildcard string (123) in the detection group, it is regarded as meeting the matching rule and the match passes.
[0041] One embodiment provides a network system intrusion detection device, including: A first module, configured to extract original data, which consists of whitelist access rules, preprocess the original data to obtain preprocessed data; A second module, configured to process the preprocessed data by using a cyclic density clustering algorithm with a dynamic ε value to obtain a clustering result composed of clustering sets with different densities; A third module, configured to split the strings of the whitelist access rules in the clustering set into program path strings and file path strings according to tab characters, and split the program path strings and file path strings with / to obtain strings at different levels; A fourth module, configured to identify all sequential common substrings in the strings at each level and retain them in order, replace the positions of non-common substrings in the strings with wildcards, and merge the wildcard replacement results at all levels to obtain the mask rules of the clustering set; A fifth module, configured to store the mask rules of all clustering sets by using a hash tree, and each layer of the hash tree corresponds to a level of the mask rules; A sixth module, configured to extract the mask rules of the same layer of the hash tree, split the mask rules according to wildcards to obtain multiple detection groups, where each detection group consists of a pair of wildcards and non-wildcards; split the string of the access rule to be detected with / to obtain strings at different levels as detection strings; A seventh module, configured to match the detection groups and the detection strings layer by layer according to the hash tree structure. When all detection groups and detection strings in the same layer match successfully, it is determined that the current layer matches successfully; when all layers match successfully, it is determined that the access rule to be detected passes the detection; otherwise, the access rule to be detected is considered an illegal access rule.
[0042] On the other hand, the present invention provides a computer device, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the network system intrusion detection method provided in any one of the above embodiments are implemented. The computer device may be a server. The computer device includes a processor, a memory, a network interface, and a database connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store sample data. The network interface of the computer device is used to communicate with an external terminal through a network connection.
[0043] On the other hand, the present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the network system intrusion detection method provided in any of the above embodiments are implemented.
[0044] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database or other medium used in the embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.
[0045] Matters not described in the present invention are well-known technologies.
[0046] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0047] The above-described embodiments merely represent several implementation manners of the present application. Their descriptions are relatively specific and detailed, but they should not be construed as limiting the scope of the invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.
[0048] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A network system intrusion detection method, characterized in that, It includes the following steps: S1. Extract the original data, which consists of whitelist access rules, and preprocess the original data to obtain preprocessed data; S2. Process the preprocessed data using a cyclic density clustering algorithm with a dynamic ε value to obtain a clustering result composed of clustering sets with different densities; S3. Split the strings of the whitelist access rules in the clustering set into program path strings and file path strings according to tab characters, and split the program path strings and file path strings with / to obtain strings at different levels; S4. In the strings at each level, identify all sequential common substrings and retain them in order, replace the positions of non-common substrings in the strings with wildcards, and merge the wildcard replacement results at all levels to obtain the mask rules of the clustering set; S5. Store the mask rules of all clustering sets using a hash tree, where each layer of the hash tree corresponds to a level of the mask rules; S6. Extract the mask rules at the same level of the hash tree, split the mask rules by wildcards to obtain multiple detection groups, where each detection group consists of a pair of wildcards and non-wildcards; split the string of the access rule to be detected with / to obtain strings at different levels as detection strings; S7. Match the detection groups and detection strings layer by layer according to the hash tree structure. When all detection groups and detection strings at the same level match successfully, it is determined that the current level matches successfully; when all levels match successfully, it is determined that the access rule to be detected passes the detection; otherwise, the access rule to be detected is considered an illegal access rule.
2. The network system intrusion detection method according to claim 1, characterized in that, Preprocess the original data to obtain preprocessed data, including: S11. Retain only one copy of the whitelist access rules with exactly the same content in the original data, thereby removing duplicates from the original data; S12. Unify the path formats of the whitelist access rules in the deduplicated data to obtain preprocessed data.
3. The network system intrusion detection method according to claim 1, wherein, Process the preprocessed data using a cyclic density clustering algorithm with a dynamic ε value to obtain a clustering result composed of clustering sets with different densities, including: S21. Determine the initial ε value based on the distances between the whitelist access rules in the preprocessed data; S22. Perform density clustering on the preprocessed data using the current ε value to obtain a clustering set that meets the requirements and a discrete outlier set; S23. Increase the ε value by a preset step value, return to S22 until there is no discrete outlier set, and obtain a clustering result composed of clustering sets with different densities.
4. The network system intrusion detection method according to claim 1, characterized in that Store the mask rules of all clustering sets using a hash tree, including: S51. Let each layer of the hash tree correspond to a level of the mask rules, and each node stores the string at that level, and the strings at the next level are used as the child nodes of the current node; S52. Use the tab character as a special level node; S53. Connect the parent node and the child node through key-value pairs to complete the storage.
5. The network system intrusion detection method according to claim 1, characterized in that, In the strings at each level, identify all sequential common substrings and retain them in order, replace the positions of non-common substrings in the strings with wildcards, and merge the wildcard replacement results at all levels to obtain the mask rules of the clustering set, and also include: For the non-common substrings at the end of the string, retain the trailing characters of the longest substring, and replace the remaining parts with wildcards; For strings that cannot be replaced with wildcards at all, retain the original strings in a hard-coded manner.
6. The network system intrusion detection method according to claim 1, wherein, Match the detection group and the detection string layer by layer according to the hash tree structure, including: The detection group preferentially matches the non-wildcard content in the detection string. If the non-wildcard content exists at the corresponding position in the detection string, it is considered that the current detection group and the detection string match successfully, and the next detection group and the detection string are matched until all detection groups and detection strings are matched.
7. The network system intrusion detection method according to claim 6, wherein Match the detection group and the detection string layer by layer according to the hash tree structure, and also include: For the last detection group, if the length of the detection string is less than the wildcard length, it is determined to match successfully. If the length of the detection string is greater than the wildcard length, the sliding matching method is used for matching.
8. Network system intrusion detection device, characterized in that, Include: The first module is used to extract the original data, which consists of whitelist access rules, and preprocess the original data to obtain preprocessed data; The second module is used to process the preprocessed data by using the cyclic density clustering algorithm with a dynamic ε value to obtain a clustering result composed of clustering sets with different densities; The third module is used to split the strings of the whitelist access rules in the clustering set into program path strings and file path strings according to the tab character, and split the program path strings and file path strings with / to obtain strings at different levels; The fourth module is used to identify all sequential common substrings in the strings at each level and retain them in order, replace the positions of the non-common substrings in the strings with wildcards, and merge the wildcard replacement results at all levels to obtain the mask rules of the clustering set; The fifth module is used to store the mask rules of all clustering sets by using a hash tree, and each layer of the hash tree corresponds to a level of the mask rules; The sixth module is used to extract the mask rules of the same layer of the hash tree, split the mask rules by wildcards to obtain multiple detection groups, and the detection groups consist of a pair of wildcards and non-wildcards; split the string of the access rule to be detected with / to obtain strings at different levels as the detection string; The seventh module is used to match the detection group and the detection string layer by layer according to the hash tree structure. When all detection groups and detection strings in the same layer match successfully, it is determined that the current layer matches successfully; when all layers match successfully, it is determined that the access rule to be detected passes the detection; otherwise, the access rule to be detected is considered an illegal access rule.
9. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the network system intrusion detection method described in any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, A computer program is stored thereon, and when the computer program is executed by the processor, it implements the steps of the network system intrusion detection method described in any one of claims 1-7.
Citation Information
Patent Citations
Sensitive word filtering method and device, computer equipment and storage medium
CN109684469A
Intrusion detection method and device based on edge cloud environment
CN114826690A
Turbine set prediction method based on density clustering and cyclic fuzzy neural network
CN116227570A
Systems And Methods For Identifying Potential Duplicate Entries In A Database
US20130132410A1