Data fuzzy matching method, device and equipment

By using decision trees and target masks in security devices for fuzzy matching of data, the problems of long and low matching time and low performance in the prior art are solved, and fast and efficient matching performance is achieved.

CN119966946AActive Publication Date: 2025-05-09NEW H3C TECH CO LTD

Patent Information

Application Number
CN202510125379.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-26
Publication Date
2025-05-09
Estimated Expiration
2045-01-26

AI Technical Summary

Technical Problem

When the prior art issues a large number of fuzzy domain names, the matching time is long and the matching performance is low, and it cannot meet the actual needs.

Method used

By obtaining the target matching data, fuzzy matching data are generated, and a decision tree is used to query the target matching rules and masks from the leaf nodes, thereby reducing the number of matches and improving matching performance.

Benefits of technology

It realizes fast matching, reduces matching time, improves matching performance, and can meet actual needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119966946A_ABST
    Figure CN119966946A_ABST
Patent Text Reader

Abstract

The invention provides a data fuzzy matching method, device and equipment, and the method comprises the steps: obtaining target matching data, and obtaining fuzzy matching data according to the target matching data; taking a first data bit of the fuzzy matching data as a current data bit, and taking continuous K data bits starting from the current data bit as current to-be-matched data; querying a target leaf node corresponding to the current to-be-matched data from all leaf nodes of the established decision tree; selecting M data bits from the current data to be matched, and selecting M rule bits from the K rule bits; and if the value of each bit in the M data bits is the same as the value of the corresponding rule bit in the M rule bits, processing corresponding to successful matching is executed based on the target matching data. According to the scheme, the matching performance can be improved, and the matching time can be shortened.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of communication technology, and in particular to a data fuzzy matching method, device and equipment. Background Art

[0002] On security devices (such as firewall devices), a large number of fuzzy domain names are sent in the form of substrings. When the security device receives a DNS request message sent by a DNS (Domain Name System) server, it can resolve the domain name to be matched and the IP address (that is, the IP address corresponding to the domain name to be matched) from the DNS request message. If the domain name to be matched matches any fuzzy domain name, the IP address is stored. If the domain name to be matched fails to match all fuzzy domain names, the IP address is not stored. On this basis, the stored IP address can be used as a legal IP address, and the security device allows data packets that access the legal IP address and discards data packets that access the illegal IP address, thereby achieving access control of the data packets.

[0003] For example, if the fuzzy domain name includes the regular string "xiao", if the domain name to be matched includes the data string "xiaoyi", then the data string "xiaoyi" matches the regular string "xiao". If the domain name to be matched includes the data string "xiamen", then the data string "xiamen" does not match the regular string "xiao". Alternatively, if the fuzzy domain name includes the regular string "xia", then both the data string "xiaoyi" and the data string "xiamen" match the regular string "xia".

[0004] However, when a large number of fuzzy domain names are issued, the domain name to be matched needs to be matched with each fuzzy domain name in turn, which takes a long time and has low matching performance, and the matching performance cannot meet actual needs. Summary of the Invention

[0005] The present application provides a data fuzzy matching method, the method comprising:

[0006] Obtaining target matching data, and obtaining fuzzy matching data based on the target matching data;

[0007] The first data bit of the fuzzy matching data is used as the current data bit, and the consecutive K data bits starting from the current data bit are used as the current data to be matched;

[0008] Querying a target leaf node corresponding to the current to-be-matched data from all leaf nodes of the established decision tree, the target leaf node being used to record a target matching rule and a target mask; wherein the target matching rule includes K rule bits, and the target mask is used to indicate M rule bits among the K rule bits as exact matching bits;

[0009] Selecting M data bits from the current data to be matched, and selecting M rule bits from the K rule bits;

[0010] If the value of each bit in the M data bits is the same as the value of the corresponding rule bit in the M rule bits, performing a process corresponding to a successful match based on the target matching data;

[0011] If the value of each bit in the M data bits is different from the value of the corresponding rule bit in the M rule bits, the data bit after the current data bit is used as the current data bit, and the operation of selecting M data bits from the current data to be matched and selecting M rule bits from the K rule bits is repeated until the value of each bit in the M data bits is the same as the value of the corresponding rule bit in the M rule bits.

[0012] The present application provides a data fuzzy matching device, the device comprising:

[0013] A generation module is used to obtain target matching data and obtain fuzzy matching data based on the target matching data; the first data bit of the fuzzy matching data is used as the current data bit, and the consecutive K data bits starting from the current data bit are used as the current data to be matched;

[0014] An acquisition module is configured to query a target leaf node corresponding to the current to-be-matched data from all leaf nodes of the established decision tree, wherein the target leaf node is configured to record a target matching rule and a target mask; wherein the target matching rule includes K rule bits, and the target mask is configured to indicate M rule bits among the K rule bits as exact matching bits;

[0015] A determination module is used to select M data bits from the current data to be matched, and select M rule bits from the K rule bits; if the value of each bit in the M data bits is the same as the value of the corresponding rule bit in the M rule bits, then perform the corresponding processing of the successful match based on the target matching data; if the value of each bit in the M data bits is different from the value of the corresponding rule bit in the M rule bits, then use the next data bit after the current data bit as the current data bit, and the generation module uses the consecutive K data bits starting from the current data bit as the current data to be matched.

[0016] The present application provides an electronic device, comprising: a processor and a machine-readable storage medium, wherein the machine-readable storage medium stores machine-executable instructions that can be executed by the processor; the processor is used to execute the machine-executable instructions to implement the data fuzzy matching method of the above example of the present application.

[0017] The present application provides a computer program product, which includes a computer program. When the computer program is executed by a processor, it implements the data fuzzy matching method of the above example of the present application.

[0018] The present application provides a machine-readable storage medium, which stores machine-executable instructions that can be executed by a processor; wherein the processor is used to execute the machine-executable instructions to implement the data fuzzy matching method of the above example of the present application.

[0019] As can be seen from the above technical solutions, in the embodiment of the present application, the target leaf node corresponding to the current data to be matched is queried from all leaf nodes of the decision tree, and the current data to be matched is matched with the target matching rule in the target leaf node, without matching the current data to be matched with all matching rules, thereby reducing the number of matches, improving matching performance, reducing matching time, and matching performance can meet actual needs. For example, even if there are tens of thousands of matching rules, it is only necessary to match the current data to be matched with several matching rules in the target leaf node, which greatly reduces the number of matches. In addition, the exact matching bits are indicated by the target mask, so that only the exact matching bits need to be matched, further improving matching performance and reducing matching time. By splitting the target matching data and querying it multiple times, fuzzy matching with unfixed exact matching field positions can be achieved, thereby significantly improving the performance of fuzzy matching of massive data with unfixed exact matching field positions, and significantly improving the matching performance under massive rules, and significantly improving the performance of accurate and fuzzy matching, and making the matching performance more stable. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1 This is a flow chart of a data fuzzy matching method in one embodiment of the present application;

[0021] Figure 2 It is a flowchart of a decision tree construction method in one embodiment of the present application;

[0022] Figure 3 This is a schematic diagram of left-aligning and continuously arranging exact matching bits in one embodiment of the present application;

[0023] Figure 4A This is a schematic diagram of dividing matching rules into leaf nodes in one embodiment of the present application;

[0024] Figure 4B This is a schematic diagram of dividing matching rules into leaf nodes in one embodiment of the present application;

[0025] Figure 4C is a schematic diagram of a four-level decision tree in one embodiment of the present application;

[0026] Figure 5A is a schematic diagram of a target mask in one embodiment of the present application;

[0027] Figure 5B is a schematic diagram of a secondary mask in one embodiment of the present application;

[0028] Figure 6 This is a flow chart of a data fuzzy matching method in one embodiment of the present application;

[0029] Figure 7 This is a schematic diagram of searching for leaf nodes in one embodiment of the present application;

[0030] Figure 8 is a schematic diagram of fuzzy matching data in one embodiment of the present application;

[0031] Figure 9A This is a structural diagram of a data fuzzy matching device in one embodiment of the present application;

[0032] Figure 9B It is a hardware structure diagram of an electronic device in one embodiment of the present application. DETAILED DESCRIPTION

[0033] In the embodiment of the present application, a data fuzzy matching method is proposed, which can be applied to electronic devices, which can be security devices (such as firewall devices, etc.) or network devices (such as routers, switches, etc.). Figure 1 FIG. 5 is a flow chart of the method, which may include:

[0034] Step 101: Obtain target matching data, and obtain fuzzy matching data based on the target matching data. For example, the target matching data includes multiple data bits (such as exact matching bits and special character bits), and the fuzzy matching data includes multiple data bits and padding bits.

[0035] Step 102: The first data bit of the fuzzy matching data is used as the current data bit, and the K consecutive data bits starting from the current data bit are used as the current data to be matched.

[0036] Step 103: Query the target leaf node corresponding to the current data to be matched from all leaf nodes of the established decision tree. The target leaf node is used to record the target matching rule and the target mask.

[0037] In an example, the target matching rule includes K rule bits, and the target mask is used to indicate M rule bits among the K rule bits as exact matching bits.

[0038] Step 104: Select M data bits from the current data to be matched (e.g., select M data bits based on the indication information of the target mask), and select M rule bits from the K rule bits (i.e., the K rule bits of the target matching rule) (e.g., select M rule bits based on the indication information of the target mask), and determine whether the value of each bit in the M data bits is the same as the value of the corresponding rule bit in the M rule bits. If so, that is, the value of each bit in the M data bits is the same as the value of the corresponding rule bit in the M rule bits, then step 105 can be executed; if not, that is, the value of each bit in the M data bits is different from the value of the corresponding rule bit in the M rule bits, then step 106 can be executed.

[0039] Step 105: Execute the corresponding processing based on the target matching data. For example, obtain the IP address corresponding to the target matching data from the DNS request message and store the IP address in the security table entry. In this way, data packets accessing the IP address can pass through the security device, thereby controlling the access of the data packets.

[0040] Step 106: Use the data bit following the current data bit as the current data bit, and repeat the operation of using the consecutive K data bits starting from the current data bit as the current data to be matched, selecting M data bits from the current data to be matched, and selecting M rule bits from the K rule bits, until the value of each bit in the M data bits is the same as the value of the corresponding rule bit in the M rule bits, that is, repeating steps 102 to 106.

[0041] In one example, fuzzy matching data is obtained based on target matching data, which may include but is not limited to: determining the number of padding bits based on the number of exact matching bits and the minimum valid length in the target matching data; wherein the minimum valid length is the minimum value of the valid lengths of all matching rules corresponding to the decision tree; and adding the number of padding bits after the target matching data to obtain fuzzy matching data, wherein the padding bits are special character bits that will not appear in the target matching rules.

[0042] In one example, if the value of each bit in the M data bits is different from the value of the corresponding rule bit in the M rule bits, before using the next data bit after the current data bit as the current data bit, it is also possible to determine whether the number of matches corresponding to the target matching data is the sum of the number and 1; wherein the initial value of the number of matches corresponding to the target matching data is a preset value (such as 0). If not, the number of matches corresponding to the target matching data is increased by 1, and the next data bit after the current data bit is used as the current data bit; if so, the corresponding processing of the match failure is performed based on the target matching data. For example, when obtaining the IP address corresponding to the target matching data from the DNS request message, it is prohibited to store the IP address in the security table entry, so that the data message accessing the IP address cannot pass through the security device.

[0043] In one example, before querying the target leaf node corresponding to the current data to be matched from all leaf nodes of the established decision tree, the process of establishing the decision tree may include, but is not limited to: obtaining multiple matching rules, each of which may include K rule bits; based on the BSS value of each rule bit, using the rule bit with the maximum BSS value as the reference rule bit; dividing all matching rules into leaf nodes based on the reference rule bits to obtain the current decision tree; judging whether the current decision tree meets the end-division condition. If not, based on the BSS values ​​of the remaining rule bits other than the reference rule bit, using the rule bit with the maximum BSS value among the BSS values ​​of the remaining rule bits as the reference rule bit again; dividing all matching rules into multiple leaf nodes based on all reference rule bits to obtain the current decision tree, and repeatedly performing the operation of judging whether the current decision tree meets the end-division condition until the current decision tree meets the end-division condition and stops. If so, the current decision tree is determined to be the established decision tree, and a bit mask is generated for the decision tree, wherein the bit mask is used to indicate all reference rule bits.

[0044] In one example, the conditions for ending the division may include, but are not limited to: the target spatial factor of the current decision tree is not less than the spatial factor threshold; wherein the spatial factor threshold may be the maximum value that the configured target spatial factor can reach; or, the spatial factor threshold may be the configured maximum floating-point number. The number of leaf node rule references of the current decision tree is not greater than the leaf node rule number threshold; wherein the leaf node rule number threshold may be a configured specified value (such as 1). The matching rules within each leaf node of the current decision tree cannot be further divided into two leaf nodes.

[0045] Determining whether the current decision tree meets the end-segmentation condition may include: if the current decision tree meets any one of the end-segmentation conditions, determining that the current decision tree meets the end-segmentation condition; if the current decision tree does not meet each of the end-segmentation conditions, determining that the current decision tree does not meet the end-segmentation condition.

[0046] In an example, obtaining multiple matching rules may include but is not limited to: obtaining multiple configured original rules; for each original rule, if the number of rule bits in the original rule is less than K, wildcard bits may be supplemented to the original rule to obtain a supplemented rule, and the number of rule bits in the supplemented rule may be equal to K, K may be the nth power of 2, and n may be a positive integer; if the supplemented rule includes exact matching bits and general matching bits, all exact matching bits are left-aligned and arranged continuously or right-aligned and arranged continuously to obtain a matching rule corresponding to the original rule.

[0047] Obtaining target matching data may include but is not limited to: obtaining the original matching data that has been input; if the number of data bits in the original matching data is less than K, supplementing the original matching data with special character bits to obtain supplemented matching data, and the number of data bits in the supplemented matching data is equal to K; if the supplemented matching data includes exact matching bits and special character bits (special character bits refer to bits of characters that will not appear in the matching rules), all exact matching bits are left-aligned and arranged continuously or right-aligned and arranged continuously to obtain target matching data.

[0048] In one example, a matching rule in a leaf node is configured with a target mask, and the target mask includes a first-level mask and a second-level mask; wherein the K rule bits of the matching rule are divided into P1 data blocks, the first-level mask includes P1 bits, and the P1 bits correspond one-to-one to the P1 data blocks; for each data block, when there is an exact matching bit in the data block, the value of the bit corresponding to the data block in the first-level mask is a first value, and the first value indicates that there is an exact matching bit in the data block; when there is no exact matching bit in the data block, the value of the bit corresponding to the data block in the first-level mask is a second value, and the second value indicates that there is no exact matching bit in the data block;

[0049] Wherein, for each data block, when there is no exact matching bit in the data block, the data block does not correspond to the secondary mask; when there is an exact matching bit in the data block, the data block corresponds to the secondary mask;

[0050] Wherein, when the data block includes P2 regular bits, the secondary mask includes P2 bits, and the P2 bits correspond one-to-one to the P2 regular bits;

[0051] For each rule bit, when the rule bit is an exact match bit, the value of the bit corresponding to the rule bit in the secondary mask is the third value, and the third value represents the exact match bit; when the rule bit is a general match bit, the value of the bit corresponding to the rule bit in the secondary mask is the fourth value, and the fourth value represents the general match bit.

[0052] In an example, if a first matching rule and a second matching rule exist in a leaf node, and the rule scope of the first matching rule is larger than the rule scope of the second matching rule, then the position of the second matching rule in the leaf node is located before the position of the first matching rule in the leaf node.

[0053] As can be seen from the above technical solutions, in the embodiment of the present application, the target leaf node corresponding to the current data to be matched is queried from all leaf nodes of the decision tree, and the current data to be matched is matched with the target matching rule in the target leaf node, without matching the current data to be matched with all matching rules, thereby reducing the number of matches, improving matching performance, reducing matching time, and matching performance can meet actual needs. For example, even if there are tens of thousands of matching rules, it is only necessary to match the current data to be matched with several matching rules in the target leaf node, which greatly reduces the number of matches. In addition, the exact matching bits are indicated by the target mask, so that only the exact matching bits need to be matched, further improving matching performance and reducing matching time. By splitting the target matching data and querying it multiple times, fuzzy matching with unfixed exact matching field positions can be achieved, thereby significantly improving the performance of fuzzy matching of massive data with unfixed exact matching field positions, and significantly improving the matching performance under massive rules, and significantly improving the performance of accurate and fuzzy matching, and making the matching performance more stable.

[0054] The above technical solutions of the embodiments of the present application are described below in conjunction with specific application scenarios.

[0055] The strstr function (used to return the address of the first occurrence of a substring in a string) and the strcasestr function (used to find another string within a string) are used for string substring matching. For example, based on the strstr or strcasestr function, for the regular string "xiao", the data string "xiaoyi" matches the regular string "xiao", but the data string "xiamen" does not match the regular string "xiao". For the regular string "xia", the data string "xiaoyi" matches the regular string "xia", and the data string "xiamen" matches the regular string "xia".

[0056] On a security device (such as a firewall device), a large number of fuzzy domain name rules (i.e., part of the full domain name) and a large number of precise domain name rules (i.e., the full domain name) can be issued in the form of substrings, that is, a mixed configuration of fuzzy domain name rules and precise domain name rules is supported. When the security device receives a DNS request message sent by a DNS server, it can parse the domain name and IP address to be matched from the DNS request message. If the domain name to be matched successfully matches any fuzzy domain name rule or precise domain name rule (such as using the function strstr or the function strcasestr for matching), the IP address is stored in the security table entry. If the domain name to be matched fails to match all fuzzy domain name rules and precise domain name rules, the IP address is not stored in the security table entry.

[0057] On this basis, the security device allows data packets that access the IP addresses in the security table and discards data packets that access the IP addresses outside the security table, thereby implementing access control for data packets.

[0058] In addition to issuing a large number of fuzzy domain name rules and a large number of precise domain name rules in the form of substrings, that is, the matching rules include fuzzy domain name rules and precise domain name rules, other types of matching rules can also be issued in the form of substrings. They can be any matching rules that need to improve the performance of string substring matching. There is no restriction on the type of matching rules. For the convenience of description, in the subsequent embodiments, domain name rules are used as an example for explanation.

[0059] In the embodiment of the present application, a method for fuzzy matching of massive data with no fixed location is proposed, which may involve a decision tree construction process and a data fuzzy matching process based on the decision tree. The decision tree construction process and the data fuzzy matching process based on the decision tree are described below with reference to specific examples.

[0060] Regarding the decision tree construction process, a decision tree construction method is proposed in the embodiment of the present application, see Figure 2 FIG. 1 is a flow chart of a decision tree construction method, which may include:

[0061] Step 201: Acquire multiple matching rules, each matching rule including K rule bits.

[0062] For example, K may be a positive integer, such as K may be a power of 2, n may be a positive integer, or K may not be a power of 2. For example, when n is 1, K is 2, when n is 2, K is 4, when n is 3, K is 8, when n is 4, K is 16, when n is 8, K is 256, and so on.

[0063] For each matching rule, the matching rule may include K bits, and the bits in the matching rule may be called rule bits. Therefore, the matching rule may include K rule bits.

[0064] In one example, for step 201, the following steps may be used to obtain multiple matching rules:

[0065] Step 2011: Acquire multiple configured original rules. For each original rule, the original rule may include multiple rule bits, and the number of rule bits in the original rule is less than or equal to K.

[0066] For example, the original rules may be pre-configured, and the original rules may be pre-configured by the user, or the original rules may be pre-configured using a certain algorithm, and there is no restriction on the source of the original rules.

[0067] For example, the original rule may contain three protocol number rules: original rule 1: protocol number 38, original rule 2: protocol number 46, and original rule 3: protocol number 54 or 62. These original rules can be converted into two-process bit format as follows: original rule 1: 0010 0110, original rule 2: 0010 1110, and original rule 3: 0011*110. In these original rules, 0 represents bit 0, 1 represents bit 1, and * represents a wildcard, meaning it can represent either bit 0 or bit 1.

[0068] Step 2012: For each original rule, if the number of rule bits in the original rule is less than K, wildcard bits may be added to the original rule to obtain a supplemented rule, and the number of rule bits in the supplemented rule may be equal to K. Alternatively, if the number of rule bits in the original rule is equal to K, the original rule remains unchanged, that is, the original rule may be used as the supplemented rule.

[0069] For example, taking the domain name rule as the original rule, since the maximum length of a domain name is 253 bytes, the value of K that is closest to the maximum length of 253 bytes of the domain name is 256 bytes (K needs to be greater than or equal to the maximum length of the domain name, and K is 2 to the power of n). Therefore, for each original rule, the wildcard bit can be supplemented to the original rule to obtain the supplemented rule, and the number of rule bits in the supplemented rule can be equal to 256 bytes, that is, there are 2048 rule bits.

[0070] For example, when the wildcard bit is supplemented to the original rule to obtain the supplemented rule, the wildcard bit represents the bit of the wildcard character. The wildcard character is a special statement that can be in the form of an asterisk (*) and a question mark (?). The wildcard character is used for fuzzy search and indicates that it can match all characters.

[0071] For example, on a 32-bit system or a 64-bit system, the performance of operations based on 4 bytes or 8 bytes is the highest. Therefore, the original rule is supplemented to 2 to the power of 8, that is, 256 bytes, to obtain the supplemented rule. In this way, the performance of operations on the supplemented rule is the highest.

[0072] Step 2013: If the supplemented rule includes exact matching bits and general matching bits, all exact matching bits are left-aligned and arranged consecutively to obtain a matching rule corresponding to the original rule.

[0073] Alternatively, if the supplemented rule includes exact matching bits and general matching bits, all exact matching bits are right-aligned and arranged consecutively to obtain a matching rule corresponding to the original rule.

[0074] Obviously, for each original rule, after preprocessing the original rule (for example, supplementing wildcard bit processing and / or aligning continuous arrangement processing, etc.), the matching rule corresponding to the original rule can be obtained. In this way, multiple matching rules corresponding to multiple original rules can be obtained.

[0075] For example, if the post-supplement rule is a domain name post-supplement rule, the domain name post-supplement rule is 256 bytes, and the domain name post-supplement rule can be a fuzzy domain name rule or a precise domain name rule. If the domain name post-supplement rule includes exact match bits and universal match bits, then all exact match bits in the domain name post-supplement rule are left-aligned and arranged consecutively to ensure that within a single domain name post-supplement rule, there is no universal match bit to the left of the first exact match bit, and all universal match bits are on the right.

[0076] See also Figure 3The figure shows a diagram of all exact match bits being aligned left and arranged consecutively. For domain name supplement rule 1, abcde represents five exact match characters. The bits corresponding to these five exact match characters can be used as exact match bits. The exact match bits corresponding to these five exact match characters need to be aligned left and arranged consecutively. ***…*** (251 *) represents 251 wildcard characters. The bits corresponding to these 251 wildcard characters can be used as universal match bits. The universal match bits corresponding to these 251 wildcard characters can be located to the right of the last exact match bit. For domain name supplement rule 2, an can represent two exact match characters. The exact match bits corresponding to these two exact match characters are aligned left and arranged consecutively. The universal match bits corresponding to 254 wildcard characters can be located to the right of the last exact match bit. For domain name supplement rule 3, gaagbb can represent six exact match characters. The exact match bits corresponding to these six exact match characters are aligned left and arranged consecutively. The universal matching bits corresponding to the 250 wildcard characters can be located to the right of the last exact matching bit.

[0077] Alternatively, taking the domain name supplement rule as an example, the domain name supplement rule is 256 bytes, and the domain name supplement rule can be a fuzzy domain name rule or a precise domain name rule. If the domain name supplement rule includes exact match bits and universal match bits, then all exact match bits in the domain name supplement rule are right-aligned and arranged consecutively to ensure that within a single domain name supplement rule, there is no universal match bit to the right of the first exact match bit, and all universal match bits are on the left.

[0078] Step 202: Based on the BSS (Bit Separability Set) value of each rule bit, the rule bit with the maximum BSS value is used as a reference rule bit.

[0079] For example, after obtaining multiple matching rules, since each matching rule includes K rule bits, the BSS value of each rule bit can be calculated, that is, K rule bits correspond to K BSS values. In this way, the rule bit with the maximum BSS value can be used as the reference rule bit.

[0080] For ease of description, let's take three matching rules as an example: matching rule 1 is 0010 0110, matching rule 2 is 00101110, and matching rule 3 is 0011*110. Then, we can calculate the eight BSS values ​​corresponding to the eight rule bits. The BSS value is the product of the number of matching rules with 0 bits and the number of matching rules with 1 bits. Common matching bits* are not included in the BSS calculation. The BSS value indicates the degree of differentiation between the matching rule set (i.e., all matching rules). A larger BSS value indicates greater differentiation between the matching rule sets.

[0081] For example, 8 BSS values ​​corresponding to 8 regular bits can be shown in Table 1.

[0082] Table 1

[0083] bit0 bit1 bit2 bit3 bit4 bit5 bit6 bit7 0 0 0 2 1 0 0 0

[0084] Obviously, for the first rule bit (bit0), there are three zeros. Therefore, the number of matching rules for the 0-bit is 3, and the number of matching rules for the 1-bit is 0. The product of the two represents the BSS value, that is, the BSS value is 0. For the fourth rule bit (bit3), there are two zeros and one one. Therefore, the number of matching rules for the 0-bit is 2, and the number of matching rules for the 1-bit is 1. The product of the two represents the BSS value, that is, the BSS value is 2. For the fifth rule bit (bit4), there is one zero and one one. The universal matching bit* is not included in the calculation. Therefore, the number of matching rules for the 0-bit is 1, and the number of matching rules for the 1-bit is 1. The product of the two represents the BSS value, that is, the BSS value is 1. And so on.

[0085] For example, after obtaining 8 BSS values ​​corresponding to 8 regular bits, the regular bit with the maximum BSS value (ie, the fourth regular bit) can be used as a reference regular bit.

[0086] Step 203: Divide all matching rules into leaf nodes based on the reference rule bits to obtain a current decision tree, that is, the current decision tree is obtained after all matching rules are divided into multiple leaf nodes.

[0087] In an example, a matching rule with a reference rule bit of 0 may be divided into one leaf node, and a matching rule with a reference rule bit of 1 may be divided into another leaf node.

[0088] For example, a decision tree can be constructed. The current decision tree includes a root node, which includes all matching rules. For example, matching rule 1 is 0010 0110, matching rule 2 is 0010 1110, and matching rule 3 is 0011*110. After the fourth rule bit (bit3) is used as the reference rule bit, see Figure 4A Figure 2 shows how matching rules are grouped into leaf nodes. Matching rules whose fourth rule bit (bit 3) is 0 (matching rules 1 and 2) are grouped into the left leaf node, while matching rules whose fourth rule bit (bit 3) is 1 (matching rule 3) are grouped into the right leaf node. After performing this grouping on all matching rules, the resulting tree is called the current decision tree.

[0089] Step 204: Determine whether the current decision tree meets the end division condition.

[0090] If not, step 205 may be executed, and if so, step 207 may be executed.

[0091] In one example, the conditions for ending the partitioning may include but are not limited to at least one of the following:

[0092] Condition 1: The target spatial factor of the current decision tree is not less than the spatial factor threshold.

[0093] In one example, the target space factor may be determined based on the total number of reference rule bits, the number of matching rules in a leaf node of the current decision tree, and the total number of all matching rules.

[0094] For example, the target space factor can be determined using the following formula. Of course, this formula is only an example.

[0095]

[0096] In the above formula, SPFAC represents the target space factor. The larger the value of the target space factor, the faster the decision tree matching performance is, but the larger the memory usage is. m represents the total number of reference rule bits. Figure 4A In the example, the total number of reference rule bits is 1. When step 202 and step 203 are performed again, the total number of reference rule bits is increased by 1. Similarly, the total number of reference rule bits keeps changing.

[0097] 2 m Indicates the total number of leaf nodes in the current decision tree, N i Indicates the number of matching rules in the i-th leaf node. Figure 4A In the case where i is 0, N i Indicates the number of matching rules in the first leaf node, that is, 2. When i is 1, N i Indicates the number of matching rules in the second leaf node, which is 2.

[0098] N represents the total number of all matching rules. Figure 4A The total number of matching rules is 3.

[0099] In one example, the spatial factor threshold can be configured based on experience, and there is no restriction on the value of the spatial factor threshold. For example, the spatial factor threshold can be the configured maximum value.

[0100] For example, the target space factor is an important parameter of the decision tree. By adjusting the target space factor, we can achieve the following effect: for the leaf node where the matching rule is finally placed, no new reference rule bits can be found between any two matching rules, and only matching rules of the "big package small" and "cross inclusion" relationships exist. The matching rules of these two relationships cannot find new reference rule bits.

[0101] Here's an example to illustrate the "big includes small" and "cross-inclusion" relationships. Suppose domain 1 is abcde******, where * represents a wildcard character and there are 251 * characters following it. Domain 2 is ab******, where * represents a wildcard character and there are 254 * characters following it. Clearly, domain 2 is "big" and domain 1 is "small." Therefore, domain 1 is included in the "big includes small" relationship.

[0102] Assume that domain name 3 is ******ab, where * represents a wildcard character and there are 254 of them. Domain name 4 is uvwxyz******, where * represents a wildcard character and there are 250 of them. Thus, the first six characters of domain name 3 contain the uvwxyz of domain name 4, and the last two characters of domain name 4 contain the ab of domain name 3. These are in a "cross-inclusion" relationship. In this embodiment, by aligning all exact matching bits to the left (or right) and arranging them consecutively, there is no "cross-inclusion" relationship, and only the "large-includes-small" relationship needs to be considered.

[0103] In order to avoid leaf nodes with a "big-pack-small" relationship as much as possible, the space factor threshold can be set large enough, that is, the space factor threshold can be the configured maximum value. For example, the space factor threshold can be the maximum value that the configured target space factor can reach. If the target space factor is a signed 4-byte integer, the space factor threshold can be 0xFFFFFFFF, that is, 0xFFFFFFFF is the maximum value that the target space factor can reach. Alternatively, the space factor threshold can be the configured maximum floating-point number, the maximum floating-point number specified by IEEE 754 (3.40282347e+38), the maximum floating-point number specified in the C language, the maximum floating-point number specified in the Java language, etc. Of course, the above are just examples of space factor thresholds.

[0104] Condition 2: The number of leaf node rule references in the current decision tree is not greater than the leaf node rule number threshold.

[0105] In an example, based on the number of matching rules in each leaf node, the maximum number of matching rules can be used as the reference number of leaf node rules. Figure 4A In the example, there are 2 leaf nodes. The number of matching rules in the first leaf node is 2, and the number of matching rules in the second leaf node is 2. Therefore, the maximum number of matching rules is 2. In this way, the number of leaf node rule references can be 2.

[0106] In one example, the leaf node rule number threshold can be configured based on experience, and there is no restriction on the leaf node rule number threshold. For example, the leaf node rule number threshold can be a configured specified value.

[0107] For example, the leaf node rule count threshold can be denoted as BINTH. The smaller the leaf node rule count threshold, the shorter the decision tree height, the fewer matching rules for the leaf node, and the faster matching performance, but at the expense of increased memory usage. When the number of rules for the leaf node with the largest number of matching rules (i.e., the number of reference leaf node rules) is less than or equal to the leaf node rule count threshold, the reference rule bit for filtering the cutting rules is terminated.

[0108] For example, the number of reference rules for leaf nodes is a crucial parameter in decision trees. By adjusting the number of reference rules for leaf nodes, the goal is to ensure that, within the last leaf node where matching rules are placed, no new reference rule bits are found between any two matching rules. Only matching rules with the "large-include-small" and "cross-inclusion" relationships exist, and no new reference rule bits are found for matching rules with these two relationships. To minimize the presence of leaf nodes with "large-include-small" relationships, assuming there are no "cross-inclusion" relationships, the leaf node rule count threshold can be set to 1, that is, BINTH is set to 1.

[0109] Condition 3: The matching rules within each leaf node of the current decision tree cannot be further divided into two leaf nodes. For example, for each leaf node, if no new reference rule bits can be found between any two matching rules within the leaf node, and only matching rules with the "large-include-small" and "cross-inclusion" relationships exist, then the matching rules cannot be further divided into two leaf nodes.

[0110] Of course, the above three conditions are just examples, and more end-partitioning conditions can be configured. For example, the end-partitioning conditions can also include: the number of reference rule bits reaches a quantity threshold, such as 6, 8, 10, etc., that is, when the number of reference rule bits reaches 6, the current decision tree also meets the end-partitioning conditions.

[0111] In one example, when judging whether the end-partitioning conditions are met, if the current decision tree meets any one of the end-partitioning conditions, it is determined that the current decision tree has met the end-partitioning conditions; if the current decision tree does not meet each of the end-partitioning conditions, it is determined that the current decision tree does not meet the end-partitioning conditions.

[0112] Step 205: Based on the BSS values ​​of the remaining rule bits except the reference rule bit, the rule bit with the maximum BSS value among the BSS values ​​of the remaining rule bits is used as the reference rule bit again, that is, one reference rule bit is added.

[0113] In an example, the leaf node with the most matching rules can be selected from all leaf nodes. The leaf node has multiple matching rules. Based on this, the BSS value of each rule bit can be calculated. On the basis of excluding the reference rule bit, the rule bit with the maximum BSS value among the BSS values ​​of the remaining rule bits is used again as the reference rule bit.

[0114] See also Figure 4A As shown in Figure 2, the number of matching rules for the left leaf node is greater than that for the right leaf node. Therefore, the left leaf node is selected. For this leaf node, excluding the reference rule bit (bit 3), the seven BSS values ​​corresponding to the seven rule bits can be seen in Table 2.

[0115] Table 2

[0116] bit0 bit1 bit2 bit3 bit4 bit5 bit6 bit7 0 0 0 * 1 0 0 0

[0117] As can be seen from Table 2, excluding the reference rule bit (bit3), the first rule bit (bit0) contains two zeros. Therefore, the number of matching rules for bit 0 is 2, and the number of matching rules for bit 1 is 0. The product of the two represents the BSS value, which is 0. The fifth rule bit (bit4) contains one zero and one one. Therefore, the number of matching rules for bit 0 is 1, and the number of matching rules for bit 1 is 1. The product of the two represents the BSS value, which is 1. And so on.

[0118] It can be seen from Table 2 that after obtaining the 7 BSS values ​​corresponding to the 7 regular bits, the regular bit with the maximum BSS value (ie, the 5th regular bit) can be used as the reference regular bit again.

[0119] In another example, based on all matching rules of the root node, the BSS value of each rule bit can be calculated. After excluding the reference rule bit, the rule bit with the maximum BSS value among the BSS values ​​of the remaining rule bits is used as the reference rule bit again.

[0120] Step 206: Divide all matching rules into multiple leaf nodes based on all reference rule bits to obtain the current decision tree. For example, divide all matching rules into 2n leaf nodes, where n represents the number of reference rule bits.

[0121] For example, taking two reference rule bits as an example, the matching rule with the reference rule bit being 00 (i.e., bit3 is 0 and bit4 is 0, bit3 represents the first reference rule bit, and bit4 represents the second reference rule bit) can be divided into the first leaf node, the matching rule with the reference rule bit being 01 (i.e., bit3 is 0 and bit4 is 1) can be divided into the second leaf node, the matching rule with the reference rule bit being 10 (i.e., bit3 is 1 and bit4 is 0) can be divided into the third leaf node, and the matching rule with the reference rule bit being 11 (i.e., bit3 is 1 and bit4 is 1) can be divided into the fourth leaf node, thereby obtaining the current decision tree, that is, the current decision tree includes a total of 4 leaf nodes.

[0122] For example, taking three reference rule bits as an example, the matching rule with the reference rule bit of 000 is divided into the first leaf node, the matching rule with the reference rule bit of 001 is divided into the second leaf node, the matching rule with the reference rule bit of 010 is divided into the third leaf node, the matching rule with the reference rule bit of 011 is divided into the fourth leaf node, the matching rule with the reference rule bit of 100 is divided into the fifth leaf node, the matching rule with the reference rule bit of 101 is divided into the sixth leaf node, the matching rule with the reference rule bit of 110 is divided into the seventh leaf node, and the matching rule with the reference rule bit of 111 is divided into the eighth leaf node, thereby obtaining the current decision tree, that is, the current decision tree includes a total of 8 leaf nodes.

[0123] In summary, the current decision tree can be obtained. The current decision tree includes the root node, which includes all matching rules. After taking the 4th rule bit (bit3) and the 5th rule bit (bit4) as the reference rule bits, see Figure 4B The figure shows a schematic diagram of dividing the matching rules into leaf nodes.

[0124] Matching rule 1 with bit 3 set to 0 and bit 4 set to 0 can be assigned to the first leaf node, while matching rule 2 with bit 3 set to 0 and bit 4 set to 1 can be assigned to the second leaf node. Considering that * in matching rule 3 can be either 0 or 1, if * in matching rule 3 is 0, matching rule 3 with bit 3 set to 1 and bit 4 set to 0 can be assigned to the third leaf node. If * in matching rule 3 is 1, matching rule 3 with bit 3 set to 1 and bit 4 set to 1 can be assigned to the fourth leaf node.

[0125] In one example, after step 206, the operation of determining whether the current decision tree meets the end-partitioning condition can be repeatedly performed until the current decision tree meets the end-partitioning condition and stops. For example, after step 206, step 204 can be performed to determine whether the current decision tree meets the end-partitioning condition. If not, steps 205 and 206 are re-executed, i.e., a reference rule bit is added, and all matching rules are divided into more leaf nodes (2 to the power of n leaf nodes, where n represents the reference rule bit), and so on, until the current decision tree meets the end-partitioning condition and stops.

[0126] Step 207: Determine the current decision tree as the established decision tree (ie, the final decision tree), and generate a bit mask for the decision tree. The bit mask is used to indicate all reference rule bits.

[0127] In one example, the final decision tree can be a two-level decision tree, with the first level being the root node and the second level being 2n-power leaf nodes, where n represents the reference rule bit. On this basis, the current decision tree (see Figure 4B As shown in FIG, the final decision tree is established. In this way, a bit mask corresponding to the final decision tree can also be generated. For example, the bit mask can be 00011000, where the fourth "1" indicates that the fourth rule bit (bit3) is used as the reference rule bit, the fifth "1" indicates that the fifth rule bit (bit4) is used as the reference rule bit, and the remaining 0s indicate that the bits are not used as reference rule bits.

[0128] In an example, if the current decision tree does not have a leaf node of the specified type, the current decision tree is used as the final decision tree that has been established. The final decision tree is a two-level decision tree, with the first level being the root node and the second level being 2 to the power of n leaf nodes. Alternatively, if the current decision tree has a leaf node of the specified type, the leaf node of the specified type can be divided into multiple leaf nodes, that is, the final decision tree is a multi-level decision tree, with the first level being the root node, the second level including multiple leaf nodes, the third level including multiple leaf nodes, and so on. Figure 4CAs shown, it is a schematic diagram of a four-level decision tree, where the first level is the root node, the second level includes multiple leaf nodes, the third level includes multiple leaf nodes, and the fourth level includes multiple leaf nodes.

[0129] Regarding the method of dividing the root node into second-level leaf nodes, refer to steps 201-207. A bit mask of the second-level leaf node can be generated. The bit mask is used to indicate the reference rule bits from the root node to the second-level leaf node. The second-level leaf node corresponding to the data can be queried based on the bit mask.

[0130] If there is a specified type of leaf node in the second-level leaf node, that is, the specified type of leaf node includes multiple matching rules and the multiple matching rules can be divided, then based on the multiple matching rules of the specified type of leaf node, refer to steps 201-207, the multiple matching rules of the specified type of leaf node can be divided into third-level leaf nodes. This process will not be repeated, and a bit mask of the third-level leaf node is generated. The bit mask is used to indicate the reference rule bit from the second-level leaf node to the third-level leaf node. The third-level leaf node corresponding to the data can be queried based on the bit mask.

[0131] If there is a specified type leaf node in the third-level leaf node, that is, the specified type leaf node includes multiple matching rules and the multiple matching rules can be divided, then based on the multiple matching rules of the specified type leaf node, refer to steps 201-207, the multiple matching rules of the specified type leaf node can be divided into fourth-level leaf nodes, and a bit mask of the fourth-level leaf node is generated.

[0132] In one example, for each leaf node of an established decision tree (i.e., a final decision tree), each matching rule within the leaf node may be configured with a target mask, i.e., a target mask is configured for each matching rule within the leaf node. For example, for each matching rule, the matching rule may include K rule bits (e.g., 256*8 rule bits), and the target mask is used to indicate M rule bits among the K rule bits as exact matching bits, such as the first M rule bits as exact matching bits.

[0133] For example, taking the matching rule of a fuzzy domain name as an example, the matching rule may include 256 bytes (i.e., 256*8 rule bits), but the exact matching part of the fuzzy domain name in the actual production environment is relatively short. For example, the exact matching part is the name of an organization, manufacturer, trademark, etc., and there are a lot of wildcard characters. In order to avoid meaningless wildcard comparisons, in this embodiment, by configuring a target mask for the matching rule, the target mask is used to indicate the exact matching bits in K rule bits, thereby solving the performance problem of exact matching.

[0134] In one example, the target mask may include a first-level mask and a second-level mask. The K rule bits of the matching rule are divided into P1 data blocks, where P1 is a positive integer and can be configured according to actual needs, such as 16, 32, 64, etc. Based on this, the first-level mask may include P1 bits, and P1 bits correspond one-to-one to P1 data blocks. For example, taking the example that P1 data blocks can be 32 data blocks, and the size of each data block is 8 bytes (64 bits), that is, each data block can include 64 rule bits, and 32 data blocks correspond to 2048 rule bits (i.e., K rule bits).

[0135] For each data block, if there is an exact matching bit in the data block, the value of the bit corresponding to the data block in the first-level mask is a first value (e.g., 1), and the first value indicates that there is an exact matching bit in the data block. If there is no exact matching bit in the data block, the value of the bit corresponding to the data block in the first-level mask is a second value (0), and the second value indicates that there is no exact matching bit in the data block.

[0136] For each data block, if there are no exact matching bits in the data block, then the data block does not correspond to the secondary mask; if there are exact matching bits in the data block, then the data block corresponds to the secondary mask. If the data block includes P2 regular bits, then the secondary mask corresponding to the data block includes P2 bits, and the P2 bits correspond one-to-one with the P2 regular bits. For example, if the P2 regular bits are 64 regular bits, that is, the data block size is 8 bytes, then the secondary mask includes 64 bits.

[0137] For each rule bit, if the rule bit is an exact match bit, the value of the bit corresponding to the rule bit in the secondary mask is a third value (e.g., 1), and the third value represents an exact match bit. If the rule bit is a universal match bit, the value of the bit corresponding to the rule bit in the secondary mask is a fourth value (e.g., 0), and the fourth value represents a universal match bit.

[0138] For example, suppose the matching rule is fuzzy domain name abcabc****…*** (a total of 250 *), see Figure 5AThe following figure shows a target mask for the fuzzy domain name abcabc****…***. The fuzzy domain name abcabc****…*** can be divided into 32 data blocks. The first data block consists of the first 8 bytes, i.e., "abcabc**," which means it contains 6 exact matching characters and 2 fuzzy matching characters. The second data block consists of the second 8 bytes, i.e., "********," which means it contains 8 fuzzy matching characters. The third data block consists of the third 8 bytes, i.e., "********," and so on. The following 31 data blocks all contain 8-byte fuzzy matching characters.

[0139] The first-level mask can be 4 bytes, that is, 32 bits, and the 32 bits correspond to 32 data blocks one by one. For the first data block, since the data block has an exact matching bit (such as the 48 bits corresponding to abcabc), the value of the first bit in the first-level mask is the first value (such as 1). For the second data block, since the data block does not have an exact matching bit, the value of the second bit in the first-level mask is the second value (such as 0). For the third data block, since the data block does not have an exact matching bit, the value of the third bit in the first-level mask is the second value (such as 0), and so on. In this way, the first-level mask can be 100000000000000000000000000000000000.

[0140] In addition, the first data block corresponds to the secondary mask, while the remaining data blocks do not. For the secondary mask corresponding to the first data block, the secondary mask can be 8 bytes, i.e., 64 bits. The 8 bytes (64 bits) of the secondary mask correspond one-to-one with the 8 bytes (64 regular bits) of the data block.

[0141] The first byte (8 rule bits) of the data block is "a", indicating that the 8 rule bits are exact match bits, and the value of the first byte (8 bits) of the secondary mask is the third value (such as 1). The second byte (8 rule bits) of the data block is "b", indicating that the 8 rule bits are exact match bits, and the value of the second byte (8 bits) of the secondary mask is the third value (such as 1). Similarly, the sixth byte (8 rule bits) of the data block is "c", indicating that the 8 rule bits are exact match bits, and the value of the sixth byte (8 bits) of the secondary mask is the third value (such as 1). The seventh byte (8 rule bits) of the data block is "*", indicating that the 8 rule bits are universal match bits, and the value of the seventh byte (8 bits) of the secondary mask is the fourth value (such as 0). The 8th byte (8 rule bits) of the data block is "*", indicating that the 8 rule bits are universal matching bits, and the value of the 8th byte (8 bits) of the secondary mask is the fourth value (such as 0).

[0142] On this basis, when performing exact matching, the data block with exact matching characters is first found through the first-level mask, and then the fuzzy domain name and the matching domain name are matched through the second-level mask.

[0143] For example, the above secondary mask is not only suitable for fuzzy matching at the byte level, but also for fuzzy matching at the bit level. For example, the sixth byte "c" above is assumed to correspond to 8 bits 0110 0101. The first 6 bits are exact match bits, and the last 2 01 bits are set as wildcard bits. In this way, the secondary mask can be transformed into Figure 5B shown.

[0144] In an example, for a leaf node of an established decision tree (i.e., the final decision tree), if the leaf node contains a first matching rule and a second matching rule, and the rule scope of the first matching rule is larger than the rule scope of the second matching rule, then within the leaf node, the position of the second matching rule within the leaf node can be located in front of the position of the first matching rule within the leaf node.

[0145] For example, the rule scope of the first matching rule is greater than the rule scope of the second matching rule, which means that the first matching rule encompasses the second matching rule. For example, if the target matching data successfully matches the first matching rule, the target matching data may not successfully match the second matching rule; if the target matching data successfully matches the second matching rule, the target matching data must successfully match the first matching rule.

[0146] Referring to the above embodiment, the first matching rule envelops the second matching rule, meaning that the first matching rule and the second matching rule have a "big envelops small" relationship. For example, assuming domain name 1 is abcde****** and domain name 2 is ab******, domain name 2 is the "big" rule and domain name 1 is the "small" rule. In this case, domain name 2 is the first matching rule and domain name 1 is the second matching rule. Obviously, if the target matching data matches domain name 2, the target matching data does not necessarily match domain name 1; conversely, if the target matching data matches domain name 1, the target matching data must match domain name 2.

[0147] In view of the data fuzzy matching process based on the decision tree, a data fuzzy matching method is proposed in the embodiment of the present application, see Figure 6 FIG. 5 is a flow chart of the method, which may include:

[0148] Step 601: Obtain target matching data, where the target matching data includes K data bits.

[0149] For example, the target matching data may include K bits, and the bits in the target matching data may be called data bits. Therefore, the target matching data may include K data bits.

[0150] In one example, for step 601, the target matching data may be obtained by:

[0151] Step 6011: Obtain input original matching data, which may include multiple data bits, and the number of data bits in the original matching data may be less than or equal to K. For example, the original matching data may be obtained from a DNS request message, i.e., the original matching data input by the DNS server.

[0152] Step 6012: If the number of data bits in the original matching data is less than K, special character bits (special character bits are bits of characters that do not appear in the matching rules) may be supplemented to the original matching data to obtain supplemented matching data, and the number of data bits in the supplemented matching data may be equal to K. Alternatively, if the number of data bits in the original matching data is equal to K, the original matching data remains unchanged, i.e., the original matching data is used as the supplemented matching data.

[0153] For example, taking the original matching data as a domain name, since the maximum length of a domain name is 253 bytes, the value of K that is closest to the maximum length of 253 bytes of a domain name is 256 bytes. Therefore, the original matching data can be supplemented with special character bits to obtain 256 bytes of supplemented matching data.

[0154] Step 6013: If the supplemented matching data includes exact matching bits and special character bits, all exact matching bits are left-aligned and arranged continuously to obtain target matching data.

[0155] Alternatively, if the supplemented matching data includes exact matching bits and special character bits, all exact matching bits are right-aligned and arranged consecutively to obtain target matching data.

[0156] Step 602: Generate fuzzy matching data based on target matching data. The target matching data may include multiple data bits, and the fuzzy matching data may include multiple data bits and padding bits.

[0157] In one example, the number of padding bits S can be determined based on the number of exact matching bits and the minimum valid length in the target matching data, and this number of padding bits (i.e., S padding bits) can be added after the target matching data to obtain fuzzy matching data. On this basis, the target matching data can include K data bits, and the fuzzy matching data can include K data bits and S padding bits. For the S padding bits, the padding bits are special character bits that will not appear in the matching rules, that is, special character bits that will not appear in the matching rules are used as padding bits. In this embodiment, the content of these special character bits is not restricted.

[0158] The number of padding bits, S, can be determined based on the number of exact match bits and the minimum valid length. For example, if the target matching data may include exact match bits and special character bits, the number of exact match bits is the number of exact match bits in the target matching data. For example, if the target matching data is the ambiguous domain name abcabc****…*** (a total of 250 * characters), the number of exact match bits is 6 bytes (48 bits).

[0159] The minimum effective length is the minimum effective length of all matching rules corresponding to the decision tree, see Figure 3 As shown, there are matching rules (domain name 1 and domain name N) with a valid length of 6 bytes (48 bits), and there is a matching rule (domain name 2) with a valid length of 2 bytes (16 bits). Therefore, the minimum value of the valid length is 2 bytes (16 bits), and the minimum valid length can be 16 bits.

[0160] For example, the number S of padding bits may be determined using the following formula: S=DR, where D represents the number of exact matching bits in the target matching data, and R represents the minimum valid length.

[0161] The target matching data can be stored continuously in a memory of length (M+DR), with the first byte placed at the leftmost low address (array subscript 0), and the bytes that do not need to be matched are padded with special characters (characters that do not appear in the matching rules). In the memory (M+DR), M represents the maximum valid length, that is, the maximum value of the valid length supported by the decision tree. The maximum valid length can be 256 bytes, that is, the maximum valid length is the length K of the target matching data in the above embodiment, and (DR) represents the number of padding bits S. Obviously, (M+DR) means adding (DR) padding bits after the target matching data. Therefore, (M+DR) can represent the length of the fuzzy matching data.

[0162] Step 603: The first data bit of the fuzzy matching data is used as the current data bit, and the K consecutive data bits starting from the current data bit are used as the current data to be matched.

[0163] In this example, assuming the number of exact matching bits in the target matching data is D and the minimum valid length is R, D is first compared with the minimum valid length R. If D is less than R, the target matching data is directly determined to not match the target matching rule in the target leaf node, i.e., the match fails. If D is not less than R, the target matching data needs to be matched against the target matching rule, and step 603 is executed.

[0164] Regarding the number D of exact matching bits of the target matching data, the target matching data may include exact matching bits and special character bits. For example, if the target matching data is the fuzzy domain name abcabc****…*** (a total of 250 *), the number D of exact matching bits is 6 bytes (48 bits).

[0165] In one example, if no special characters can be found or a bit-granular match is performed, it is also necessary to record the number of exact matching bits of the target matching data (recorded as the first exact matching length) and the number of exact matching bits of the target matching rule (recorded as the second exact matching length). When performing an exact match, if the first exact matching length is inconsistent with the second exact matching length, it is directly determined that the target matching data does not match the target matching rule, that is, the match fails. If the first exact matching length is consistent with the second exact matching length, the target matching data is matched with the target matching rule, and step 603 is executed.

[0166] For example, if the target matching data is the fuzzy domain name abcd000000, that is, the exact matching bits of the target matching data are 4 bytes (32 bits), that is, the first exact matching length is 4 bytes, and the target matching rule is the fuzzy domain name abcd00****, that is, the exact matching bits of the target matching rule are 6 bytes (48 bits), that is, the second exact matching length is 6 bytes, then, when matching the target matching data with the target matching rule, the 6-byte (48-bit) fuzzy domain name abcd00 is matched with the 6-byte (48-bit) fuzzy domain name abcd00. Obviously, the above matching result is a successful match, but in fact, the exact matching character of the target matching data "abcd000000" is "abcd", which fails to match the target matching rule "abcd00", that is, the above matching result is wrong.

[0167] Based on this, in this embodiment, the first exact match length (4 bytes) and the second exact match length (6 bytes) are compared. Obviously, since the first exact match length is inconsistent with the second exact match length, it is directly determined that the target matching data fails to match the target matching rule. The 6-byte fuzzy domain name abcd00 is no longer matched with the 6-byte fuzzy domain name abcd00, and no incorrect matching result is obtained.

[0168] Step 604: Query the target leaf node corresponding to the current data to be matched from all leaf nodes of the established decision tree. The target leaf node is used to record the target matching rule and the target mask.

[0169] In one example, the target matching rule may include K rule bits, and the target mask is used to indicate M of the K rule bits as exact match bits. For example, the target mask may include a primary mask and a secondary mask. The primary mask is used to indicate which data block of the target matching rule includes the exact match bits. The secondary mask is used to indicate which rule bits of the data block are exact match bits. For details about the primary mask and the secondary mask, please refer to the above embodiment.

[0170] After obtaining the current data to be matched, the bit mask (the bit mask is used to indicate all reference rule bits) can be queried to obtain the reference data bits in the current data to be matched. For example, if all reference rule bits include bit3 and bit4, then the reference data bits are bit3 and bit4. Therefore, the data bit of bit3 (i.e., the fourth data bit) and the data bit of bit4 (i.e., the fifth data bit) are selected from the current data to be matched. In this way, the target leaf node corresponding to the current data to be matched can be queried from all leaf nodes through the data bit of bit3 and the data bit of bit4.

[0171] See also Figure 4B As shown in the figure, if the data bit of bit3 is 0 and the data bit of bit4 is 0, the first leaf node is used as the target leaf node corresponding to the current data to be matched. If the data bit of bit3 is 0 and the data bit of bit4 is 1, the second leaf node is used as the target leaf node corresponding to the current data to be matched. If the data bit of bit3 is 1 and the data bit of bit4 is 0, the third leaf node is used as the target leaf node corresponding to the current data to be matched. If the data bit of bit3 is 1 and the data bit of bit4 is 1, the fourth leaf node is used as the target leaf node corresponding to the current data to be matched.

[0172] In summary, the target leaf node can be queried from all leaf nodes of the decision tree.

[0173] In an example, in a decision tree, each leaf node contains a pointer to a block of memory containing matching rules. Each non-leaf node contains the target mask of the rule bits and a pointer to an array of all leaf nodes. Based on this, the target leaf node corresponding to the current data to be matched can be queried from all leaf nodes of the decision tree. See Figure 7 The figure shows a schematic diagram of searching for leaf nodes.

[0174] The Packet Header represents the incoming message tuple information, expressed in binary. That is, the current data to be matched based on the message, which can be a domain name, etc. Here, the current data to be matched includes 10 data bits, and these 10 data bits are respectively recorded as B0, B1, ..., B9.

[0175] Bitmask indicates the reference data bits selected during the tree building process. Here, the second, sixth, seventh, and ninth bits are used as reference data bits.

[0176] The Index for the Next Node represents the calculated subscript of the target leaf node. That is, the index consisting of B1, B5, B6, and B8 corresponds to the target leaf node. For example, assuming B1 is 1, B5 is 0, B6 is 1, and B8 is 0, then the leaf node corresponding to index 0101 is the target leaf node. Obviously, the bit value corresponding to the 1-bit bitmask in the Packet Header can be taken out and added to the end of the all-0 data to obtain the Index for the Next Node, and then the target leaf node can be found.

[0177] Step 605: Select M data bits from the current data to be matched, select M rule bits from the K rule bits of the target matching rule, and determine whether the value of each bit in the M data bits is the same as the value of the corresponding rule bit in the M rule bits.

[0178] If so, that is, the value of each bit in the M data bits is the same as the value of the corresponding rule bit in the M rule bits, execute step 606; if not, that is, the value of each bit in the M data bits is different from the value of the corresponding rule bit in the M rule bits, execute step 607.

[0179] Step 606: Determine whether the target matching data successfully matches the target matching rule, and perform corresponding processing based on the target matching data. For example, the IP address corresponding to the target matching data is obtained from the DNS request message, and the IP address is stored in the security table entry. In this way, data packets accessing the IP address can pass through the security device, thereby implementing access control for the data packets.

[0180] Step 607: Determine whether the number of matches corresponding to the target matching data is the sum of the number of padding bits and 1. The number of padding bits can be determined based on the minimum valid length and the number of exact matching bits, such as S = DR. In step 607, it can be determined whether the number of matches corresponding to the target matching data has reached D - R + 1. If so, step 608 can be executed; if not, step 609 can be executed.

[0181] For example, the initial value of the number of matches corresponding to the target matching data is a preset value (such as 0). When the current data to be matched is selected from the fuzzy matching data for the first time (that is, the current data to be matched is matched with the target matching rule), the number of matches corresponding to the target matching data is updated to 1. Each subsequent time the current data to be matched is selected from the fuzzy matching data, the number of matches corresponding to the target matching data is increased by 1.

[0182] Step 608: Determine whether the target matching data fails to match the target matching rule, and perform processing corresponding to the match failure based on the target matching data. For example, when obtaining the IP address corresponding to the target matching data from the DNS request message, the IP address is prohibited from being stored in the security table. In this way, data packets accessing the IP address cannot pass through the security device.

[0183] Step 609: add 1 to the number of matches corresponding to the target matching data, and use the data bit after the current data bit as the current data bit, and repeat the operation of using K consecutive data bits starting from the current data bit as the current data to be matched, selecting M data bits from the current data to be matched, and selecting M rule bits from the K rule bits, until the value of each bit in the M data bits is the same as the value of the corresponding rule bit in the M rule bits, that is, re-execute steps 603-609.

[0184] The following describes the above steps in detail in combination with specific application scenarios.

[0185] Assuming the target matching data is the domain name abcdefgh.com****…*** (a total of 244 *), the number of exact matching bits D in the target matching data is 12 bytes (96 bits). Assuming the minimum valid length R is 2 bytes (16 bits), the number of padding bits S is 10 bytes (80 bits). Based on this, the target matching data can be stored continuously in a memory with a length of (M+DR), where M represents the maximum valid length, i.e., 256 bytes, and (M+DR) represents the length of the fuzzy matching data, i.e., the length of the fuzzy matching data is 266 bytes (2128 bits).

[0186] See also Figure 8 The figure below is a schematic diagram of fuzzy matching data. The fuzzy matching data includes 12 bytes of exact matching characters (abcdefgh.com) and 254 bytes of special characters. In this way, the target matching data "abcdefgh.com" can be placed in a 266-byte array, and the target matching data "abcdefgh.com" is placed on the left, and the remaining positions are filled with special characters that do not appear in the domain name. Figure 8 In the figure, byte granularity is used as an example. The fuzzy matching data can also be expressed in bit granularity, that is, each exact matching character in the fuzzy matching data can correspond to 8 exact matching bits.

[0187] For the first matching process, the first exact matching character "a" of the fuzzy matching data is used as the current exact matching character (such as the exact matching bit as the current data bit), and the continuous K data bits (such as 256 bytes) starting from the current exact matching character are used as the current data to be matched, that is, the current data to be matched is the character between two 1s. The first M data bits are selected from the current data to be matched, and the first M rule bits are selected from the target matching rule. If the M data bits are the same as the M rule bits, it is determined that the target matching data and the target matching rule are successfully matched, and the matching process ends. If the M data bits are different from the M rule bits, since the number of matches corresponding to the target matching data does not reach the number of padding bits, the second matching process is continued.

[0188] For the second matching process, the second exact matching character "b" of the fuzzy matching data is used as the current exact matching character, and the K consecutive data bits starting from the current exact matching character are used as the current data to be matched, that is, the current data to be matched is the character between two 2s. The first M data bits are selected from the current data to be matched, and the first M rule bits are selected from the target matching rule. If the M data bits are the same as the M rule bits, it is determined that the target matching data and the target matching rule are successfully matched, and the matching process ends. If the M data bits are different from the M rule bits, since the number of matches corresponding to the target matching data does not reach the number of padding bits, the third matching process continues.

[0189] For the third matching process, the current data to be matched is the character between two 2s, and so on, until the number of matches corresponding to the target matching data reaches the number of padding bits plus 1, it is determined that the target matching data fails to match the target matching rule, or when M data bits are the same as M rule bits, it is determined that the target matching data successfully matches the target matching rule, and the matching process ends.

[0190] In summary, the matching process between the target matching data and the target matching rule can be completed until it is determined whether the target matching data and the target matching rule are matched successfully or failed.

[0191] In the above process, the value of M can be determined based on the target mask. The target mask can include a first-level mask and a second-level mask. The first-level mask is used to indicate which data block of the target matching rule / target matching data includes the exact matching bit, and the second-level mask is used to indicate which rule bits / data bits of the data block are the exact matching bits. In this way, the value of M can be determined based on the first-level mask and the second-level mask, and then the first M data bits are selected from the current data to be matched, and the first M rule bits are selected from the target matching rule, and a comparison is made to see whether the M data bits are the same as the M rule bits.

[0192] As can be seen from the above, for the first matching process, the matching domain name pointed to by start pointer 0 and end pointer 255 is used for the query. For the second matching process, the matching domain name pointed to by start pointer 1 and end pointer 256 is used for the query. Similarly, for the final matching process, the matching domain name pointed to by start pointer 10 and end pointer 265 is used for the query. In this way, a maximum of 11 queries are performed. As long as there is a hit, the hit result is output and the query is terminated. If all possible matching fuzzy domain names need to be found, all 11 queries must be completed to obtain the final query result.

[0193] In one example, when matching target matching data with target matching rules, there are three possible matching results: hitting any one target matching rule, failing to hit any target matching rules, and finding all target matching rules that can be hit. In this embodiment, for the target leaf node corresponding to the target matching data, the target leaf node may have no target matching rule, one target matching rule, or multiple target matching rules in a "large-package-small" relationship. Based on this, the following matching processing methods can be used for the above three situations.

[0194] 1. Find all target matching rules that can be hit.

[0195] In the first case, the target matching data does not hit the target matching rule, the match fails, and the query ends.

[0196] In the second case, the first-level mask and the second-level mask are used to perform an exact match between the target matching data and the target matching rule. The target matching data may hit the target matching rule, the match is successful, and the query ends. Alternatively, the target matching data may not hit the target matching rule, the match fails, and the query ends.

[0197] In the third case, considering that hitting the "small" rule will definitely hit the "big" rule, hitting the "big" rule may not necessarily hit the "small" rule, therefore, the "small" rule is placed in front and the "big" rule is placed in the back. Based on this, if the target leaf node includes the first matching rule and the second matching rule, and the first matching rule envelops the second matching rule, then the first matching rule is the "big" rule and the second matching rule is the "small" rule, and the second matching rule is placed in front of the first matching rule. On this basis, after matching the first "small" rule and hitting it, you can stop matching, and the subsequent rules will definitely hit, and the hit "small" rule and all the rules following it can be output at the same time. For example, if the target matching data matches the first matching rule, if the target matching data hits the first matching rule, the match is successful, and the query ends. The first matching rule and the second matching rule can be output at the same time, indicating that multiple rules are hit at the same time.

[0198] 2. All target matching rules are unmatched.

[0199] In the first case, the target matching data does not hit the target matching rule, the match fails, and the query ends.

[0200] In the second case, the first-level mask and the second-level mask are used to perform an exact match between the target matching data and the target matching rule. The target matching data may hit the target matching rule, the match is successful, and the query ends. Alternatively, the target matching data may not hit the target matching rule, the match fails, and the query ends.

[0201] In the third case, considering that hitting a "small" rule will definitely hit a "large" rule, but hitting a "large" rule may not necessarily hit a "small" rule, the "large" rule is placed first and the "small" rule is placed last. If the first "large" rule is not matched, the matching can be stopped, and subsequent "small" rules will definitely not be matched. Based on this, if the target leaf node includes the first matching rule and the second matching rule, and the first matching rule envelops the second matching rule, the first matching rule is placed before the second matching rule. Based on this, if the target matching data matches the second matching rule, if the target matching data does not hit the second matching rule, the match fails, the query ends, and all matching rules fail.

[0202] 3. Just match the rules for hitting any target.

[0203] In the first case, the target matching data does not hit the target matching rule, the match fails, and the query ends.

[0204] In the second case, the first-level mask and the second-level mask are used to perform an exact match between the target matching data and the target matching rule. The target matching data may hit the target matching rule, the match is successful, and the query ends. Alternatively, the target matching data may not hit the target matching rule, the match fails, and the query ends.

[0205] In the third case, considering that hitting the "small" rule will definitely hit the "large" rule, but hitting the "large" rule may not necessarily hit the "small" rule, the "large" rule is placed first and the "small" rule is placed last. After matching the first "large" rule, you can stop matching and end the query. Based on this, if the target leaf node includes the first matching rule and the second matching rule, and the first matching rule envelops the second matching rule, then the first matching rule is placed before the second matching rule. If the target matching data matches the second matching rule, if the target matching data hits the second matching rule, the match is successful and the query ends.

[0206] As can be seen from the above technical solutions, in the embodiments of the present application, it is not necessary to match the current data to be matched with all matching rules, which reduces the number of matches, improves matching performance, greatly improves the performance of precise and fuzzy matching, reduces matching time, and the matching performance can meet actual needs. By indicating the precise matching bits by the target mask, the precise matching is optimized, and only the precise matching bits need to be matched, which further improves the matching performance and reduces the matching time. By splitting the target matching data and querying it multiple times, fuzzy matching of the precise matching field position is achieved, which significantly improves the performance of fuzzy matching of massive data with non-fixed precise matching field position, significantly improves the matching performance under massive rules, significantly improves the performance of precise and fuzzy matching, and makes the matching performance more stable.

[0207] It can significantly improve the fuzzy matching performance of massive data where the exact match field position is not fixed. Taking the fuzzy matching of Internet English domain names as the experimental object (this algorithm is not limited to the fuzzy matching of Internet English domain names, it is suitable for any scenario of fuzzy matching of massive data where the exact match field position is not fixed), the time taken for a single domain name to match any subdomain in the presence of massive fuzzy subdomains is measured.

[0208] The experimental conditions for this experiment are as follows: The maximum length of a domain name is 253 bytes. The domain name rule is expanded to 256 bytes, which is 2 to the power of 8, for ease of calculation. The supplementary part is a wildcard character. The characters used in the domain name are letters (az, AZ), numbers (0-9), hyphens (-), and separators (.). The format of the regular domain name used for matching is ***xyz***, where the leading and trailing * represent wildcard characters, and xyz represents the subdomain. The total number of wildcard characters and subdomains is 256, and the number of leading and trailing wildcard characters can be 0. Here, xyz is just an example. It is randomly generated in the experiment, with lengths ranging from 1 to 255. The number of regular domain names is: 8K, 16K, 32K, 100K, and 1K represents 1024. The data domain name used for matching is the domain name of a well-known website.

[0209] The experimental group used the above-mentioned technical solution of this embodiment, while the control group used the substring matching library function strcasestr to perform traversal search. Under the same device environment, the same number of regular domain names, and the same data domain names, multiple sets of data were measured and averaged. The comparison results can be seen in Table 3.

[0210] Table 3

[0211]

[0212] As shown in Table 3, the algorithm's single-matching time ranges from kilonanoseconds to 10,000 nanoseconds, a difference of only one order of magnitude. The control group's time ranges from 100,000 nanoseconds to 10 million nanoseconds, a difference of two orders of magnitude. Since this is the average of multiple experiments, in the original data before averaging, the algorithm's performance remains in the kilonanosecond to 10,000 nanosecond range, while the control group's performance ranges from a good kilonanosecond to a bad 10 million nanoseconds, a difference of up to four orders of magnitude. Therefore, the matching performance of the algorithm is more stable and less volatile than that of the control group.

[0213] Compared with strcasestr's traversal matching, this algorithm improves performance by 10 to 1000 times, regardless of whether it matches or not. The more rule domain names there are, the more obvious the improvement.

[0214] Based on the same application concept as the above method, a data fuzzy matching device is proposed in the embodiment of the present application. Figure 9A FIG. 1 is a schematic diagram of the structure of the device, which may include:

[0215] A generation module 911 is configured to obtain target matching data and obtain fuzzy matching data based on the target matching data; use the first data bit of the fuzzy matching data as the current data bit, and use K consecutive data bits starting from the current data bit as the current data to be matched;

[0216] An acquisition module 912 is configured to query a target leaf node corresponding to the current to-be-matched data from all leaf nodes of the established decision tree, wherein the target leaf node is configured to record a target matching rule and a target mask; wherein the target matching rule includes K rule bits, and the target mask is configured to indicate M rule bits among the K rule bits as exact matching bits;

[0217] The determination module 913 is used to select M data bits from the current data to be matched, and select M rule bits from the K rule bits; if the value of each bit in the M data bits is the same as the value of the corresponding rule bit in the M rule bits, then the corresponding processing of the match is performed based on the target matching data; if the value of each bit in the M data bits is different from the value of the corresponding rule bit in the M rule bits, then the data bit after the current data bit is used as the current data bit, and the generation module 911 uses the consecutive K data bits starting from the current data bit as the current data to be matched.

[0218] In one example, when the generation module 911 obtains the fuzzy matching data based on the target matching data, it is specifically used to: determine the number of padding bits based on the number of exact matching bits and the minimum valid length in the target matching data; wherein the minimum valid length is the minimum value of the valid lengths of all matching rules corresponding to the decision tree; add the number of padding bits after the target matching data to obtain the fuzzy matching data, and the padding bits are special character bits that will not appear in the target matching rules; the determination module 913 is also used to determine whether the number of matches corresponding to the target matching data is the sum of the number and 1 if the value of each bit in the M data bits is different from the value of the corresponding rule bit in the M rule bits; wherein the initial value of the number of matches corresponding to the target matching data is a preset value; if not, the number of matches corresponding to the target matching data is increased by 1, and the data bit after the current data bit is used as the current data bit; if so, the corresponding processing of the matching failure is performed based on the target matching data.

[0219] In one example, the device also includes: an establishment module for establishing the decision tree; when the establishment module establishes the decision tree, it is specifically used to: obtain multiple matching rules, each matching rule includes K rule bits; based on the bit separable set BSS value of each rule bit, the rule bit with the maximum BSS value is used as the reference rule bit; based on the reference rule bit, all matching rules are divided into leaf nodes to obtain the current decision tree; determine whether the current decision tree meets the end division condition; if not, based on the BSS values ​​of the remaining rule bits except the reference rule bit, the rule bit with the maximum BSS value in the BSS values ​​of the remaining rule bits is used again as the reference rule bit; based on all reference rule bits, all matching rules are divided into multiple leaf nodes to obtain the current decision tree, and repeatedly perform the operation of determining whether the current decision tree meets the end division condition until the current decision tree meets the end division condition and stops; if so, the current decision tree is determined as the established decision tree, and a bit mask is generated for the decision tree, and the bit mask is used to indicate all reference rule bits.

[0220] In one example, the end division condition includes: the target spatial factor of the current decision tree is not less than the spatial factor threshold; wherein the spatial factor threshold is the maximum value that the configured target spatial factor can reach; or, the spatial factor threshold is the configured maximum floating-point number; the number of leaf node rule references of the current decision tree is not greater than the leaf node rule number threshold; wherein the leaf node rule number threshold is a configured specified value; the matching rules within each leaf node of the current decision tree cannot be further divided into two leaf nodes; wherein, the establishment module is specifically used to determine whether the current decision tree meets the end division condition: if the current decision tree meets any one of the end division conditions, it is determined that the current decision tree has met the end division condition; if the current decision tree does not meet each of the end division conditions, it is determined that the current decision tree does not meet the end division condition.

[0221] In one example, when the establishment module obtains multiple matching rules, it is specifically used to: obtain multiple configured original rules; for each original rule, if the number of rule bits in the original rule is less than K, the wildcard bit is supplemented to the original rule to obtain a supplemented rule, and the number of rule bits in the supplemented rule is equal to K, K is 2 to the power of n, and n is a positive integer; if the supplemented rule includes exact matching bits and general matching bits, all exact matching bits are left-aligned and arranged continuously or right-aligned and arranged continuously to obtain the matching rule corresponding to the original rule; when the acquisition module 912 obtains target matching data, it is specifically used to: obtain the input original matching data; if the number of data bits in the original matching data is less than K, the special character bit is supplemented to the original matching data to obtain supplemented matching data, and the number of data bits in the supplemented matching data is equal to K; if the supplemented matching data includes exact matching bits and special character bits, all exact matching bits are left-aligned and arranged continuously or right-aligned and arranged continuously to obtain the target matching data.

[0222] In one example, the matching rule in the leaf node is configured with a target mask, and the target mask includes a primary mask and a secondary mask;

[0223] The K rule bits of the matching rule are divided into P1 data blocks, the first-level mask includes P1 bits, and the P1 bits correspond one-to-one to the P1 data blocks; for each data block, when there is an exact matching bit in the data block, the value of the bit corresponding to the data block in the first-level mask is a first value, and the first value indicates that there is an exact matching bit in the data block; when there is no exact matching bit in the data block, the value of the bit corresponding to the data block in the first-level mask is a second value, and the second value indicates that there is no exact matching bit in the data block;

[0224] Wherein, for each data block, when there is no exact matching bit in the data block, the data block does not correspond to the secondary mask; when there is an exact matching bit in the data block, the data block corresponds to the secondary mask;

[0225] Wherein, when the data block includes P2 regular bits, the secondary mask includes P2 bits, and the P2 bits correspond one-to-one to the P2 regular bits;

[0226] For each rule bit, when the rule bit is an exact match bit, the value of the bit corresponding to the rule bit in the secondary mask is the third value, and the third value represents the exact match bit; when the rule bit is a general match bit, the value of the bit corresponding to the rule bit in the secondary mask is the fourth value, and the fourth value represents the general match bit.

[0227] In an example, if the leaf node contains a first matching rule and a second matching rule, and the rule scope of the first matching rule is larger than the rule scope of the second matching rule, then the position of the second matching rule in the leaf node is located in front of the position of the first matching rule in the leaf node.

[0228] Based on the same application concept as the above method, an electronic device is proposed in the embodiment of the present application, see Figure 9B As shown, it includes: a processor 921 and a machine-readable storage medium 922, the machine-readable storage medium 922 stores machine-executable instructions that can be executed by the processor 921; the processor 921 is used to execute the machine-executable instructions to implement the data fuzzy matching method disclosed in the above example of this application.

[0229] Based on the same application concept as the above method, an embodiment of the present application also provides a machine-readable storage medium, on which a number of computer instructions are stored. When the computer instructions are executed by a processor, the data fuzzy matching method disclosed in the above example of the present application can be implemented.

[0230] The machine-readable storage medium may be any electronic, magnetic, optical, or other physical storage device that may contain or store information, such as executable instructions, data, and the like. For example, the machine-readable storage medium may be: RAM (Random Access Memory), volatile memory, non-volatile memory, flash memory, a storage drive (such as a hard disk drive), a solid-state drive, any type of storage disk (such as a CD, DVD, etc.), or similar storage media, or a combination thereof.

[0231] Based on the same application concept as the above method, an embodiment of the present application further provides a computer program product, which may include a computer program. When the computer program is executed by a processor, it implements the data fuzzy matching method disclosed in the above example of the present application.

[0232] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the embodiments of the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0233] The foregoing is merely an embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application should all be included within the scope of the claims of the present application.

Claims

1. A data fuzzy matching method, characterized in that: The method comprises: Acquire target matching data, and obtain fuzzy matching data according to the target matching data; The first data bit of the fuzzy matching data is used as the current data bit, and the continuous K data bits starting from the current data bit are used as the current data to be matched; Querying the target leaf node corresponding to the current to-be-matched data from all leaf nodes of the established decision tree, wherein the target leaf node is used to record the target matching rule and the target mask; wherein the target matching rule includes K rule bits, and the target mask is used to indicate M rule bits among the K rule bits as exact matching bits; Selecting M data bits from the current data to be matched, and selecting M rule bits from the K rule bits; If the value of each bit in the M data bits is the same as the value of the corresponding rule bit in the M rule bits, performing a process corresponding to a successful match based on the target matching data; If the value of each of the M data bits is different from the value of the corresponding rule bit in the M rule bits, the data bit after the current data bit is used as the current data bit, and the operation of taking the consecutive K data bits starting from the current data bit as the current data to be matched, selecting M data bits from the current data to be matched, and selecting M rule bits from the K rule bits is repeated until the value of each of the M data bits is the same as the value of the corresponding rule bit in the M rule bits.

2. The method according to claim 1, characterized in that The obtaining of fuzzy matching data according to the target matching data specifically includes: Determine the number of padding bits based on the number of exact matching bits in the target matching data and the minimum effective length; wherein the minimum effective length is the minimum value of the effective lengths of all matching rules corresponding to the decision tree; Adding the number of padding bits after the target matching data to obtain the fuzzy matching data, wherein the padding bits are special character bits that will not appear in the target matching rule; If the value of each bit in the M data bits is different from the value of the corresponding regular bit in the M regular bits, the method further includes: Determine whether the number of matches corresponding to the target matching data is the sum of the number and 1; wherein an initial value of the number of matches corresponding to the target matching data is a preset value; If not, the number of matches corresponding to the target matching data is increased by 1, and the data bit after the current data bit is used as the current data bit; If so, a process corresponding to the matching failure is performed based on the target matching data.

3. The method according to claim 1 or 2, characterized in that: Before querying the target leaf node corresponding to the current to-be-matched data from all leaf nodes of the established decision tree, the method further includes: Acquire multiple matching rules, each matching rule including K rule bits; Based on the bit separable set BSS value of each rule bit, the rule bit with the maximum BSS value is used as the reference rule bit; Based on the reference rule bits, all matching rules are divided into leaf nodes to obtain the current decision tree; Determine whether the current decision tree meets the end division conditions; If not, based on the BSS values ​​of the remaining rule bits except the reference rule bit, the rule bit with the maximum BSS value among the BSS values ​​of the remaining rule bits is used as the reference rule bit again; Based on all reference rule bits, all matching rules are divided into multiple leaf nodes to obtain the current decision tree, and the operation of judging whether the current decision tree meets the end division condition is repeatedly performed until the current decision tree meets the end division condition and stops; If so, the current decision tree is determined as the established decision tree, and a bit mask is generated for the decision tree, where the bit mask is used to indicate all reference rule bits.

4. The method according to claim 3, characterized in that The end division conditions include: The target spatial factor of the current decision tree is not less than a spatial factor threshold; wherein the spatial factor threshold is the maximum value that the configured target spatial factor can reach; or the spatial factor threshold is the configured maximum floating point number; The reference number of leaf node rules of the current decision tree is not greater than a threshold number of leaf node rules; wherein the threshold number of leaf node rules is a configured specified value; The matching rules in each leaf node of the current decision tree cannot be further divided into two leaf nodes; The step of judging whether the current decision tree meets the end-partitioning condition specifically includes: If the current decision tree satisfies any one of the end-partitioning conditions, it is determined that the current decision tree has satisfied the end-partitioning condition; If the current decision tree does not satisfy each of the end-segmentation conditions, it is determined that the current decision tree does not satisfy the end-segmentation condition.

5. The method according to claim 3, characterized in that: The obtaining of multiple matching rules specifically includes: Get multiple configured original rules; For each original rule, if the number of rule bits in the original rule is less than K, then the original rule is supplemented with wildcard bits to obtain a supplemented rule, where the number of rule bits in the supplemented rule is equal to K, where K is 2 to the power of n, and n is a positive integer; If the supplemented rule includes exact matching bits and general matching bits, all exact matching bits are arranged in a left-aligned or right-aligned sequence to obtain a matching rule corresponding to the original rule; The obtaining of target matching data specifically includes: Get the input raw matching data; If the number of data bits in the original matching data is less than K, the special character bits are supplemented to the original matching data to obtain supplemented matching data, where the number of data bits in the supplemented matching data is equal to K; If the supplemented matching data includes exact matching bits and special character bits, all exact matching bits are left-aligned and continuously arranged or right-aligned and continuously arranged to obtain the target matching data.

6. The method according to claim 3, characterized in that The matching rule in the leaf node is configured with a target mask, and the target mask includes a primary mask and a secondary mask; Among them, the K rule bits of the matching rule are divided into P1 data blocks, the first-level mask includes P1 bits, and the P1 bits correspond to the P1 data blocks one by one; for each data block, when there is an exact matching bit in the data block, the value of the bit corresponding to the data block in the first-level mask is a first value, and the first value indicates that there is an exact matching bit in the data block; when there is no exact matching bit in the data block, the value of the bit corresponding to the data block in the first-level mask is a second value, and the second value indicates that there is no exact matching bit in the data block; Wherein, for each data block, when there is no exact matching bit in the data block, the data block does not correspond to the secondary mask; when there is an exact matching bit in the data block, the data block corresponds to the secondary mask; Wherein, when the data block includes P2 regular bits, the secondary mask includes P2 bits, and the P2 bits correspond one-to-one to the P2 regular bits; For each rule bit, when the rule bit is an exact match bit, the value of the bit corresponding to the rule bit in the secondary mask is a third value, and the third value represents an exact match bit; when the rule bit is a general match bit, the value of the bit corresponding to the rule bit in the secondary mask is a fourth value, and the fourth value represents a general match bit.

7. The method according to claim 3, characterized in that If the first matching rule and the second matching rule exist in the leaf node, and the rule scope of the first matching rule is larger than the rule scope of the second matching rule, the position of the second matching rule in the leaf node is located in front of the position of the first matching rule in the leaf node.

8. A data fuzzy matching device, characterized in that: The device comprises: A generation module is used to obtain target matching data, and obtain fuzzy matching data according to the target matching data; the first data bit of the fuzzy matching data is used as the current data bit, and the continuous K data bits starting from the current data bit are used as the current data to be matched; An acquisition module, used to query the target leaf node corresponding to the current to-be-matched data from all leaf nodes of the established decision tree, wherein the target leaf node is used to record the target matching rule and the target mask; wherein the target matching rule includes K rule bits, and the target mask is used to indicate M rule bits among the K rule bits as exact matching bits; A determination module is used to select M data bits from the current data to be matched, and select M rule bits from the K rule bits; if the value of each bit in the M data bits is the same as the value of the corresponding rule bit in the M rule bits, then the corresponding processing of the match is performed based on the target matching data; if the value of each bit in the M data bits is different from the value of the corresponding rule bit in the M rule bits, then the next data bit after the current data bit is used as the current data bit, and the generation module uses the continuous K data bits starting from the current data bit as the current data to be matched.

9. The device according to claim 8, characterized in that The generating module is specifically used to obtain the fuzzy matching data according to the target matching data: determine the number of padding bits based on the number of exact matching bits and the minimum effective length in the target matching data; wherein the minimum effective length is the minimum value of the effective lengths of all matching rules corresponding to the decision tree; add the number of padding bits after the target matching data to obtain the fuzzy matching data, wherein the padding bits are special character bits that will not appear in the target matching rules; The determination module is also used to determine whether the number of matches corresponding to the target matching data is the sum of the number and 1 if the value of each bit in the M data bits is different from the value of the corresponding rule bit in the M rule bits; wherein the initial value of the number of matches corresponding to the target matching data is a preset value; if not, the number of matches corresponding to the target matching data is increased by 1, and the next data bit after the current data bit is used as the current data bit; if so, the corresponding processing of the matching failure is performed based on the target matching data.

10. The device according to claim 8 or 9, characterized in that The device further comprises: a building module, which is used to build the decision tree; when the building module builds the decision tree, it is specifically used to: Acquire multiple matching rules, each matching rule includes K rule bits; based on the bit separable set BSS value of each rule bit, use the rule bit with the maximum BSS value as the reference rule bit; divide all matching rules into leaf nodes based on the reference rule bits to obtain the current decision tree; Determine whether the current decision tree meets the end division condition; if not, based on the BSS values ​​of the remaining rule bits except the reference rule bit, use the rule bit with the maximum BSS value among the BSS values ​​of the remaining rule bits as the reference rule bit again; divide all matching rules into multiple leaf nodes based on all reference rule bits to obtain the current decision tree, and repeatedly perform the operation of determining whether the current decision tree meets the end division condition until the current decision tree meets the end division condition and stops; If so, the current decision tree is determined as the established decision tree, and a bit mask is generated for the decision tree, where the bit mask is used to indicate all reference rule bits.

11. The device according to claim 10, characterized in that The end division conditions include: the target spatial factor of the current decision tree is not less than the spatial factor threshold; wherein the spatial factor threshold is the maximum value that the configured target spatial factor can reach; or, the spatial factor threshold is the configured maximum floating point number; the number of leaf node rule references of the current decision tree is not greater than the leaf node rule number threshold; wherein the leaf node rule number threshold is a configured specified value; the matching rules in each leaf node of the current decision tree cannot be further divided into two leaf nodes; The establishment module determines whether the current decision tree meets the end division condition and is specifically used for: If the current decision tree satisfies any one of the end-partitioning conditions, it is determined that the current decision tree has satisfied the end-partitioning condition; if the current decision tree does not satisfy any of the end-partitioning conditions, it is determined that the current decision tree does not satisfy the end-partitioning condition.

12. The device according to claim 10, characterized in that When the establishment module obtains multiple matching rules, it is specifically used to: obtain multiple configured original rules; for each original rule, if the number of rule bits in the original rule is less than K, then add wildcard bits to the original rule to obtain a supplemented rule, and the number of rule bits in the supplemented rule is equal to K, K is 2 to the power of n, and n is a positive integer; if the supplemented rule includes an exact matching bit and a general matching bit, then all the exact matching bits are left-aligned and continuously arranged or right-aligned and continuously arranged to obtain a matching rule corresponding to the original rule; When acquiring the target matching data, the acquisition module is specifically used to: acquire the input original matching data; if the number of data bits in the original matching data is less than K, supplement the original matching data with special character bits to obtain supplemented matching data, and the number of data bits in the supplemented matching data is equal to K; if the supplemented matching data includes exact matching bits and special character bits, all exact matching bits are left-aligned and arranged continuously or right-aligned and arranged continuously to obtain the target matching data.

13. The device according to claim 10, characterized in that The matching rule in the leaf node is configured with a target mask, and the target mask includes a primary mask and a secondary mask; Among them, the K rule bits of the matching rule are divided into P1 data blocks, the first-level mask includes P1 bits, and the P1 bits correspond to the P1 data blocks one by one; for each data block, when there is an exact matching bit in the data block, the value of the bit corresponding to the data block in the first-level mask is a first value, and the first value indicates that there is an exact matching bit in the data block; when there is no exact matching bit in the data block, the value of the bit corresponding to the data block in the first-level mask is a second value, and the second value indicates that there is no exact matching bit in the data block; Wherein, for each data block, when there is no exact matching bit in the data block, the data block does not correspond to the secondary mask; when there is an exact matching bit in the data block, the data block corresponds to the secondary mask; Wherein, when the data block includes P2 regular bits, the secondary mask includes P2 bits, and the P2 bits correspond one-to-one to the P2 regular bits; For each rule bit, when the rule bit is an exact match bit, the value of the bit corresponding to the rule bit in the secondary mask is a third value, and the third value represents an exact match bit; when the rule bit is a general match bit, the value of the bit corresponding to the rule bit in the secondary mask is a fourth value, and the fourth value represents a general match bit.

14. The device according to claim 10, characterized in that If the first matching rule and the second matching rule exist in the leaf node, and the rule scope of the first matching rule is larger than the rule scope of the second matching rule, the position of the second matching rule in the leaf node is located in front of the position of the first matching rule in the leaf node.

15. An electronic device, characterized in that: include: a processor and a machine-readable storage medium storing machine-executable instructions executable by the processor; The processor is used to execute machine executable instructions to implement the method described in any one of claims 1-7.

Citation Information

Patent Citations

  • Data packet matching method and device, network equipment and storage medium

    CN110708317A

  • Data packet classification method based on information entropy

    CN114637773A

  • Rule storage method and device, electronic equipment and storage medium

    CN115834340A

  • Message processing method and device, equipment and medium

    CN115834515A

  • Dictionary tree domain name matching method and device, equipment and storage medium

    CN116346777A

Cited By

  • Data storage method, device and equipment

    CN120710932A