Decision tree generation method and apparatus, electronic device, and storage medium

CN116743474BActive Publication Date: 2026-10-09BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310786625.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-06-29
Publication Date
2026-10-09
Estimated Expiration
2043-06-29

AI Technical Summary

Technical Problem

由于,现有技术中线上策略是人工挖掘得到的,所以现有技术存在生成线上策略的周期较长、浪费人力、且无法保证生成的线上策略的准确性的问题

Benefits of technology

[0019] According to a fifth aspect of this disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements the decision tree generation method of the first aspect described above.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116743474B_ABST
    Figure CN116743474B_ABST
Patent Text Reader

Abstract

The disclosure provides a decision tree generation method and device, electronic equipment and a storage medium, relates to the technical field of computers, and in particular to the technical field of Internet, big data and the like. The specific implementation scheme is: determining the attribute of at least one target identifier in at least one target identifier dimension included in each first traffic data based on the attribute of multiple reference identifiers under multiple candidate identifier dimensions; aggregating multiple first traffic data based on the attribute of at least one target identifier in at least one target identifier dimension included in each first traffic data to obtain one or more samples; determining the label of each sample; generating a decision tree based on each sample and its label, and the decision tree is used to obtain one or more combination rules capable of detecting whether the to-be-tested traffic data is abnormal traffic data. The technical scheme provided by the disclosure can efficiently and accurately generate a decision tree, and then can efficiently and accurately mine combination rules for detecting to-be-tested traffic data through the decision tree.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of computer technology, and in particular to the fields of the Internet and big data. Background Technology

[0002] In existing technologies, online strategies are typically used to detect abnormal traffic, such as web crawler traffic. These online strategies are usually combinations of sub-rules obtained through manual discovery. Because these online strategies are manually discovered, existing technologies suffer from problems such as long generation cycles, wasted manpower, and inability to guarantee the accuracy of the generated strategies. Summary of the Invention

[0003] This disclosure provides a decision tree generation method, apparatus, electronic device, and storage medium.

[0004] According to a first aspect of this disclosure, a method for generating decision trees is provided, comprising:

[0005] Based on the attributes of multiple reference identifiers under multiple candidate identifier dimensions, determine the attributes of at least one target identifier under at least one target identifier dimension contained in each of the multiple first traffic data;

[0006] Based on the attribute of at least one target identifier under at least one target identifier dimension contained in each of the first traffic data, the multiple first traffic data are aggregated to obtain one or more samples, wherein each of the one or more samples includes one or more first traffic data;

[0007] Determine the label for each sample, wherein the label for each sample is used to indicate whether each sample is an abnormal traffic sample;

[0008] Based on each sample and its annotation, a decision tree is generated, wherein the decision tree is used to obtain one or more combined rules, and the one or more combined rules are used to detect whether the traffic data to be tested is abnormal traffic data.

[0009] According to a second aspect of this disclosure, a decision tree generation apparatus is provided, comprising:

[0010] The first attribute determination module is used to determine the attribute of at least one target identifier under at least one target identifier dimension contained in each of the multiple first traffic data based on the attributes of multiple reference identifiers under multiple candidate identifier dimensions.

[0011] The sample acquisition module is used to aggregate the multiple first traffic data based on the attributes of at least one target identifier under at least one target identifier dimension contained in each first traffic data to obtain one or more samples, wherein each of the one or more samples includes one or more first traffic data;

[0012] The labeling determination module is used to determine the label of each sample, wherein the label of each sample is used to indicate whether each sample is an abnormal traffic sample;

[0013] The decision tree generation module is used to generate a decision tree based on each sample and the label of each sample, wherein the decision tree is used to obtain one or more combined rules, and the one or more combined rules are used to detect whether the traffic data to be tested is abnormal traffic data.

[0014] According to a third aspect of this disclosure, an electronic device is provided, comprising:

[0015] At least one processor; and

[0016] A memory communicatively connected to the at least one processor; wherein,

[0017] The memory stores instructions that can be executed by the at least one processor, which, when executed by the at least one processor, enables the at least one processor to perform the decision tree generation method described in the first aspect.

[0018] According to a fourth aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions for causing the computer to perform the decision tree generation method of the first aspect described above.

[0019] According to a fifth aspect of this disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements the decision tree generation method of the first aspect described above.

[0020] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description.

[0021] The technical solution provided in this disclosure automatically generates a decision tree based on the attributes of the target identifier under the target identifier dimension contained in each first traffic data. This decision tree can generate combined rules for detecting whether the traffic data to be tested is abnormal traffic data. In this way, the decision tree can be generated efficiently and accurately, and then the combined rules for detecting the traffic data to be tested can be efficiently and accurately mined through the decision tree. This avoids the problems of long generation cycle, waste of manpower, and inability to guarantee the accuracy of the generated online strategy caused by manually mining the combination of sub-rules in the prior art. Attached Figure Description

[0022] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein:

[0023] Figure 1 This is a flowchart illustrating a decision tree generation method according to an embodiment of the present disclosure;

[0024] Figure 2 This is a schematic diagram of the structure of a decision tree provided according to an embodiment of the present disclosure;

[0025] Figure 3 This is a schematic block diagram of a decision tree generation apparatus provided according to an embodiment of the present disclosure;

[0026] Figure 4 This is another schematic block diagram of a decision tree generation apparatus provided according to an embodiment of the present disclosure;

[0027] Figure 5 This is a block diagram of an electronic device provided in an embodiment of this disclosure. Detailed Implementation

[0028] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.

[0029] The first aspect of this disclosure provides a method for generating a decision tree, such as... Figure 1 As shown, it includes:

[0030] S101, based on the attributes of multiple reference identifiers under multiple candidate identifier dimensions, determine the attributes of at least one target identifier under at least one target identifier dimension contained in each of the multiple first traffic data;

[0031] S102, based on the attribute of at least one target identifier under at least one target identifier dimension contained in each first traffic data, the multiple first traffic data are aggregated to obtain one or more samples, wherein each of the one or more samples includes one or more first traffic data;

[0032] S103, determine the label of each sample, wherein the label of each sample is used to indicate whether each sample is an abnormal traffic sample;

[0033] S104, Based on each sample and the label of each sample, a decision tree is generated, wherein the decision tree is used to obtain one or more combined rules, and the one or more combined rules are used to detect whether the traffic data to be tested is abnormal traffic data.

[0034] The decision tree generation method described above can be implemented by an electronic device. For example, the electronic device can be a terminal or server with computing and / or processing capabilities.

[0035] The technical solution provided in this disclosure automatically generates a decision tree based on the attributes of the target identifier under the target identifier dimension contained in each first traffic data. This decision tree can generate combined rules for detecting whether the traffic data to be tested is abnormal traffic data, realizing efficient and accurate generation of decision trees. Then, based on the decision tree, combined rules for detecting the traffic data to be tested can be efficiently and accurately mined, avoiding the problems of long cycle, waste of manpower, and inability to guarantee the accuracy of the generated online strategy caused by manually mining the combination of sub-rules in the prior art.

[0036] In some possible implementations, the method further includes: obtaining multiple reference identifiers under multiple candidate identifier dimensions of the multiple second traffic data based on multiple second traffic data, wherein each second traffic data is different from each first traffic data; performing feature analysis on each of the multiple reference identifiers to obtain at least one feature corresponding to each reference identifier; and determining the attributes of the multiple reference identifiers under the multiple candidate identifier dimensions based on the at least one feature corresponding to each reference identifier.

[0037] The processing timing for the attributes of multiple reference identifiers under multiple candidate identifier dimensions can be: before determining the attributes of at least one target identifier under at least one target identifier dimension contained in each of the multiple first traffic data based on the attributes of multiple reference identifiers under multiple candidate identifier dimensions.

[0038] The multiple second traffic data can be obtained through the first part of the logs of the target business line. The target business line can be at least one of multiple business lines, and the first part of the logs can be logs generated before the target business line obtains the first traffic. The methods for obtaining the multiple second traffic data may include: standardizing the first part of the logs of the target business line to obtain standardized first part logs of the target business line, and obtaining the multiple second traffic data based on the standardized first part logs. The standardization process may include at least one of the following: data cleaning, field extraction, database storage, etc., without exhaustive or limited examples.

[0039] For example, the fact that each of the plurality of second traffic data is different from each of the first traffic data can mean that: the data of the plurality of second traffic data and the plurality of first traffic data are all historical business traffic data of the target business line, but the time domain range of the plurality of second traffic data can be different from the time domain range of the plurality of first traffic data. For example, the generation time of the plurality of second traffic data can be earlier than the generation time of the plurality of first traffic data; or, for example, the duration corresponding to the plurality of second traffic data can be the same as or different from the duration corresponding to the plurality of first traffic data.

[0040] For example, the time domain range of the plurality of second traffic data may be different from the time domain range of the plurality of first traffic data. For example, the plurality of second traffic data may be N days of historical traffic data of the target business line, and the plurality of first traffic data may be M days of historical traffic data of the target business line. N and M are both positive integers, and N and M may be the same or different. As long as N days are different from M days and N days are earlier than M days, they are within the protection scope of this embodiment.

[0041] The step of obtaining multiple reference identifiers under multiple candidate identifier dimensions based on multiple second traffic data may specifically include: extracting all identifiers contained in the multiple second traffic data to obtain multiple reference identifiers under multiple candidate identifier dimensions of the multiple second traffic data, wherein the identifier dimension to which any one of the multiple reference identifiers belongs is any one of the multiple candidate identifier dimensions.

[0042] Any candidate identifier (ID) dimension can be any of the following: IP (Internet Protocol) related dimensions, UA (user agent) dimensions, fingerprint identification dimensions, UID (user identification) dimensions, etc. It should be understood that any candidate identifier dimension can also include other identifier dimensions known in the art besides the above-mentioned dimensions, and this is not limited thereto. For example, the aforementioned multiple candidate identifier dimensions can include at least two of the following dimensions: IP related dimensions, UA dimensions, fingerprint identification dimensions, and UID dimensions. The IP related dimension can be an IP address dimension or an IPC address dimension. The IPC address refers to the first three segments of the IP address; for example, if the IP address is 1.2.3.4, then the first three segments of the IP address are 1.2.3. The fingerprint identification dimension can be, but is not limited to, the JA3 fingerprint dimension. JA3 fingerprinting is a method for fingerprinting transport layer security applications; JA3 fingerprints can uniquely identify the corresponding browser.

[0043] The candidate identifier dimension is equivalent to the identifier category, and any candidate identifier dimension can have one or more reference identifiers. For example, when a candidate identifier dimension is the IP address dimension in the IP-related dimensions, the reference identifiers under this candidate identifier dimension can include at least one of the following: first IP address, second IP address, third IP address, etc.

[0044] The step of performing feature analysis on each of the plurality of reference identifiers to obtain at least one feature corresponding to each reference identifier may specifically include: aggregating multiple second traffic data based on each of the plurality of reference identifiers to obtain multiple clusters, wherein different clusters correspond to different reference identifiers, each cluster includes one or more second traffic data, and different clusters include different second traffic data; performing feature analysis on the one or more second traffic data contained in each cluster to obtain at least one feature corresponding to the reference identifier of each cluster.

[0045] The step of performing feature analysis based on one or more second traffic data contained in each cluster to obtain at least one feature corresponding to the reference identifier of each cluster may include: performing feature analysis on all second traffic data contained in the m-th cluster among the plurality of clusters to obtain at least one feature corresponding to the m-th reference identifier of the m-th cluster, wherein the m-th cluster is one of the plurality of clusters, and m is a positive integer. The aggregation process may be performed using a big data computing engine, such as the Spark platform.

[0046] The m-th reference identifier may include at least one of the following features: temporal features, statistical features, and algorithm model features.

[0047] The temporal feature can be a temporal access sequence. For example, when the m-th reference identifier is the first IP address, the temporal feature corresponding to the m-th reference identifier can be the first temporal access sequence of the user corresponding to the first IP address in the first time period, wherein the first time period includes multiple sub-time periods. Specifically, when the first time period is 1 day and each sub-time period in the first time period is 1 minute, the first temporal access sequence is a feature with 24*60=1440 dimensions. Each value in this 24*60=1440-dimensional feature can represent the number of accesses to the target business line within the corresponding 1 minute.

[0048] The statistical features can be features set based on statistical rules for overclocking ratios. For example, a statistical rule for overclocking ratios is: if an IP address (i.e., any reference identifier) ​​makes more than 1,000 requests in one minute, then the IP address is judged as an abnormal IP address; correspondingly, the statistical feature can be the number of requests made by the IP address in one minute.

[0049] The algorithm model features refer to the features output by the algorithm model. For example, the input to the algorithm model is each access request, and the output of the algorithm model is behavioral features such as click interval, mouse trajectory, and touch screen trajectory. These behavioral features are used as algorithm model features.

[0050] The step of determining the attributes of multiple reference identifiers under the multiple candidate identifier dimensions based on at least one feature corresponding to each reference identifier can specifically include: determining the attributes of the m-th reference identifier under the x-th candidate identifier dimension based on at least one feature corresponding to the m-th reference identifier under the x-th candidate identifier dimension. Wherein, the x-th candidate identifier dimension is one of the multiple candidate identifier dimensions, and the m-th reference identifier is one of the multiple reference identifiers under the x-th candidate identifier dimension, where x and m are positive integers. When determining the attributes of the m-th reference identifier under the x-th candidate identifier dimension, it can be determined based on only one feature among the at least one features corresponding to the m-th reference identifier under the x-th candidate identifier dimension, or it can be determined based on multiple features among the at least one features corresponding to the m-th reference identifier under the x-th candidate identifier dimension. The attributes of any reference identifier can be described as abnormal or normal, or as black or white; no limitation is made here. For ease of explanation, in this embodiment, the attributes of any reference identifier are exemplarily described as black or white.

[0051] When the attribute of any reference identifier is described as black or white, determining the attributes of multiple reference identifiers under the multiple candidate identifier dimensions can be understood as qualitatively classifying the multiple reference identifiers under the multiple candidate identifier dimensions as black or white. The attributes of the multiple reference identifiers under the multiple candidate identifier dimensions can be stored in a list. This list can store the attributes of multiple reference identifiers under the multiple candidate identifier dimensions. For example, the list may include: 15219 black IPs and 1937 white IPs under the IP dimension, 3259 black IPCs and 5187 white IPCs under the IPC dimension, and 3658 black UAs and 3903 white UAs under the UA dimension.

[0052] The above technical solution can determine the attributes of multiple reference identifiers under multiple candidate identifier dimensions based on multiple second traffic data, and then determine the attributes of at least one target identifier under at least one target identifier dimension contained in each of the multiple first traffic data. Therefore, the above technical solution can provide more accurate attributes of multiple reference identifiers under multiple candidate identifier dimensions, ensuring the accuracy of subsequently determining the attributes of at least one target identifier under at least one target identifier dimension contained in each of the first traffic data.

[0053] In some possible implementations, determining the attributes of at least one target identifier under at least one target identifier dimension contained in each of the multiple first traffic data based on the attributes of multiple reference identifiers under multiple candidate identifier dimensions may specifically include: matching the k-th target identifier contained in the q-th first traffic data with each of the multiple reference identifiers under the multiple candidate identifier dimensions; if the k-th target identifier has a matching reference identifier, using the candidate identifier dimension corresponding to the matching reference identifier as the target identifier dimension of the k-th target identifier of the q-th first traffic data, and using the attributes of the matching reference identifier as the attributes of the k-th target identifier of the q-th first traffic data. Wherein, the q-th first traffic data is any one of the multiple first traffic data, the k-th target identifier is one of the at least one target identifier, and q and k are positive integers.

[0054] The plurality of first traffic data can be obtained through the second part of the logs of the target business line. The first part of the logs and the second part of the logs of the target business line have different generation times, and the generation time of the first part of the logs is earlier than the generation time of the second part of the logs; the method of obtaining the plurality of first traffic data can be similar to the method of obtaining the plurality of second traffic data, and will not be described in detail here.

[0055] In some possible implementations, the method further includes: determining the sub-rule hit status of each first traffic data based on multiple preset sub-rules, wherein the sub-rule hit status of each first traffic data is used to indicate whether the features of each first traffic data hit each of the multiple preset sub-rules. Correspondingly, the step of aggregating the multiple first traffic data to obtain one or more samples based on the attributes of at least one target identifier under at least one target identifier dimension contained in each first traffic data includes: aggregating the multiple first traffic data to obtain the one or more samples based on the attributes of at least one target identifier under at least one target identifier dimension contained in each first traffic data and the sub-rule hit status of each first traffic data.

[0056] Any one of the preset sub-rules can be a sub-rule used to identify whether the first traffic data is abnormal traffic data. The number and content of multiple preset sub-rules are not limited in this embodiment.

[0057] The step of determining the sub-rule hit status of each first traffic data based on multiple preset sub-rules can specifically include: matching the feature of each first traffic data corresponding to the g-th preset sub-rule among the multiple preset sub-rules with the g-th preset sub-rule to determine the hit status of the g-th preset sub-rule; determining the hit status of each preset sub-rule among the multiple preset sub-rules based on the hit status of the g-th preset sub-rule; and determining the sub-rule hit status of each first traffic data based on the hit status of each preset sub-rule among the multiple preset sub-rules; where g is a positive integer.

[0058] Each of the one or more samples includes one or more first traffic data sets, and the first traffic data sets included in different samples are not the same. For example, the one or more first traffic data sets included in sample 0 are different from the one or more first traffic data sets included in sample 1. The one or more first traffic data sets included in the same sample have the same target identifier attributes and sub-rule hit status under the target identifier dimension. For example, each row in Table 1 represents a sample, and row 0 in Table 1 indicates that the JA3 fingerprint of each of the one or more first traffic data sets included in sample 0 is white, the UA is white, the IPC address is white, the IP address is black, the UID is white, the preset sub-rule 1 is hit, the preset sub-rule 2 is not hit, and the preset sub-rule 3 is not hit.

[0059] Table 1 Sample Example Table

[0060]

[0061] The aggregation process can be the same as the aforementioned process, and will not be repeated here.

[0062] By adopting the above technical solution, when generating the decision tree, multiple first traffic data are aggregated based on the attributes of at least one target identifier under at least one target identifier dimension contained in each first traffic data and the sub-rule hit status of each first traffic data, resulting in one or more samples. Therefore, compared to constructing a decision tree based solely on the one or more samples obtained by aggregating multiple first traffic data based on the attributes of at least one target identifier under at least one target identifier dimension contained in each first traffic data, the accuracy of the combined rules produced by the decision tree is improved.

[0063] In some possible implementations, determining the label of each sample includes at least one of the following: if the number of deduplicated users of the j-th sample is less than a preset threshold, determining that the label of the j-th sample indicates that the j-th sample is an abnormal traffic sample, wherein the j-th sample is one of the one or more samples, and j is a positive integer; if the number of deduplicated users of the j-th sample is greater than or equal to the preset threshold, determining that the label of the j-th sample indicates that the j-th sample is a normal traffic sample.

[0064] The j-th sample mentioned above can be any one of the one or more samples. Since each sample can undergo the same processing as the j-th sample, this embodiment will not elaborate on each one.

[0065] The preset threshold can be determined based on the actual scenario or user experience, and is not limited here. For example, the preset threshold can be 10 or 20, but is not limited to 10 or 20. This embodiment does not exhaustively list them all.

[0066] Before determining the label of the j-th sample, the number of duplicate users in the j-th sample can be counted in advance. The method for counting the number of duplicate users in the j-th sample can be to use a deduplication algorithm to remove duplicate users from the j-th sample, and then calculate the number of users in the j-th sample after deduplication, thus obtaining the number of duplicate users in the j-th sample. This embodiment does not exhaustively list or limit this method.

[0067] In this embodiment, the labeling of any sample can be represented in many ways. For example, as shown in Table 1, taking j=0 as an example, the label of the 0th sample can be "abnormal" or "normal"; for another example, the label of any sample can also be "abnormal traffic sample" or "normal traffic sample", etc., without limitation.

[0068] It should be noted that if any sample is labeled as an abnormal traffic sample, then every first traffic data point in that sample is also abnormal traffic data. For example, in Table 1, when j=0, the label of the 0th sample is abnormal, which means that all first traffic data points included in the 0th sample are abnormal traffic data.

[0069] Through the above technical solution, if the number of deduplicated users in a sample is less than a preset threshold, the sample is labeled as an abnormal traffic sample; or if the number of deduplicated users in a sample is greater than or equal to the preset threshold, the sample is labeled as a normal traffic sample. This achieves automatic labeling of samples and ensures the accuracy of the final generated decision tree when using the sample for decision tree training.

[0070] In some possible implementations, before generating the decision tree based on each sample and its annotation, relevant settings for generating the decision tree can be pre-configured. For example, this may include determining the input and the construction algorithm for this decision tree. The input includes a training set D, a preset Gini index threshold, and a preset sample size threshold. The training set D may include each sample and its annotation from one or more samples obtained by the above method; for example, the training set D may include all the contents of Table 1.

[0071] After completing the aforementioned processing of the relevant settings for generating the decision tree, the process of generating the decision tree based on each sample and its annotation is executed.

[0072] The decision tree construction algorithms mainly include ID3 (Iterative Dichotomy 3) algorithm and CART (Classification and Regression Tree) algorithm.

[0073] In a preferred example, the construction algorithm can be set to the CART algorithm. The following explanation uses the CART algorithm to construct a CART decision tree as an example to illustrate the process of generating a decision tree based on each sample and its annotation:

[0074] Based on the training set D, starting from the root node, recursively perform the following operations on each node to generate a CART decision tree:

[0075] The first step is to calculate the Gini index of any attribute item in the training set D. All attribute items in the training set D can be obtained based on the attributes of at least one target identifier under at least one target identifier dimension in each first traffic data and the sub-rule hit status of each first traffic data. For example, when the training set D includes all the content in Table 1, the attribute items are the content shown in the header of Table 1, such as one attribute item being "JA3 fingerprint is white". For each attribute item, it can be denoted as attribute item A, and the Gini index of attribute item A is:

[0076]

[0077] Where D represents the training set D; D1 represents the training subset D1 divided into "yes" and "no" based on each sample's judgment result for A=a; D2 represents the training subset D2 divided into "no" and "no" based on each sample's judgment result for A=a; A represents attribute A; a is the possible value of attribute A, such as "yes" or "no" in Table 1; GiniIndex(D|A=a) represents the Gini index of attribute A in the training set D; Gini(D) represents the uncertainty of the training set D; and Gini(D,A) represents the uncertainty of the training set D after the A=a partition.

[0078] The second step is to select the attribute with the smallest Gini index and its corresponding value from all attribute items and their values ​​as the optimal attribute item and optimal value. Based on the optimal attribute item and optimal value, two sub-decision nodes are generated from the current decision node, and the training set D is distributed among the two sub-decision nodes. The current decision node is the node corresponding to the optimal attribute item. It should be noted that a decision node is any node in the decision tree other than the leaf nodes.

[0079] The third step involves recursively executing the first and second steps for the two sub-decision nodes until a stopping condition is met, generating a CART decision tree. The stopping condition may include at least one of the following: maximum tree depth, minimum number of samples required to split a decision node, minimum number of samples required for a leaf node, minimum impurity of a decision node split, maximum number of leaf nodes, or no more features.

[0080] In some possible implementations, the method further includes one of the following: generating one or more combined rules based on one or more paths contained in the decision tree, wherein different paths in the one or more paths are used to generate different combined rules, and the different paths in the one or more paths include multiple nodes from the root node to different leaf nodes of the decision tree; and filtering to obtain one or more combined rules based on one or more paths contained in the decision tree.

[0081] Taking the nth path as any one of one or more paths as an example, the following describes the multiple nodes from the root node to different leaf nodes of the decision tree in different paths of the one or more paths. The multiple nodes included in the nth path can include at least: the root node and the nth leaf node; furthermore, the multiple nodes included in the nth path can also include one or more other nodes between the root node (excluding the root node) and the nth leaf node (excluding the nth leaf node). Here, n is a positive integer.

[0082] Taking the nth path as an example, each node (including each other node and / or the root node) among the multiple nodes on the nth path, except for the nth leaf node, represents a judgment on an attribute item. Any branch (i.e., a sub-path from the upstream node to the downstream node in any two adjacent nodes on the nth path) represents the judgment result on an attribute item; the nth leaf node represents the classification result. Specifically, each node (including each other node and / or the root node) can represent a judgment on the attribute of at least one target identifier under at least one target identifier dimension contained in each first traffic data, or a judgment on the sub-rule hit status of each first traffic data; any branch represents the judgment result of yes or no on the attribute of at least one target identifier under at least one target identifier dimension contained in each first traffic data, or the judgment result of yes or no on the sub-rule hit status of each first traffic data; each leaf node represents the classification result as abnormal or normal.

[0083] For example, Figure 2 In this multi-node system, each node (excluding leaf nodes) corresponds to a rectangular bounding box, representing the judgment of an IP address being black, a UA being white, a JA3 fingerprint being black, an IPC address being white, and a UID being white, respectively. Each node has a branch representing the connection between nodes, indicating a yes or no judgment result. Each leaf node corresponds to an elliptical bounding box, representing a normal or abnormal classification result. Figure 2 Taking the nth path from root node 21, through node 22 to leaf node 23 as an example, root node 21 represents the judgment that the IP address is black, branch 22 represents the judgment that the IP address is black is yes, node 23 represents the judgment that the UA is white, branch 24 represents the judgment that the UA is white is yes, and leaf node 25 represents the classification result as normal.

[0084] The above example only uses the nth path as an illustration. Since the nth path is any one of one or more paths, the description of each path is the same as that of the nth path. The only difference is that different paths contain different leaf nodes and at least some of the other nodes are different, so they will not be described in detail.

[0085] Through the above technical solution, after generating the decision tree, each path contained in the decision tree can be used to generate a combination rule to obtain one or more combination rules; or one or more paths contained in the decision tree can be filtered, and one or more suitable paths can be selected according to the needs to obtain one or more combination rules to meet different user needs.

[0086] In some possible implementations, the step of filtering one or more combined rules based on one or more paths contained in the decision tree includes: if the i-th path in the one or more paths contained in the decision tree satisfies a preset condition, obtaining one of the one or more combined rules based on the i-th path; wherein the preset condition includes at least one of the following: the Gini index of the leaf node of the i-th path is 0; at least one node among the multiple nodes of the i-th path is used to determine the attribute of the at least one target identifier under the at least one target identifier dimension; the traffic scale corresponding to the leaf node of the i-th path is greater than a preset scale; i is a positive integer.

[0087] The above scheme provides three preset conditions, the first of which is that the Gini index of the leaf node of the i-th path is 0. In this first condition, the method for calculating the Gini index of the leaf node of the i-th path is similar to the method for calculating the Gini index of attribute item A, and will not be repeated here. A Gini index of 0 for the leaf node of the i-th path indicates that the purity of the i-th path is 100%. Therefore, based on the first condition, it is possible to filter out combination rules with high purity.

[0088] The second condition in the preset conditions provided in the above scheme is: at least one node among the multiple nodes of the i-th path is used to determine the attribute of the at least one target identifier under the at least one target identifier dimension. That is, any selected path must satisfy the condition of the attribute of the at least one target identifier under the at least one target identifier dimension. The combined rules obtained based on such paths can include the determination of the attribute of the at least one target identifier under the at least one target identifier dimension, thereby making the detection accuracy of the combined rules higher and more in line with user expectations. Therefore, based on the second condition, it is possible to filter out combined rules that meet user expectations and have high accuracy.

[0089] The third condition in the above scheme is that the traffic volume corresponding to the leaf node of the i-th path is greater than a preset volume. In this third condition, the traffic volume corresponding to the leaf node of the i-th path refers to the traffic volume of all first traffic data covering the i-th path among the multiple first traffic data used to construct the decision tree. Based on this third condition, it is possible to filter out combined rules generated from a larger scale of first traffic data, thus achieving the selection of combined rules with better universality.

[0090] In some possible examples, any one of the three conditions mentioned above can be used as the preset condition.

[0091] For example, the first condition described above can be used as a preset condition. Correspondingly, the step of filtering one or more combination rules based on one or more paths contained in the decision tree can be: if the i-th path in the one or more paths contained in the decision tree satisfies that the Gini index of the leaf node of the i-th path is 0, then one of the one or more combination rules is generated based on the i-th path. The above example only illustrates the use of the first condition as a preset condition. Since the explanations for using the second or third condition as a preset condition are similar to those for using the first condition, they will not be elaborated upon.

[0092] In some possible examples, any two of the above three conditions can be used as preset conditions.

[0093] For example, the first and second conditions described above can be used as preset conditions. Correspondingly, the step of filtering one or more combination rules based on one or more paths contained in the decision tree can be: if the i-th path in the decision tree satisfies that the Gini index of the leaf node of the i-th path is 0, and at least one node among the multiple nodes of the i-th path is used to determine the attribute of the at least one target identifier under the at least one target identifier dimension, then one of the one or more combination rules is generated based on the i-th path. When using the first and second conditions described above as preset conditions to filter and obtain combination rules, the order of judgment of the preset conditions is not limited. For example, the judgment order can be: for the i-th path in one or more paths, first determine whether the i-th path satisfies the first condition; if the first condition is satisfied, then determine whether the i-th path satisfies the second condition. For example, the judgment order can also be: for the i-th path among one or more paths, first determine whether the i-th path satisfies the second condition; if it satisfies the second condition, then determine whether the i-th path satisfies the first condition. Alternatively, the judgment order can also be: for the i-th path among one or more paths, simultaneously determine whether the i-th path satisfies both the first and second conditions, obtaining two judgment results; if both results are yes, determine that the i-th path simultaneously satisfies both the first and second conditions. The above is only an example of using the first and second conditions as preset conditions. Since the explanation of using the second and third conditions as preset conditions or using the first and third conditions as preset conditions is similar to that of using the first and second conditions as preset conditions, it will not be elaborated upon further.

[0094] In some possible examples, the above three conditions can be used as preset conditions.

[0095] The step of filtering one or more combined rules based on one or more paths contained in the decision tree can be as follows: If the i-th path in the decision tree satisfies the following conditions: the Gini index of the leaf node of the i-th path is 0, and at least one node among the multiple nodes of the i-th path is used to determine the attribute of the at least one target identifier under the at least one target identifier dimension, and the traffic scale corresponding to the leaf node of the i-th path is greater than a preset scale, then one of the one or more combined rules is generated based on the i-th path. Similarly, when using the above three conditions as preset conditions to filter and obtain combined rules, the order of judgment of the preset conditions is not limited. For example, the judgment order can be to judge each of the three conditions one by one; it can also be to judge all three conditions simultaneously; it can also be to judge one condition first, and then judge two conditions simultaneously; or it can also be to judge two conditions simultaneously first, and then judge one condition.

[0096] For example, judging each of the three conditions one by one could be as follows: For the i-th path among one or more paths, first determine whether the i-th path satisfies the first condition; if it does, then determine whether it satisfies the second condition; if it does, then determine whether it satisfies the third condition. The above description is merely illustrative; the order of judging the three conditions can also differ from the above description. Due to space limitations, these differences will not be elaborated upon here.

[0097] For example, simultaneously judging the three conditions can be: for the i-th path in one or more paths, simultaneously judging whether the i-th path satisfies the first condition, the second condition, and the third condition, obtaining three judgment results, and if all three judgment results are yes, determining that the i-th path simultaneously satisfies the first condition, the second condition, and the third condition.

[0098] Since the methods of first judging one condition and then judging two conditions simultaneously, and the methods of first judging two conditions simultaneously and then judging one condition, can be illustrated by combining examples of judging using the first and second conditions as preset conditions and examples of judging each of the three conditions one by one when using three conditions as preset conditions, it is undoubtedly clear that, in order to save space, they will not be elaborated on here.

[0099] In one specific implementation, the first, second, and third conditions mentioned above are used as preset conditions, and the three conditions are judged simultaneously. The processing of obtaining multiple combination rules based on the generated decision tree includes: filtering the 54 paths contained in the decision tree based on the Gini index of the leaf node being 0, resulting in 32 paths, i.e., 32 combination rules; then, filtering the 32 paths based on the attribute of at least one node among the multiple nodes used to judge the at least one target identifier dimension, resulting in 18 paths, i.e., 18 combination rules; finally, filtering the 18 paths based on the traffic scale corresponding to the leaf node being greater than a preset scale, resulting in 12 paths, i.e., 12 combination rules, which are the final combination rules.

[0100] Through the above technical solution, after generating the decision tree, one or more paths contained in the decision tree can be filtered based on preset conditions. One or more paths that meet the preset conditions are filtered to generate one or more combination rules. This achieves the goal of filtering out combination rules with at least one of the following advantages: high purity, meeting user expectations, and good universality, thereby ensuring the accuracy and effectiveness of the final combination rules.

[0101] A second aspect of this disclosure provides a decision tree generation apparatus, such as... Figure 3 As shown, it includes:

[0102] The first attribute determination module 301 is used to determine the attribute of at least one target identifier under at least one target identifier dimension contained in each of the multiple first traffic data based on the attributes of multiple reference identifiers under multiple candidate identifier dimensions.

[0103] The sample acquisition module 302 is used to aggregate the multiple first traffic data based on the attributes of at least one target identifier under at least one target identifier dimension contained in each first traffic data to obtain one or more samples, wherein each of the one or more samples includes one or more first traffic data.

[0104] The labeling determination module 303 is used to determine the label of each sample, wherein the label of each sample is used to indicate whether each sample is an abnormal traffic sample;

[0105] The decision tree generation module 304 is used to generate a decision tree based on each sample and the annotation of each sample, wherein the decision tree is used to obtain one or more combined rules, and the one or more combined rules are used to detect whether the traffic data to be tested is abnormal traffic data.

[0106] In some possible implementations, the sample acquisition module 302 is used to determine the sub-rule hit status of each first traffic data based on multiple preset sub-rules, wherein the sub-rule hit status of each first traffic data is used to indicate whether the features of each first traffic data hit each of the multiple preset sub-rules; and to aggregate the multiple first traffic data based on the attributes of at least one target identifier under at least one target identifier dimension contained in each first traffic data and the sub-rule hit status of each first traffic data to obtain the one or more samples.

[0107] like Figure 4 As shown, in some possible implementations, the apparatus further includes a rule generation module 305. The rule generation module 305 is configured to perform one of the following: generate one or more combined rules based on one or more paths contained in the decision tree; wherein different paths in the one or more paths are used to generate different combined rules, and the different paths in the one or more paths include multiple nodes from the root node to different leaf nodes of the decision tree; and filter to obtain one or more combined rules based on one or more paths contained in the decision tree.

[0108] In some possible implementations, the rule generation module 305 is used to obtain one of the one or more combined rules based on the i-th path when the i-th path in one or more paths contained in the decision tree satisfies a preset condition; wherein the preset condition includes at least one of the following: the Gini index of the leaf node of the i-th path is 0; at least one node of the plurality of nodes of the i-th path is used to determine the attribute of the at least one target identifier under the at least one target identifier dimension; the traffic scale corresponding to the leaf node of the i-th path is greater than a preset scale; i is a positive integer.

[0109] In some possible implementations, the labeling determination module 303 is configured to perform at least one of the following: if the number of deduplicated users in the j-th sample is less than a preset threshold, determine that the label of the j-th sample indicates that the j-th sample is an abnormal traffic sample, wherein the j-th sample is one of the one or more samples, and j is a positive integer; if the number of deduplicated users in the j-th sample is greater than or equal to the preset threshold, determine that the label of the j-th sample indicates that the j-th sample is a normal traffic sample.

[0110] Please refer to it again. Figure 4In some possible implementations, the apparatus further includes a second attribute determination module 306. The second attribute determination module 306 is configured to: obtain multiple reference identifiers under multiple candidate identifier dimensions of the multiple second traffic data based on multiple second traffic data, wherein each second traffic data is different from each first traffic data; perform feature analysis on each of the multiple reference identifiers to obtain at least one feature corresponding to each reference identifier; and determine the attributes of the multiple reference identifiers under the multiple candidate identifier dimensions based on the at least one feature corresponding to each reference identifier.

[0111] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0112] Figure 5 A schematic block diagram of an example electronic device 500 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0113] like Figure 5 As shown, the electronic device 500 includes a computing unit 501, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 502 or a computer program loaded from a storage unit 508 into a random access memory (RAM) 503. The RAM 503 may also store various programs and data required for the operation of the electronic device 500. The computing unit 501, ROM 502, and RAM 503 are interconnected via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.

[0114] Multiple components in electronic device 500 are connected to I / O interface 505, including: input unit 506, such as keyboard, mouse, etc.; output unit 507, such as various types of monitors, speakers, etc.; storage unit 508, such as disk, optical disk, etc.; and communication unit 509, such as network card, modem, wireless transceiver, etc. Communication unit 509 allows electronic device 500 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0115] The computing unit 501 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 501 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 501 performs the various methods and processes described above. For example, in some embodiments, the various methods described above can be implemented as computer software programs tangibly contained in a machine-readable medium, such as storage unit 508. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 500 via ROM 502 and / or communication unit 509. When the computer program is loaded into RAM 503 and executed by the computing unit 501, one or more steps of the various methods described above can be performed. Alternatively, in other embodiments, the computing unit 501 can be configured to perform the various methods described above by any other suitable means (e.g., by means of firmware).

[0116] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0117] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0118] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0119] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0120] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with embodiments of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.

[0121] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, servers in distributed systems, or servers incorporating blockchain technology.

[0122] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.

[0123] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.

Claims

1. A decision tree generation method, comprising: Based on the attributes of multiple reference identifiers under multiple candidate identifier dimensions, determine the attributes of at least one target identifier under at least one target identifier dimension contained in each of the multiple first traffic data; Based on the attribute of at least one target identifier under at least one target identifier dimension contained in each of the first traffic data, the multiple first traffic data are aggregated to obtain one or more samples, wherein each of the one or more samples includes one or more first traffic data; Determine the label for each sample, wherein the label for each sample is used to indicate whether each sample is an abnormal traffic sample; Based on each sample and the label of each sample, a decision tree is generated, wherein the decision tree is used to obtain one or more combined rules, and the one or more combined rules are used to detect whether the traffic data to be tested is abnormal traffic data; The method further includes: determining the sub-rule hit status of each first traffic data based on multiple preset sub-rules, wherein the sub-rule hit status of each first traffic data is used to indicate whether the features of each first traffic data hit each preset sub-rule among the multiple preset sub-rules; The step of aggregating the multiple first traffic data based on the attributes of at least one target identifier under at least one target identifier dimension contained in each first traffic data to obtain one or more samples includes: aggregating the multiple first traffic data based on the attributes of at least one target identifier under at least one target identifier dimension contained in each first traffic data and the sub-rule hit status of each first traffic data to obtain the one or more samples.

2. The method according to claim 1, wherein, The method also includes one of the following: Based on one or more paths contained in the decision tree, one or more combination rules are generated; wherein, different paths in the one or more paths are used to generate different combination rules, and the different paths in the one or more paths include multiple nodes from the root node to different leaf nodes of the decision tree. Based on one or more paths contained in the decision tree, one or more combined rules are obtained through filtering.

3. The method according to claim 2, wherein, The step of filtering one or more combined rules based on one or more paths contained in the decision tree includes: If the i-th path in one or more paths contained in the decision tree satisfies a preset condition, one of the one or more combination rules is obtained based on the i-th path; wherein the preset condition includes at least one of the following: the Gini index of the leaf node of the i-th path is 0; at least one node in the plurality of nodes of the i-th path is used to determine the attribute of the at least one target identifier under the at least one target identifier dimension; the traffic scale corresponding to the leaf node of the i-th path is greater than a preset scale; i is a positive integer.

4. The method according to any one of claims 1-3, wherein, Determining the label of each sample includes at least one of the following: If the number of unique users in the j-th sample is less than a preset threshold, the label of the j-th sample is determined to indicate that the j-th sample is an abnormal traffic sample, wherein the j-th sample is one of the one or more samples, and j is a positive integer; If the number of unique users in the j-th sample is greater than or equal to the preset threshold, the label of the j-th sample is determined to indicate that the j-th sample is a normal traffic sample.

5. The method according to any one of claims 1-3, wherein, The method further includes: Based on multiple second traffic data, multiple reference identifiers are obtained under multiple candidate identifier dimensions of the multiple second traffic data, wherein each of the multiple second traffic data is different from each of the first traffic data; Perform feature analysis on each of the plurality of reference identifiers to obtain at least one feature corresponding to each reference identifier; Based on at least one feature corresponding to each reference identifier, the attributes of multiple reference identifiers under the multiple candidate identifier dimensions are determined.

6. A decision tree generation apparatus, comprising: The first attribute determination module is used to determine the attribute of at least one target identifier under at least one target identifier dimension contained in each of the multiple first traffic data based on the attributes of multiple reference identifiers under multiple candidate identifier dimensions. The sample acquisition module is used to aggregate the multiple first traffic data based on the attributes of at least one target identifier under at least one target identifier dimension contained in each first traffic data to obtain one or more samples, wherein each of the one or more samples includes one or more first traffic data; The labeling determination module is used to determine the label of each sample, wherein the label of each sample is used to indicate whether each sample is an abnormal traffic sample; The decision tree generation module is used to generate a decision tree based on each sample and the label of each sample, wherein the decision tree is used to obtain one or more combined rules, and the one or more combined rules are used to detect whether the traffic data to be tested is abnormal traffic data; The sample acquisition module is further configured to determine the sub-rule hit status of each first traffic data based on multiple preset sub-rules, wherein the sub-rule hit status of each first traffic data is used to indicate whether the features of each first traffic data hit each preset sub-rule among the multiple preset sub-rules; and to aggregate the multiple first traffic data based on the attributes of at least one target identifier under at least one target identifier dimension contained in each first traffic data and the sub-rule hit status of each first traffic data to obtain the one or more samples.

7. The apparatus according to claim 6, wherein, The device further includes a rule generation module for performing one of the following: Based on one or more paths contained in the decision tree, one or more combination rules are generated; wherein, different paths in the one or more paths are used to generate different combination rules, and the different paths in the one or more paths include multiple nodes from the root node to different leaf nodes of the decision tree. Based on one or more paths contained in the decision tree, one or more combined rules are obtained through filtering.

8. The apparatus according to claim 7, wherein, The rule generation module is used to obtain one of the one or more combined rules based on the i-th path when the i-th path in one or more paths contained in the decision tree satisfies a preset condition; wherein the preset condition includes at least one of the following: the Gini index of the leaf node of the i-th path is 0; at least one node in the plurality of nodes of the i-th path is used to determine the attribute of the at least one target identifier under the at least one target identifier dimension; the traffic scale corresponding to the leaf node of the i-th path is greater than a preset scale; i is a positive integer.

9. The apparatus according to any one of claims 6-8, wherein, The annotation determination module is configured to perform at least one of the following: If the number of unique users in the j-th sample is less than a preset threshold, the label of the j-th sample is determined to indicate that the j-th sample is an abnormal traffic sample, wherein the j-th sample is one of the one or more samples, and j is a positive integer; If the number of unique users in the j-th sample is greater than or equal to the preset threshold, the label of the j-th sample is determined to indicate that the j-th sample is a normal traffic sample.

10. The apparatus according to any one of claims 6-8, wherein, The device further includes: a second attribute determination module, configured to obtain multiple reference identifiers under multiple candidate identifier dimensions of the multiple second traffic data based on multiple second traffic data, wherein each second traffic data is different from each first traffic data; perform feature analysis on each of the multiple reference identifiers to obtain at least one feature corresponding to each reference identifier; and determine the attributes of the multiple reference identifiers under the multiple candidate identifier dimensions based on the at least one feature corresponding to each reference identifier.

11. An electronic device, comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-5.

12. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-5.

13. A computer program product comprising a computer program that, when executed by a processor, implements the method according to any one of claims 1-5.

Citation Information

Patent Citations

  • Traffic processing method and system, terminal equipment and storage medium

    CN115022050A

  • Feature combination determination method and device, equipment, storage medium and program product

    CN115221948A