Anomaly Detection Method and Device
By performing feature dimension classification and random forest decision tree analysis on e-commerce platform user data, the problem that e-commerce platforms cannot accurately detect fake data manufacturing behavior is solved, and comprehensive detection of abnormal accounts is achieved to ensure data accuracy.
Patent Information
- Application Number
- CN202210404122.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-18
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2042-04-18
AI Technical Summary
E-commerce platforms cannot fully and accurately detect the behavior of creating fake data, resulting in inaccurate data and affecting user transactions.
By obtaining the data to be detected, including the data of the first user and the second user associated therewith, classifying it according to the preset feature dimension, and inputting the classified feature data into the preset decision tree, the decision tree generated by the random forest model training is used for abnormal detection, and determining whether there is an abnormality for the user.
It realizes comprehensive and accurate detection of abnormal accounts, effectively reduces the occurrence of e-commerce data fraud and ensures the accuracy of data.
Smart Images

Figure CN114677150B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of data processing, and in particular to an anomaly detection method and device. Background Art
[0002] The rapid development of the e-commerce industry has led to some e-commerce stores creating fake "high-quality data" by means of false shipments and false evaluations to improve the sales volume and evaluation results displayed externally by the store. In response to such behaviors, currently, e-commerce platforms mainly regard stores as the objects of control and identify stores with false data by checking aspects such as the conversion rate of the store, the concentrated shopping time, and the search situation. This method lacks supervision at the consumer user level and thus cannot effectively reduce the behavior of creating false data through false transactions. Summary of the Invention
[0003] Therefore, this application provides an anomaly detection method and device to solve the problem that the behavior of creating false data in the e-commerce platform cannot be detected comprehensively and accurately, resulting in inaccurate e-commerce data and affecting user transactions.
[0004] To achieve the above object, the first aspect of this application provides an anomaly detection method, which includes:
[0005] Obtain the data to be detected, where the data to be detected includes the data of a first user and the data of a second user associated with the first user;
[0006] Classify the data to be detected according to a preset feature dimension to obtain classified feature data;
[0007] Input the classified feature data into the corresponding preset decision tree to obtain the classification sub-results of each preset decision tree;
[0008] Determine the anomaly detection result of the first user according to multiple classification sub-results, where the anomaly detection result is used to indicate whether the first user has an anomaly.
[0009] Furthermore, the preset decision tree is a decision tree obtained by training a random forest model, and the random forest model is a model based on a random forest and used to generate the preset decision tree for anomaly detection.
[0010] Furthermore, the method further includes:
[0011] Determine multiple preset feature dimensions;
[0012] Classify the training data according to the preset feature dimension to obtain classified feature training data;
[0013] Calculate the Gini coefficient of the training data for each categorical feature, and determine each level of classification node based on the Gini coefficient.
[0014] Generate corresponding preset decision trees based on each level of classification node, and obtain the random forest model.
[0015] Further, the calculating the Gini coefficient of the training data for each categorical feature and determining each level of classification node based on the Gini coefficient includes:
[0016] Calculate the Gini coefficient of the training data for each of the categorical features, and determine the minimum Gini coefficient at the first level.
[0017] Take the categorical feature corresponding to the minimum Gini coefficient at the first level as the first-level classification node, and divide the training data of the categorical features into multiple first-level branch training data according to the first-level classification node.
[0018] For each (n - 1)-th level branch training data, calculate the Gini coefficient of the (n - 1)-th level branch training data, determine the minimum Gini coefficient at the n-th level, take the categorical feature corresponding to the minimum Gini coefficient at the n-th level as the n-th level classification node, and divide the (n - 1)-th level branch training data into multiple n-th level branch training data, where n is an integer greater than or equal to 2;
[0019] When the preset stop condition is met, stop the data partitioning operation to obtain each level of classification node.
[0020] Further, the determining the anomaly detection result of the first user according to the multiple categorical sub-results includes:
[0021] Determine the weight value of each of the preset decision trees;
[0022] Obtain the anomaly detection result of the first user according to the weight value of each of the preset decision trees and the categorical sub-results.
[0023] Further, the first user is a seller user, and the second user is a seller user having a transaction behavior with the first user;
[0024] Or, the first user is a buyer user, and the second user is a seller user having a transaction behavior with the first user.
[0025] Further, after the determining the anomaly detection result of the first user according to the multiple categorical sub-results, it further includes:
[0026] Based on a review terminal, review the anomaly detection result to obtain a review result.
[0027] Further, the anomaly detection result indicates that there is an anomaly with the first user;
[0028] After determining the anomaly detection result based on multiple classification sub-results, the following steps are further included:
[0029] Obtain the transaction data of the first user within a preset transaction time period as supplementary detection data;
[0030] Based on the supplementary detection data and the preset decision tree, obtain the supplementary detection result of the first user.
[0031] Further, after determining the anomaly detection result of the first user based on multiple classification sub-results, the following steps are further included:
[0032] In the case where the number of times an anomaly is detected for the first user within a preset time period exceeds a preset number threshold, process the first user according to a preset penalty rule.
[0033] To achieve the above object, a second aspect of the present application provides an anomaly detection device, which includes:
[0034] A data acquisition module, configured to acquire data to be detected, where the data to be detected includes data of a first user and data of a second user associated with the first user;
[0035] A feature classification module, configured to classify the data to be detected according to a preset feature dimension to obtain classified feature data;
[0036] A sub-result acquisition module, configured to input the classified feature data into a corresponding preset decision tree to obtain classification sub-results of each preset decision tree;
[0037] A result determination module, configured to determine an anomaly detection result of the first user based on multiple classification sub-results, where the anomaly detection result is used to indicate whether there is an anomaly with the first user.
[0038] The present application has the following advantages:
[0039] The anomaly detection method and device provided by the present application acquire data to be detected, where the data to be detected includes data of a first user and data of a second user associated with the first user; classify the data to be detected according to a preset feature dimension to obtain classified feature data; input the classified feature data into a corresponding preset decision tree to obtain classification sub-results of each preset decision tree; determine an anomaly detection result of the first user based on multiple classification sub-results, so as to determine whether there is an anomaly with the first user according to the anomaly detection result. This method can detect abnormal accounts more comprehensively and accurately, thereby effectively reducing the occurrence of e-commerce data fraud behavior and ensuring the accuracy of e-commerce data. Brief Description of the Drawings
[0040] The drawings are used to provide a further understanding of the present application and form a part of the specification. Together with the following detailed implementation manners, they are used to explain the present application, but do not constitute a limitation to the present application.
[0041] Figure 1 It is a flowchart of an anomaly detection method provided by an embodiment of the present application;
[0042] Figure 2 It is a flowchart of an anomaly detection method provided by an embodiment of the present application;
[0043] Figure 3 It is a schematic diagram of the working process of an anomaly detection method provided by an embodiment of the present application;
[0044] Figure 4 It is a block diagram of an anomaly detection device provided by an embodiment of the present application;
[0045] Figure 5 It is a schematic diagram of an anomaly detection system provided by an embodiment of the present application;
[0046] Figure 6 It is a block diagram of an electronic device provided by an embodiment of the present application. Detailed Implementation Manners
[0047] The following details the specific implementation manners of the present application with reference to the drawings. It should be understood that the specific implementation manners described herein are only for explaining and understanding the present application, and are not used to limit the present application.
[0048] As used in the present application, the term "and / or" includes any and all combinations of one or more related listed items.
[0049] The terms used in the present application are only for describing specific embodiments and are not intended to limit the present application. As used in the present application, the singular forms "a" and "the" are also intended to include the plural forms unless the context clearly indicates otherwise.
[0050] When the terms "include" and / or "consist of" are used in the present application, it is specified that the stated features, wholes, steps, operations, elements, and / or components exist, but do not exclude the existence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their groups.
[0051] Unless otherwise defined, all terms (including technical and scientific terms) used in this application shall have the same meaning as commonly understood by those of ordinary skill in the art. It will also be understood that terms such as those defined in common dictionaries shall be interpreted as having a meaning consistent with their meaning in the relevant art and the context of this application, and shall not be interpreted as having an idealized or overly formal meaning, unless explicitly defined as such in this application.
[0052] In a first aspect, an embodiment of this application provides an anomaly detection method.
[0053] Figure 1 FIG. is a flowchart of an anomaly detection method provided by an embodiment of this application. This anomaly detection method can be applied to an anomaly detection server. As Figure 1 shown, this anomaly detection method includes the following steps:
[0054] Step S101, obtain data to be detected.
[0055] Among them, the data to be detected is data related to the user to be detected. The user to be detected is a user of an e-commerce platform, including both seller users and buyer users. In other words, the anomaly detection method provided by the embodiment of this application can detect both seller users with anomalies and buyer users with anomalies.
[0056] Under normal circumstances, whether detecting seller users or buyer users, both parties may be involved. Therefore, the data to be detected can include data of both parties.
[0057] In some possible implementation manners, the data to be detected includes data of a first user and data of a second user associated with the first user. Among them, the first user is the user to be detected.
[0058] In one example, the first user is a seller user, and the second user is a seller user having a transaction behavior with the first user; the first user is a buyer user, and the second user is a seller user having a transaction behavior with the first user.
[0059] In some possible implementation manners, the data of the seller user includes account basic information, commodity information, sales information, etc.
[0060] In some possible implementation manners, the data of the buyer user includes account basic information, commodity browsing information, commodity purchase information, etc.
[0061] It should be noted that the above is only an example of the data to be detected, and the content of the data to be detected can be determined according to actual detection requirements or empirical data, etc. The embodiment of this application does not limit this.
[0062] It should also be noted that the data to be detected can be obtained from information sources such as the database of the e-commerce platform and the database of the user to be detected. The embodiments of the present application do not limit the source of the data to be detected.
[0063] Step S102: Classify the data to be detected according to a preset feature dimension to obtain classified feature data.
[0064] Among them, the preset feature dimension is a feature dimension used to classify the data to be detected. For different types of users or different application scenarios, the preset feature dimension can be different to obtain good classification results, thereby improving the accuracy and rationality of anomaly detection.
[0065] It should be noted that in some possible implementation manners, before classifying the data to be detected, it further includes: preprocessing the data to be detected, where the preprocessing includes operations such as removing outliers and invalid values, filling in missing values, unifying the value units, and dimension conversion.
[0066] Step S103: Input the classified feature data into the corresponding preset decision tree to obtain the classification sub-results of each preset decision tree.
[0067] Among them, the preset decision tree is a tree structure composed of multiple nodes, which grows from top to bottom, and through this tree structure, the hidden rules in the data can be presented more intuitively.
[0068] In some possible implementation manners, the prediction decision tree includes a root node, intermediate nodes, and leaf nodes. Among them, the root node represents the input node of the classified feature data, all non-leaf nodes are used for conditional judgment to determine the classification features corresponding to the corresponding input data, and the leaf nodes represent the final classification sub-results of the preset decision tree.
[0069] In some possible implementation manners, there is a one-to-one correspondence between the preset feature dimension, the classified feature data, and the preset decision tree. Therefore, a certain classified feature data should be input into the corresponding preset decision tree to obtain accurate classification sub-results.
[0070] In some possible implementation manners, the preset decision tree can be constructed immediately or pre-constructed. Among them, pre-constructing the preset decision tree can be implemented through a random forest model.
[0071] In one example, the preset decision tree is a decision tree obtained by training a random forest model, which is a model based on a random forest and used to generate a preset decision tree for anomaly detection. Among them, a random forest is an ensemble algorithm composed of decision trees and belongs to the Bagging (Bootstrap Aggregation) method in ensemble learning. A random forest is composed of many decision trees, and there is no association between different decision trees. When dealing with a classification task, input the data to be processed, and each decision tree in the random forest model makes judgments and classifications respectively. Each decision tree will obtain a classification result. Whichever classification has the most occurrences in the classification results of the decision trees, the random forest model takes this classification result as the final result.
[0072] In some possible implementation manners, the method for obtaining the random forest model includes: First, determine multiple preset feature dimensions; Second, classify the training data according to the preset feature dimensions to obtain classified feature training data; Third, calculate the Gini coefficient of each classified feature training data, and determine each level of classification nodes based on the Gini coefficient; Finally, generate corresponding preset decision trees based on each level of classification nodes to obtain the random forest model.
[0073] In some possible implementation manners, calculating the Gini coefficient of each classified feature training data and determining each level of classification nodes based on the Gini coefficient includes:
[0074] Calculate the Gini coefficient of each classified feature training data and determine the minimum Gini coefficient of the first level;
[0075] Take the classified feature corresponding to the minimum Gini coefficient of the first level as the first-level classification node, and divide the classified feature training data into multiple first-level branch training data according to the first-level classification node;
[0076] For each (n - 1)-th level branch training data, calculate the Gini coefficient of the (n - 1)-th level branch training data, determine the minimum Gini coefficient of the n-th level, take the classified feature corresponding to the minimum Gini coefficient of the n-th level as the n-th level classification node, and divide the (n - 1)-th level branch training data into multiple n-th level branch training data, where n is an integer greater than or equal to 2;
[0077] When the preset stop condition is met, stop the data division operation to obtain each level of classification nodes.
[0078] Among them, the Gini coefficient is an index for measuring the degree of difference or purity of the dataset samples. Generally, the larger the Gini coefficient, the more diverse the types of the dataset, and thus the more classification results are obtained and the lower the purity. On the contrary, the smaller the Gini coefficient, the fewer the types of the dataset, and thus the fewer classification results are obtained and the higher the purity.
[0079] The preset stop conditions include a purity threshold, a branching times threshold, etc. In other words, during the above process, the decision tree keeps growing through branching until there are no more features to classify, or the overall purity index reaches the optimum, or the branching times reach the preset threshold, etc. In such cases, all classification nodes are determined and the decision tree stops growing.
[0080] It should be noted that to solve the problem of overfitting, pruning operations can also be performed on the preset decision tree. For example, pruning the obtained preset decision tree based on error reduction pruning method, pessimistic pruning method, cost complexity pruning method, etc.
[0081] Step S104, determine the anomaly detection result of the first user according to multiple classification sub-results, and the anomaly detection result is used to represent whether the first user has an anomaly.
[0082] In some possible implementation manners, when the percentage of a certain classification sub-result in all classification sub-results exceeds a preset percentage threshold, determine the classification sub-result as the final anomaly detection result of the first user.
[0083] For example, the number of classification sub-results is 10, and the preset percentage threshold is 60%. When 7 of the classification sub-results are no anomaly and 3 are with anomaly, the percentage of no anomaly in all classification sub-results is 70%, which is greater than the preset percentage threshold of 60%. Therefore, it is determined that the anomaly detection result of the first user is no anomaly.
[0084] In some possible implementation manners, first, determine the weight values of each preset decision tree; second, obtain the anomaly detection result of the first user according to the weight values of each preset decision tree and the classification sub-results. Among them, the weight value of the preset decision tree can represent the significance degree of the corresponding classification feature data in reflecting whether the first user is abnormal.
[0085] It should be noted that after determining that the first user has an anomaly, other users with a close association relationship with the first user can also be found as users with a greater anomaly risk, and anomaly detection can be carried out on these users.
[0086] It should also be noted that in some possible implementation manners, the e-commerce platform can store information such as operation data in the blockchain network. When performing anomaly detection, relevant data is retrieved from the blockchain network, and the obtained anomaly detection result is also stored in the blockchain network, so as to ensure the confidentiality, security, and traceability of the data.
[0087] In an embodiment of the present application, obtain data to be detected, where the data to be detected includes data of a first user and data of a second user associated with the first user; classify the data to be detected according to a preset feature dimension to obtain classified feature data; input the classified feature data into a corresponding preset decision tree to obtain classification sub-results of each preset decision tree; determine an anomaly detection result of the first user according to multiple classification sub-results, so as to determine whether there is an anomaly for the first user according to the anomaly detection result. This method can detect abnormal accounts more comprehensively and accurately, thereby effectively reducing the occurrence of e-commerce data fraud behavior and ensuring the accuracy of e-commerce data.
[0088] Figure 2 FIG. is a flowchart of an anomaly detection method provided by an embodiment of the present application, and this anomaly detection method can be applied to an anomaly detection server. As Figure 2 shown, this anomaly detection method includes the following steps:
[0089] Step S201, obtain data to be detected.
[0090] Step S202, classify the data to be detected according to a preset feature dimension to obtain classified feature data.
[0091] Step S203, input the classified feature data into a corresponding preset decision tree to obtain classification sub-results of each preset decision tree.
[0092] Step S204, determine an anomaly detection result of the first user according to multiple classification sub-results.
[0093] The content of steps S201 to S204 in this embodiment is the same as that of steps S101 to S104 in the previous embodiment of the present application, and will not be elaborated here.
[0094] Step S205, based on a review terminal, review the anomaly detection result to obtain a review result.
[0095] After obtaining the anomaly detection result of the first user, to ensure the accuracy of this anomaly detection result and avoid misjudgment and other situations, the obtained anomaly detection result can be reviewed based on a review terminal, and a review result can be obtained. When the review result is consistent with the anomaly detection result, it indicates that the anomaly detection result of the first user is accurate; when the review result is inconsistent with the anomaly detection result, it indicates that the anomaly detection result of the first user is inaccurate, and a new anomaly detection can be initiated or a new round of review can be carried out.
[0096] In some possible implementation manners, the review terminal is a terminal corresponding to a review operator.
[0097] Step S206, when the anomaly detection result indicates that there is an anomaly for the first user, obtain the transaction data of the first user within a preset transaction time period as supplementary detection data.
[0098] Step S207, obtain the supplementary detection result of the first user according to the supplementary detection data and the preset decision tree.
[0099] When performing anomaly detection, to reduce the data processing intensity and complexity and quickly obtain the anomaly detection result, usually only the data of the first user within a relatively short time period and the data of the second user associated with the first user are obtained to initiate the anomaly detection. When the anomaly detection result indicates that there is an anomaly for the first user, supplementary detection can be performed to improve the detection accuracy. Moreover, when performing supplementary detection, the preset transaction time period is relatively longer than the time period for obtaining the data to be detected, so as to obtain more comprehensive supplementary detection data.
[0100] Moreover, when the supplementary detection result is consistent with the anomaly detection result, it indicates that the anomaly detection result for the first user is accurate. When the supplementary detection result is inconsistent with the anomaly detection result, it indicates that the anomaly detection result for the first user is inaccurate, and a new anomaly detection can be initiated or a new round of supplementary detection can be performed.
[0101] Step S208, when the number of anomalies detected for the first user within a preset time period exceeds the preset number threshold, process the first user according to the preset penalty rule.
[0102] Figure 3 It is a schematic diagram of the working process of an anomaly detection method provided by an embodiment of the present application. This anomaly detection method can be applied to an anomaly detection server. As Figure 3 shown, this anomaly detection method includes the following steps:
[0103] Step S301, train a random forest model.
[0104] Taking the first user as the buyer user as an example for illustration. First, determine the account basic information, pre-order behavior information, and post-order behavior information as the preset feature dimensions. After classifying the training data according to the above three dimensions, obtain the classified feature training data, calculate the Gini coefficient of each classified feature training data, determine each level of classification node based on the Gini coefficient, and finally generate a preset decision tree according to each level of classification node.
[0105] In some possible implementation manners, the preset decision tree is in the form of a binary tree, then for each classification node, there are only two classification results.
[0106] Assume there are k categories, then the Gini coefficient is defined as:
[0107] Among them, p i represents the probability that the sample point belongs to the i-th category.
[0108] For the data set D,
[0109] Among them, Gini(D|A) represents the Gini coefficient corresponding to classifying the data set D using the feature A. D1 and D2 are two subsets formed by dividing D using the feature A. Gini(D1) represents the Gini coefficient corresponding to the subset D1, and Gini(D2) represents the Gini coefficient corresponding to the subset D2.
[0110] In some possible implementation manners, the classification feature with the smallest Gini coefficient and its value are selected as the optimal feature and the optimal splitting point (i.e., the classification node), and the classification feature training data is divided into two parts. The left and right child nodes repeatedly select the optimal feature and the splitting point. When there is no feature, it ends, and a corresponding preset decision tree is constructed based on all the classification nodes.
[0111] In some possible implementation manners, after obtaining the preset decision tree, corresponding weight values can also be set for each preset decision tree. When determining the anomaly detection result, a weighted sum is performed according to the weight value and the corresponding classification sub-result. When the proportion of determining the first user as an abnormal user based on the weighted sum is higher than that of determining the first user as a non-abnormal user, it is determined that the first user has an anomaly; otherwise, it is determined that the first user is a normal user.
[0112] Step S302, obtain the data to be detected.
[0113] In some possible implementation manners, the data to be detected includes the data of the first user and the data of the second user associated with the first user. The first user can be a buyer user or a seller user.
[0114] Step S303, classify the data to be detected according to the preset feature dimension to obtain classification feature data.
[0115] In some possible implementation manners, the first user is a buyer user, and the data of the first user is classified according to the account basic information, the behavior information before placing an order, and the behavior information after placing an order to obtain the corresponding classification feature data.
[0116] For example, the account basic information includes the account registration duration, the shopping records within three months of the account, whether the account has ever been determined as an abnormal account, etc.; the behavior information before placing an order includes the shopping entrance type (product link, store link, product keyword search, store name search, favorite store, etc.), the browsing time of the purchased product, the browsing time of similar products, etc.; the behavior information after placing an order includes the product delivery duration, the evaluation record (good review, bad review), whether the evaluation includes a returned picture, whether there is a return or exchange, etc.
[0117] It should be noted that for the above three types of classification feature data, the number of corresponding preset decision trees is also three, and there is a one-to-one correspondence between them. That is, the basic account information is input into the corresponding preset decision tree, the behavior information before placing an order is input into the corresponding preset decision tree, and the behavior information after placing an order is input into the corresponding preset decision tree.
[0118] Step S304: Input the classification feature data into the random forest model to obtain the anomaly detection result of the first user.
[0119] It should be noted that after the classification feature data is input into the random forest model, different classification feature data are processed by different preset decision trees, and the classification sub-results of each preset decision tree are obtained. Finally, based on multiple classification sub-results, the anomaly detection result of the first user is determined.
[0120] Step S305: Visualize the obtained anomaly detection result.
[0121] After obtaining the anomaly detection results of multiple users, a visualization method can be used to display the anomaly detection results of multiple users and related data, so as to facilitate more intuitive analysis of existing problems and corresponding reasons.
[0122] The step division of the above various methods is only for clear description. When implemented, they can be combined into one step or some steps can be split into multiple steps. As long as the same logical relationship is included, it is within the protection scope of this patent; adding insignificant modifications to the algorithm or process or introducing insignificant designs, but not changing the core design of its algorithm and process, are all within the protection scope of this patent.
[0123] In a second aspect, an embodiment of the present application provides an anomaly detection device.
[0124] Figure 4 It is a block diagram of an anomaly detection device provided by an embodiment of the present application. This anomaly detection device can be integrated in an anomaly detection server. As Figure 4 shown, this anomaly detection device includes the following modules:
[0125] A data acquisition module 401, configured to acquire data to be detected.
[0126] Among them, the data to be detected includes the data of the first user and the data of the second user associated with the first user.
[0127] A feature classification module 402, configured to classify the data to be detected according to a preset feature dimension to obtain classification feature data.
[0128] The sub-result acquisition module 403 is configured to input the classified feature data into corresponding preset decision trees to obtain classification sub-results of each preset decision tree.
[0129] The result determination module 404 is configured to determine the anomaly detection result of the first user according to multiple classification sub-results.
[0130] Wherein, the anomaly detection result is used to characterize whether there is an anomaly for the first user.
[0131] In some possible implementation manners, the first user is a seller user, and the second user is a seller user having a transaction behavior with the first user; the first user is a buyer user, and the second user is a seller user having a transaction behavior with the first user.
[0132] In some possible implementation manners, the preset decision tree is a decision tree obtained by training a random forest model, and the random forest model is a model established based on a random forest and used to generate a preset decision tree for anomaly detection.
[0133] In some possible implementation manners, the anomaly detection device further includes: a model acquisition module. The model acquisition module includes: a feature dimension determination unit, a training classification unit, a node determination unit, and a model construction unit.
[0134] Wherein, the feature dimension determination unit is configured to determine multiple preset feature dimensions; the training classification unit is configured to classify the training data according to the preset feature dimensions to obtain classified feature training data; the node determination unit is configured to calculate the Gini coefficient of each classified feature training data and determine classification nodes at all levels based on the Gini coefficient; the model construction unit is configured to generate corresponding preset decision trees based on the classification nodes at all levels to obtain a random forest model.
[0135] In some possible implementation manners, the node determination unit is configured to: calculate the Gini coefficient of each classified feature training data, and determine the minimum Gini coefficient at the first level; use the classified feature corresponding to the minimum Gini coefficient at the first level as the classification node at the first level, and divide the classified feature training data into multiple first-level branch training data according to the classification node at the first level; for each (n - 1)-level branch training data, calculate the Gini coefficient of the (n - 1)-level branch training data, determine the minimum Gini coefficient at the nth level, use the classified feature corresponding to the minimum Gini coefficient at the nth level as the classification node at the nth level, and divide the (n - 1)-level branch training data into multiple n-level branch training data, where n is an integer greater than or equal to 2; when a preset stop condition is satisfied, stop the data division operation to obtain classification nodes at all levels.
[0136] In some possible implementation manners, the result determination module includes: a weight determination unit and a result determination unit. Among them, the weight determination unit is configured to determine the weight values of each preset decision tree; the result determination unit is configured to obtain the anomaly detection result of the first user according to the weight values of each preset decision tree and the classification sub-results.
[0137] In some possible implementation manners, the anomaly detection device further includes: a review module, configured to review the anomaly detection result based on a review terminal to obtain a review result.
[0138] In some possible implementation manners, the anomaly detection device further includes: a supplementary detection module. The supplementary detection module includes: a supplementary data acquisition unit, configured to, when the anomaly detection result indicates that the first user has an anomaly, acquire the transaction data of the first user within a preset transaction time period as supplementary detection data; a supplementary detection unit, configured to obtain the supplementary detection result of the first user according to the supplementary detection data and the preset decision tree.
[0139] In some possible implementation manners, the anomaly detection device further includes: a processing module, configured to, when the number of times the first user is detected to have an anomaly within a preset time period exceeds a preset number threshold, process the first user according to a preset penalty rule.
[0140] In a third aspect, an embodiment of the present application provides an anomaly detection system.
[0141] Figure 5 is a schematic diagram of an anomaly detection system provided by an embodiment of the present application. As Figure 5 shown, the anomaly detection system includes: e-commerce platforms 511, e-commerce platforms 512,..., e-commerce platforms 51N, a blockchain network 520, an anomaly detection server 530, and a visualization display platform 540, where N is an integer greater than or equal to 1.
[0142] Among them, each e-commerce platform is connected to the blockchain network 520 and transmits information such as operation data into the blockchain network 520 for storage in the form of blocks. The anomaly detection server 530 is built with an anomaly detection module, configured to initiate anomaly detection by using the anomaly detection method of any embodiment of the present application, and transmit the obtained anomaly detection result into the blockchain network 520 for storage. The visualization display platform 540 is connected to the blockchain network 520, and can at least visually display the user data and the anomaly detection result in the blockchain network. System users can perform operations such as viewing and querying information through the visualization display platform.
[0143] The functions or modules included in the device provided by the embodiments of the present application can be used to execute the methods described in the method embodiments of the first aspect above. For the specific implementation and technical effects, reference can be made to the descriptions in the method embodiments above. For the sake of brevity, they will not be repeated here.
[0144] It should be noted that each module involved in this embodiment is a logical module. In actual applications, a logical unit can be a physical unit, a part of a physical unit, or a combination of multiple physical units. In addition, to highlight the innovative part of the present application, units not closely related to solving the technical problems proposed in the present application are not introduced in this embodiment, but this does not mean that there are no other units in this embodiment.
[0145] Refer to Figure 6 , the embodiments of the present application provide an electronic device, which includes:
[0146] One or more processors 601;
[0147] A memory 602, on which one or more programs are stored. When the one or more programs are executed by the one or more processors, the one or more processors implement the anomaly detection method of any one of the above;
[0148] One or more I / O interfaces 603, connected between the processor and the memory, configured to implement information interaction between the processor and the memory.
[0149] Among them, the processor 601 is a device with data processing capabilities, including but not limited to a central processing unit (CPU), etc.; the memory 602 is a device with data storage capabilities, including but not limited to a random access memory (RAM, more specifically such as SDRAM, DDR, etc.), a read-only memory (ROM), an electrically erasable programmable read-only memory (EEPROM), a flash memory (FLASH); the I / O interface (read / write interface) 603 is connected between the processor 601 and the memory 602 and can implement information interaction between the processor 601 and the memory 602, including but not limited to a data bus (Bus), etc.
[0150] In some embodiments, the processor 601, the memory 602, and the I / O interface 603 are interconnected through a bus and then connected to other components of the computing device.
[0151] This embodiment also provides a computer-readable medium, on which a computer program is stored. When the program is executed by a processor, it implements the anomaly detection method provided in this embodiment. To avoid repeated description, the specific steps of the anomaly detection method will not be elaborated here.
[0152] Those of ordinary skill in the art will understand that all or some of the steps in the above-invented method, and the functional modules / units in the system and device, can be implemented as software, firmware, hardware, and their appropriate combinations. In the hardware implementation, the division between the functional modules / units mentioned above does not necessarily correspond to the division of physical components; for example, a physical component can have multiple functions, or a function or step can be executed by several physical components in cooperation. Some or all of the physical components can be implemented as software executed by a processor, such as a central processing unit, a digital signal processor, or a microprocessor, or implemented as hardware, or implemented as an integrated circuit, such as an application-specific integrated circuit. Such software can be distributed on a computer-readable medium, which can include computer storage media (or non-transitory media) and communication media (or transitory media). As is well known to those of ordinary skill in the art, the term computer storage media includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information, such as computer-readable instructions, data structures, program modules, or other data. Computer storage media includes, but is not limited to, RAM, ROM, EEPROM, flash memory, or other memory technologies, CD-ROM, digital versatile discs (DVDs), or other optical disc storage, magnetic cassettes, tapes, magnetic disk storage, or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by a computer. In addition, as is well known to those of ordinary skill in the art, communication media typically contains computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transmission mechanism, and can include any information delivery medium.
[0153] It should be noted that, in this article, the term "comprising", "including", or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article, or device that includes a series of elements not only includes those elements but also includes other elements not explicitly listed, or further includes elements inherent to such a process, method, article, or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article, or device that includes the element.
[0154] Those skilled in the art can understand that although some of the embodiments described herein include certain features included in other embodiments but not others, the combination of features of different embodiments means that it is within the scope of this embodiment and forms different embodiments.
[0155] It is understandable that the above embodiments are merely exemplary embodiments adopted to illustrate the principles of the present application. However, the present application is not limited thereto. For those of ordinary skill in the art, various modifications and improvements can be made without departing from the spirit and essence of the present application, and these modifications and improvements are also regarded as the protection scope of the present application.
Claims
1. An anomaly detection method, characterized in that, The method includes: Obtaining data to be detected, where the data to be detected includes data of a first user and data of a second user associated with the first user. The first user is a seller user, and the second user is a seller user who has a transaction behavior with the first user, or the first user is a buyer user, and the second user is a seller user who has a transaction behavior with the first user; Classifying the data to be detected according to a preset feature dimension to obtain classified feature data; Inputting the classified feature data into corresponding preset decision trees to obtain classification sub-results of each of the preset decision trees; Determining an anomaly detection result of the first user according to multiple classification sub-results, where the anomaly detection result is used to indicate whether the first user has an anomaly; When the anomaly detection result indicates that the first user has an anomaly, searching for other users associated with the first user based on the first user and performing anomaly detection on the other users; Wherein, the preset decision tree is a decision tree obtained by training a random forest model, and the random forest model is a model established based on a random forest and used to generate the preset decision tree for anomaly detection; Wherein, the method further includes: Determining multiple preset feature dimensions; Classifying training data according to the preset feature dimension to obtain classified feature training data; Calculating the Gini coefficient of each classified feature training data, and determining classification nodes at all levels based on the Gini coefficient; Generating corresponding preset decision trees based on the classification nodes at all levels to obtain the random forest model; Wherein, the calculating the Gini coefficient of each classified feature training data and determining classification nodes at all levels based on the Gini coefficient includes: Calculating the Gini coefficient of each classified feature training data and determining the minimum Gini coefficient at the first level; Taking the classified feature corresponding to the minimum Gini coefficient at the first level as the classification node at the first level, and dividing the classified feature training data into multiple first-level branch training data according to the classification node at the first level; For each (n - 1)-th level branch training data, calculating the Gini coefficient of the (n - 1)-th level branch training data, determining the minimum Gini coefficient at the n-th level, taking the classified feature corresponding to the minimum Gini coefficient at the n-th level as the classification node at the n-th level, and dividing the (n - 1)-th level branch training data into multiple n-th level branch training data, where n is an integer greater than or equal to 2; When a preset stop condition is satisfied, stopping the data division operation to obtain classification nodes at all levels.
2. The anomaly detection method according to claim 1, characterized in that The determining the anomaly detection result of the first user according to multiple classification sub-results includes: Determining the weight value of each preset decision tree; Obtaining the anomaly detection result of the first user according to the weight value of each preset decision tree and the classification sub-result.
3. The anomaly detection method according to claim 1 or 2, characterized in that After determining the anomaly detection result of the first user according to multiple classification sub-results, it further includes: Rechecking the anomaly detection result based on a rechecking terminal to obtain a rechecking result.
4. The anomaly detection method according to claim 1, wherein The anomaly detection result indicates that the first user has an anomaly; After determining the anomaly detection result according to the multiple classification sub-results, the method further includes: Obtaining the transaction data of the first user within a preset transaction time period as supplementary detection data; Obtaining a supplementary detection result of the first user according to the supplementary detection data and the preset decision tree.
5. The anomaly detection method according to claim 1, wherein After determining the anomaly detection result of the first user according to the multiple classification sub-results, the method further includes: In the case where the number of times the first user is detected to have an anomaly within a preset time period exceeds a preset number threshold, processing the first user according to a preset penalty rule.
6. An anomaly detection device, characterized in that, The apparatus includes: A data acquisition module, configured to acquire data to be detected, where the data to be detected includes data of a first user and data of a second user associated with the first user; A feature classification module, configured to classify the data to be detected according to a preset feature dimension to obtain classified feature data; A sub-result acquisition module, configured to input the classified feature data into a corresponding preset decision tree to obtain classification sub-results of each of the preset decision trees; A result determination module, configured to determine an anomaly detection result of the first user according to the multiple classification sub-results, where the anomaly detection result is used to indicate whether the first user has an anomaly; The anomaly detection apparatus is further configured to, in the case where the anomaly detection result indicates that the first user has an anomaly, search for other users associated with the first user based on the first user and perform anomaly detection on the other users; Wherein, the anomaly detection apparatus is configured to execute the anomaly detection method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Abnormal user identification method and device, electronic equipment and storage medium
CN111641608A
Transaction platform scalping detection system
CN112598431A