Abnormity recognition method and device, electronic equipment and computer readable medium
By building frequent pattern trees in business scenarios and generating preset frequent item sets, the problem of impossible to accurately identify business risks in the prior art is solved, and efficient and accurate abnormal identification is achieved.
Patent Information
- Application Number
- CN202510044898.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-10
- Publication Date
- 2025-05-09
AI Technical Summary
The prior art cannot accurately and timely identify and determine the risks behind the business in standardized business scenarios, resulting in low efficiency and low accuracy of abnormal identification.
By obtaining business scenario identification, obtaining fingerprint item data, and obtaining historical exception data in real time, determining the support for exception items, building a frequent pattern tree, and generating a preset frequent item set. When the attribute data of the preset frequent item set matches the fingerprint item data, an exception identifier is marked and the relevant data is output.
It realizes accurate and timely identification of risks behind the business, and improves the efficiency, accuracy and timeliness of abnormal identification.
Smart Images

Figure CN119961832A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to an abnormality identification method, device, electronic device and computer-readable medium. Background Art
[0002] At present, in the standardized business scenario domain, by analyzing the specific behaviors of entities such as users, merchants, and delivery, we can extract abnormal behaviors that do not match entities and businesses, and then align them through risk control strategies for control and constraints. Most of the customization of manual risk control strategies uses the interests in business scenarios as cheating motivations, or analyzes the laws of abnormal data feedback from customer complaints, and then gradually determines the formed risk control strategies. It is impossible to accurately and timely identify and determine the risks behind the business, and the efficiency and accuracy of abnormal identification are low. Summary of the invention
[0003] In view of this, the embodiments of the present application provide an anomaly identification method, device, electronic device and computer-readable medium, which can solve the problem that the existing use of manual risk control strategies for anomaly identification cannot accurately and timely identify and determine the risks behind the business, and the anomaly identification efficiency and accuracy are low.
[0004] To achieve the above object, according to one aspect of an embodiment of the present application, a method for identifying anomalies is provided, comprising:
[0005] Obtain a business scenario identifier according to the received exception identification request, and obtain corresponding fingerprint item data according to the business scenario identifier;
[0006] According to the business scenario identifier, the corresponding historical abnormal data is obtained in real time, the support of the abnormal items in the historical abnormal data is determined, a frequent pattern tree is constructed based on the abnormal items and the support, and a preset frequent item set is generated based on the frequent pattern tree;
[0007] When there is one or more attribute data of preset frequent item sets that completely matches the attribute data in the acquired fingerprint item data, an abnormal flag is marked for the fingerprint item data;
[0008] Output the fingerprint item data corresponding to all abnormal identifications and the business scenario identification corresponding to the fingerprint item data marked with abnormal identifications.
[0009] Optionally, a frequent pattern tree is constructed based on the outliers and the support, including:
[0010] Generate frequent item header table and sorted data set based on abnormal items, support and preset sorting method;
[0011] Build a frequent pattern tree based on the sorted dataset.
[0012] Optionally, generate preset frequent itemsets, including:
[0013] Generate a node list based on the frequent item header table and frequent pattern tree;
[0014] Based on the node linked list, determine the conditional pattern base of each frequent item in the frequent item header table;
[0015] Generate preset frequent item sets based on each conditional pattern base.
[0016] Optionally, based on each conditional pattern base, a preset frequent item set is generated, including:
[0017] Generate candidate frequent item sets based on each conditional pattern base;
[0018] Based on the preset item set length and the preset support threshold, the candidate frequent item sets are screened to obtain the preset frequent item sets.
[0019] Optionally, before determining the support of the abnormal item in the historical abnormal data, the method further includes:
[0020] The information entropy of each abnormal item in the historical abnormal data is calculated, and the abnormal items corresponding to the information entropy less than the preset information entropy threshold are eliminated from the historical abnormal data.
[0021] Optionally, before constructing the frequent pattern tree based on the abnormal items and the support, the method further includes:
[0022] The abnormal items corresponding to the support that is not within the preset support range are eliminated from the historical abnormal data.
[0023] In addition, the present application also provides an abnormality identification device, including:
[0024] An acquisition unit is configured to acquire a business scenario identifier according to the received abnormal identification request, and acquire corresponding fingerprint item data according to the business scenario identifier;
[0025] A frequent item set generation unit is configured to obtain corresponding historical abnormal data in real time according to the business scenario identifier, determine the support of abnormal items in the historical abnormal data, build a frequent pattern tree based on the abnormal items and the support, and generate a preset frequent item set based on the frequent pattern tree;
[0026] a marking unit configured to mark an abnormality mark for the fingerprint item data when there is one or more attribute data of a preset frequent item set that completely matches the attribute data in the acquired fingerprint item data;
[0027] The output unit is configured to output the fingerprint item data corresponding to all abnormal identifications and the business scenario identification corresponding to the fingerprint item data marked with the abnormal identification.
[0028] Optionally, the frequent itemset generation unit is further configured to:
[0029] Generate frequent item header table and sorted data set based on abnormal items, support and preset sorting method;
[0030] Build a frequent pattern tree based on the sorted dataset.
[0031] Optionally, the frequent itemset generation unit is further configured to:
[0032] Generate a node list based on the frequent item header table and frequent pattern tree;
[0033] Based on the node linked list, determine the conditional pattern base of each frequent item in the frequent item header table;
[0034] Generate preset frequent item sets based on each conditional pattern base.
[0035] Optionally, the frequent itemset generation unit is further configured to:
[0036] Generate candidate frequent item sets based on each conditional pattern base;
[0037] Based on the preset item set length and the preset support threshold, the candidate frequent item sets are screened to obtain the preset frequent item sets.
[0038] Optionally, the abnormality identification device further includes a data processing unit configured to:
[0039] The information entropy of each abnormal item in the historical abnormal data is calculated, and the abnormal items corresponding to the information entropy less than the preset information entropy threshold are eliminated from the historical abnormal data.
[0040] Optionally, the abnormality identification device further includes a data processing unit configured to:
[0041] The abnormal items corresponding to the support that is not within the preset support range are eliminated from the historical abnormal data.
[0042] In addition, the present application also provides an abnormality identification electronic device, including: one or more processors; a storage device for storing one or more programs, when the one or more programs are executed by one or more processors, the one or more processors implement the abnormality identification method as described above.
[0043] In addition, the present application also provides a computer-readable medium on which a computer program is stored, and when the program is executed by a processor, the above-mentioned abnormality identification method is implemented.
[0044] To achieve the above objective, according to another aspect of the embodiments of the present application, a computer program product is provided.
[0045] A computer program product according to an embodiment of the present application includes a computer program, and when the program is executed by a processor, the abnormality identification method provided by the embodiment of the present application is implemented.
[0046] One embodiment of the above invention has the following advantages or beneficial effects: the present application obtains the business scenario identifier according to the received abnormal identification request, obtains the corresponding fingerprint item data according to the business scenario identifier; obtains the corresponding historical abnormal data in real time according to the business scenario identifier, determines the support of the abnormal items in the historical abnormal data, builds a frequent pattern tree based on the abnormal items and the support, and generates a preset frequent item set based on the frequent pattern tree; when there is one or more preset frequent item sets whose attribute data completely matches the attribute data in the acquired fingerprint item data, marks the fingerprint item data with abnormal identifiers; outputs the fingerprint item data corresponding to all abnormal identifiers and the business scenario identifier corresponding to the fingerprint item data marked with abnormal identifiers. Accurately and timely identification and determination of the risks behind the business is achieved, and the efficiency, accuracy and timeliness of abnormal identification are improved.
[0047] The further effects of the above-mentioned non-conventional optional manner will be described below in conjunction with the specific implementation manner. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] The accompanying drawings are used to better understand the present application and do not constitute an improper limitation on the present application.
[0049] Figure 1 is a schematic diagram of the main process of an abnormality identification method according to an embodiment of the present application;
[0050] Figure 2 is a schematic diagram of the main process of an abnormality identification method according to an embodiment of the present application;
[0051] Figure 3 It is a schematic diagram of the main process of an abnormality identification method according to an embodiment of the present application;
[0052] Figure 4 is a schematic diagram of main units of an abnormality identification device according to an embodiment of the present application;
[0053] Figure 5 is an exemplary system architecture diagram to which the embodiments of the present application can be applied;
[0054] Figure 6 It is a structural diagram of a computer system of a terminal device or server suitable for implementing an embodiment of the present application. DETAILED DESCRIPTION
[0055] The following is an explanation of the exemplary embodiments of the present application in conjunction with the accompanying drawings, including various details of the embodiments of the present application to facilitate understanding, which should be considered as merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present application. Similarly, for the sake of clarity and conciseness, the description of known functions and structures is omitted in the following description. It should be noted that the acquisition, transmission, storage, use, and processing of data in the technical solution of the present application are in compliance with the relevant provisions of national laws and regulations. It should be noted that in the embodiments of the present application, certain software, components, models, and other existing solutions in the industry may be mentioned, which should be considered as exemplary, and their purpose is only to illustrate the feasibility of the implementation of the technical solution of the present application, but it does not mean that the applicant has or must have used the solution. In the technical solution of the present application, the collection, collection, update, analysis, processing, use, transmission, storage, etc. of the user personal information involved are in compliance with the provisions of relevant laws and regulations, are used for legal and reasonable purposes, and do not violate public order and good customs, are not shared, leaked or sold outside these legal uses, and are subject to supervision and management by regulatory authorities. Take necessary measures to prevent illegal access to such user personal information data, maintain the security of user personal information, network security and national security, ensure that persons who have access to personal information data comply with relevant laws and regulations, and ensure the security of user personal information. Once such user personal information data is no longer needed, the risk should be minimized by limiting or even prohibiting data collection and / or deleting data.
[0056] When used, including in certain related applications, user privacy is protected by de-identifying data, such as by removing specific identifiers when used, controlling the amount or specificity of data stored, controlling how data is stored, and / or other methods of de-identification.
[0057] Figure 1 is a schematic diagram of the main process of the abnormality identification method according to an embodiment of the present application, such as Figure 1 As shown, the abnormality identification method mainly includes the following steps S101 to S104.
[0058] Step S101, obtaining a business scenario identifier according to the received exception identification request, and obtaining corresponding fingerprint item data according to the business scenario identifier.
[0059] In this embodiment, the execution subject of the anomaly identification method (for example, it can be a server) can receive the anomaly identification request through a wired connection or a wireless connection. The anomaly identification method of the embodiment of the present application can be applied to risk control scenarios and user similarity judgment. Anomaly identification can be device anomaly identification or user behavior anomaly identification. The following example uses device anomaly identification as an example for illustration. User behavior anomaly identification is similar and will not be repeated. The execution subject can obtain a business scenario identifier when receiving the anomaly identification request. For example, the business scenario identifier is used to characterize the scenario where anomaly identification needs to be performed, such as device anomaly identification or user behavior anomaly identification or device and user behavior anomaly identification.
[0060] Device anomaly identification can be achieved through device fingerprint technology. Device fingerprint technology is based on the collection of device attribute information, through various communication encryption, client and server information coordination protection, asynchronous information collection and other technical means, to stably and reliably mark the device's identity id, which will be collectively referred to as EID in the future.
[0061] When the business scenario identifier corresponds to device abnormality identification, the corresponding fingerprint item data can be obtained according to the business scenario identifier, such as the device fingerprint data of the device item appearing in the order.
[0062] For example, the fingerprint item data may be the device fingerprint data of the device item appearing in the order. The device fingerprint data includes the collected device attributes. Specifically, the device attributes may include: device identification class, device attribute class (such as device model, device brand, etc.), device hardware (such as memory, resolution, hardware type, CPU, etc.), device network (such as client IP, network transmission speed, wifi name, wifi unique identifier, etc.), device security (such as whether it is Hooked, whether it is rooted, whether multiple applications are opened, etc.), device status (such as memory, charging, on a call, etc.), App information (such as App version), etc.
[0063] Device fingerprint technology is an important product in risk control and security attack and defense. Through device fingerprint technology, important device attributes and unique device IDs can be obtained, so as to accurately locate risks and problems in business scenarios. Device anomaly identification can be done within the standardized business scenario domain, by analyzing the specific behavior of devices of users, merchants, distribution entities, etc., extracting abnormal behaviors that do not match devices and businesses, and then controlling and restricting them.
[0064] Step S102, acquiring corresponding historical abnormal data in real time according to the business scenario identifier, determining the support of abnormal items in the historical abnormal data, constructing a frequent pattern tree based on the abnormal items and the support, and generating a preset frequent item set based on the frequent pattern tree.
[0065] When the execution subject obtains the business scenario identifier, it can quickly respond to obtain the corresponding historical abnormal data according to the business scenario identifier. The historical abnormal data can be the fingerprint item data of the device determined to be abnormal in the historical preset time period.
[0066] Determine the support of the abnormal items in the historical abnormal data, that is, the frequency of the abnormal items in the historical abnormal data, for example, the proportion of the number of times the abnormal items appear in the historical abnormal data to the total number of times the data items in the historical abnormal data appear. Specifically, the support of the abnormal item represents the frequency of the antecedent (i.e., attribute name) and the succesor (i.e., attribute value) of the abnormal item (attribute data) appearing simultaneously in a data set (i.e., historical abnormal data). A frequent pattern tree FP-tree is constructed based on the abnormal items and the support of the abnormal items, and a preset frequent item set is generated based on the frequent pattern tree FP-tree.
[0067] Specifically, before determining the support of abnormal items in the historical abnormal data, the abnormality identification method also includes: calculating the information entropy of each abnormal item in the historical abnormal data, and eliminating abnormal items corresponding to information entropy less than a preset information entropy threshold from the historical abnormal data.
[0068] In the embodiment of the present application, the smaller the information entropy, the higher the data purity, which means that the data is too high-frequency and has no dispersion, and the corresponding data will not contribute to risk identification. Such data is eliminated to improve the accuracy of abnormal identification.
[0069] Specifically, before constructing the frequent pattern tree based on the abnormal items and the support, the abnormal identification method further includes: removing abnormal items corresponding to the support that is not within the preset support range from the historical abnormal data.
[0070] A preset support range, such as [minimum support, maximum support], is used to remove abnormal items corresponding to supports that are not within the preset support range from historical abnormal data, that is, to retain support data with minimum support <= support data <= maximum support, so as to improve the accuracy of abnormality identification.
[0071] Step S103: When there is one or more preset frequent item sets whose attribute data completely matches the attribute data in the acquired fingerprint item data, an abnormal flag is marked for the fingerprint item data.
[0072] Specifically, when there is one or more preset frequent item sets whose attribute data completely matches the attribute data in the acquired fingerprint item data, an abnormal identification is marked for the fingerprint item data, including: extracting the attribute data in the fingerprint item data, and if the preset frequent item sets completely hit the attribute data, an abnormal identification is marked for the fingerprint item data.
[0073] For example, in an embodiment of the present application, the support of an abnormal item indicates the frequency of the antecedent (i.e., attribute name) and the subsequent (i.e., attribute value) of the abnormal item (attribute data) appearing simultaneously in a data set (i.e., historical abnormal data). In an embodiment of the present application, the attribute data includes an attribute name composed of antecedent (i.e., attribute name) and subsequent (i.e., attribute value): attribute value. The preset frequent item set may include attribute data and the frequency of occurrence of attribute data, wherein the frequency of occurrence of attribute data in the preset frequent item set may be obtained by multiplying the support corresponding to the attribute data (i.e., the frequency of occurrence of the abnormal item corresponding to the attribute data in the historical abnormal data) by the total number of occurrences of the data item in the historical abnormal data. For example, the preset frequent item set may include: attribute name: attribute value and attribute name: attribute value frequency of occurrence. The number of preset frequent item sets may be multiple, and when there are one or more preset frequent item sets of attribute data (e.g., attribute name: attribute value) that completely matches the attribute data (e.g., attribute name: attribute value) in the acquired fingerprint item data, the fingerprint item data is marked with an abnormal identifier. For example, when there are one or more preset frequent item sets whose attribute name:attribute value completely matches the attribute name:attribute value in the acquired fingerprint item data, an abnormal flag is marked for the fingerprint item data.
[0074] Examples of frequent itemsets in this application are as follows:
[0075] ['null', {'brand:redmin':8}, {'os:android 4.2':4}, {'changing:1':3}, {'channel:jd':2}, {'bissid:xxxxx':2}]
[0076] ['null', {'brand:redmin':8}, {'os:android 4.2':4}, {'changing:1':3}, {'timezose:us':1}]
[0077] ['null', {'brand:redmin':8}, {'os:android 4.2':4}, {'dip:1999':1}, {'nfc:1':1}
[0078] Among them, null is the root node of the tree and has no practical meaning. The key in the dictionary is "attribute name: attribute value", such as brand: redmin in the above example, and the value of the dictionary is the frequency of the corresponding "attribute name: attribute value", such as 8 after 'brand: redmin' in the above example. When there is a preset frequent item set (for example, ['null', {'brand: redmin': 8}, {'os: android 4.2': 4}, {'changing: 1': 3}, {'channel: jd': 2}, {'bissid: xxxxx': 2}]) whose attribute name: attribute value (for example, brand: redmin, os: android4.2, changing: 1, channel: jd, bissid: xxxxx) completely matches the attribute name: attribute value (for example, brand: redmin, os: android 4.2, changing: 1, channel: jd, bissid: xxxxx) in the acquired fingerprint item data, the fingerprint item data is marked with an abnormal flag. When there are multiple preset frequent item sets whose attribute data (eg, attribute name: attribute value) completely matches the attribute data (eg, attribute name: attribute value) in the acquired fingerprint item data, an abnormal flag is marked for the fingerprint item data.
[0079] Taking device anomaly identification as an example, for example, the preset frequent item set may include: device attribute name: device attribute value and device attribute name: frequency of occurrence of device attribute value. The number of preset frequent item sets may be multiple. When the device attribute data of one or more preset frequent item sets completely matches the attribute data in the acquired fingerprint item data, the fingerprint item data is marked with an abnormal identifier.
[0080] Step S104: outputting fingerprint item data corresponding to all abnormal identifications and business scenario identifications corresponding to the fingerprint item data marked with abnormal identifications.
[0081] The execution subject can obtain multiple different business scenario identifiers according to the received exception identification request. Different business scenario identifiers (i.e., representing different business scenarios) correspond to different fingerprint item data. Therefore, the output is the fingerprint item data corresponding to all exception identifiers and the business scenario identifier corresponding to the fingerprint item data marked with the exception identifier. There can also be multiple fingerprint item data marked with exception identifiers. All fingerprint item data marked with exception identifiers and the business scenario identifiers corresponding to these fingerprint item data are output to allow the processing node to more accurately perform exception processing and risk warning.
[0082] This embodiment obtains the business scenario identifier according to the received exception identification request, obtains the corresponding fingerprint item data according to the business scenario identifier; obtains the corresponding historical exception data in real time according to the business scenario identifier, determines the support of the exception items in the historical exception data, builds a frequent pattern tree based on the exception items and the support, and generates a preset frequent item set based on the frequent pattern tree; when there is one or more preset frequent item sets whose attribute data completely matches the attribute data in the acquired fingerprint item data, marks the fingerprint item data with an exception identifier; outputs the fingerprint item data corresponding to all the exception identifiers and the business scenario identifier corresponding to the fingerprint item data marked with the exception identifier. Accurately and timely identification and determination of the risks behind the business is achieved, and the efficiency, accuracy and timeliness of exception identification are improved.
[0083] Figure 2 FIG. 1 is a schematic diagram of the main flow of an abnormality identification method according to an embodiment of the present application. Figure 2 As shown, the abnormality identification method mainly includes the following steps S201-S207.
[0084] Step S201: obtaining a business scenario identifier according to the received exception identification request, and obtaining corresponding fingerprint item data according to the business scenario identifier.
[0085] The business scenario identifier is used to characterize the scenario where anomaly identification is required, such as device anomaly identification or user behavior anomaly identification or device and user behavior anomaly identification.
[0086] Device anomaly identification can be achieved through device fingerprint technology. Device fingerprint technology is based on the collection of device attribute information, through various communication encryption, client and server information coordination protection, asynchronous information collection and other technical means, to stably and reliably mark the device's identity id, which will be collectively referred to as EID in the future.
[0087] When the business scenario identifier corresponds to device abnormality identification, the corresponding fingerprint item data can be obtained according to the business scenario identifier, such as the device fingerprint data of the device item appearing in the order.
[0088] For example, the fingerprint item data may be the device fingerprint data of the device item appearing in the order. The device fingerprint data includes the collected device attributes. Specifically, the device attributes may include: device identification class, device attribute class (such as device model, device brand, etc.), device hardware (such as memory, resolution, hardware type, CPU, etc.), device network (such as client IP, network transmission speed, wifi name, wifi unique identifier, etc.), device security (such as whether it is Hooked, whether it is rooted, whether multiple applications are opened, etc.), device status (such as memory, charging, on a call, etc.), App information (such as App version), etc.
[0089] Step S202: acquiring corresponding historical abnormal data in real time according to the business scenario identifier, and determining the support degree of abnormal items in the historical abnormal data.
[0090] Historical abnormal data may be fingerprint item data of a device determined to be abnormal that appeared in a preset historical time period. Determine the support of the abnormal item in the historical abnormal data, that is, the frequency of the abnormal item appearing in the historical abnormal data, for example, the proportion of the number of times the abnormal item appears in the historical abnormal data to the total number of times the data item appears in the historical abnormal data. Specifically, the support of the abnormal item represents the frequency with which the antecedent (i.e., attribute name) and the succesor (i.e., attribute value) of the abnormal item (attribute data) appear simultaneously in a data set (i.e., historical abnormal data). In an embodiment of the present application, the attribute data includes an attribute name consisting of an antecedent (i.e., attribute name) and a succesor (i.e., attribute value): attribute value. The preset frequent item set may include attribute data and the frequency of occurrence of the attribute data, wherein the frequency of occurrence of the attribute data in the preset frequent item set may be obtained by multiplying the support corresponding to the attribute data (i.e., the frequency of occurrence of the abnormal item corresponding to the attribute data in the historical abnormal data) by the total number of occurrences of the data item in the historical abnormal data. For example, the attribute data in the abnormal item in the historical abnormal data may be as follows:
[0091] ABCEFO
[0092] ACG
[0093] EI
[0094] ACDEG
[0095] ACEGL
[0096] EJ
[0097] ABCEFP
[0098] ACD
[0099] ACEGM
[0100] ACEGN
[0101] Assuming that the minimum support = 0.2, after deleting the attribute data O, I, L, J, P, M, N in the abnormal items with less than the minimum support, the frequency of occurrence of the attribute data in the abnormal items in the remaining historical abnormal data is: A: 8, C: 8, E: 8, G: 5, B: 2, D: 2, F: 2.
[0102] Step S203: Generate a frequent item header table and a sorted data set based on the abnormal items, support and a preset sorting method.
[0103] Based on the attribute data A, C, E, G, B, D, F in the abnormal items, the frequency of occurrence of the attribute data in the abnormal items (obtained by the corresponding support calculation) 8, 8, 8, 5, 2, 2, 2 and the preset sorting method (for example, sorting method in descending frequency), generate a frequent item header table and a sorted data set.
[0104] For example, the frequent item header table is as follows:
[0105]
[0106] A sorted dataset, for example:
[0107] ACEBF
[0108] ACG
[0109] E
[0110] ACEGD
[0111] ACEG
[0112] E
[0113] ACEBF
[0114] ACD
[0115] ACEG
[0116] ACEG
[0117] Step S204: construct a frequent pattern tree based on the sorted data set.
[0118] Insert each row of the sorted data set into the tree in turn, starting from the root node of the tree. Starting from the root node, if the current node already exists in the tree, add 1 to the count attribute in the node and enter the next loop; if the current node does not exist in the tree, insert a new leaf node based on the parent node. The leaf node name is the current item value and the count attribute is assigned to 1; finally, a frequent pattern tree is obtained.
[0119] Step S205: generating a preset frequent item set based on the frequent pattern tree.
[0120] Specifically, generating a preset frequent item set includes: generating a node linked list based on a frequent item header table and a frequent pattern tree; determining a conditional pattern base of each frequent item in the frequent item header table based on the node linked list; and generating a preset frequent item set based on each conditional pattern base.
[0121] For example, according to the frequency corresponding to each letter in the frequent item header table, find the corresponding letter in the frequent pattern tree, connect them with lines to obtain a node list for easy association.
[0122] For example, starting from the bottom F node in the frequent item header table, according to the node linked list, determine that the frequency of the F node in the frequent item header table is 2. In the frequent pattern tree, there is only one node with a frequency of 2. Then start from this node and look up for its conditional pattern base, which is BECA.
[0123] For example, the frequency of node E in the frequent item header table is 8, and in the frequent pattern tree it consists of a node with a frequency of 6 and a node with a frequency of 2. Then, in the frequent pattern tree, start looking upward from these two nodes to find its conditional pattern base. The conditional pattern base of node E with a frequency of 6 is CA, and the conditional pattern base of the other node with a frequency of 2 is null. After merging, the conditional pattern base of node E in the frequent item header table is CA.
[0124] The method for determining the conditional pattern bases of other nodes in the frequent item header table is similar to the above example and will not be described in detail here.
[0125] Specifically, based on each conditional pattern base, a preset frequent item set is generated, including: based on each conditional pattern base, generating a candidate frequent item set; based on a preset item set length and a preset support threshold, screening the candidate frequent item set to obtain the preset frequent item set.
[0126] For example, we start mining from the bottom F node in the frequent item header table. It is known that the conditional pattern base of F is BECA. In the sorted data set, in the data where F appears, such as ACEBF, ACEBF, we find the frequency of BECA in turn, that is, the number of times BECA appears, and determine whether the number of times each BECA letter appears is greater than the frequency corresponding to the minimum support. If so, it is retained, otherwise the corresponding letter is deleted. Then the frequency of BECA in the data where F appears is obtained: B: 2, E: 2, C: 2, A: 2. Then, based on the frequency of BECA in the data where F appears, the largest frequent item set (5 item sets) is generated, that is, a candidate frequent item set is obtained: [A: 2, C: 2, E: 2, B: 2, F2]. The calculation method of the candidate frequent item sets corresponding to other nodes in the frequent item header table is similar to that of the F node, which will not be repeated here. Finally, a candidate frequent item set consisting of the frequent item sets corresponding to each node in the frequent item header table is obtained. Based on the preset item set length and the preset support threshold, the candidate frequent item sets are screened to accurately obtain the preset frequent item sets. This can improve the accuracy of anomaly identification.
[0127] Step S206: When there is one or more preset frequent item sets whose attribute data completely matches the attribute data in the acquired fingerprint item data, an abnormal flag is marked for the fingerprint item data.
[0128] The preset frequent item sets may include: device attributes: the values of device attributes. There may be multiple preset frequent item sets. When the device attribute data of one or more preset frequent item sets completely matches the attribute data in the acquired fingerprint item data, an abnormal flag is marked for the fingerprint item data.
[0129] Step S207: outputting the fingerprint item data corresponding to all the abnormal identifications and the business scenario identification corresponding to the fingerprint item data marked with the abnormal identification.
[0130] There may be multiple fingerprint item data marked with abnormal identification. All fingerprint item data marked with abnormal identification and the business scenario identifications corresponding to these fingerprint item data are output for the processing node to perform abnormal processing and risk warning.
[0131] The anomaly identification method of the embodiment of the present application can be applied to risk control scenarios and user similarity judgment. Anomaly identification can be device anomaly identification or user behavior anomaly identification. The following uses device anomaly identification as an example for illustration. User behavior anomaly identification is similar and will not be described in detail.
[0132] Figure 3 This is a schematic diagram of the main process of the anomaly identification method according to an embodiment of the present application. Taking device anomaly identification as an example, the relationship between device fingerprints and actual business data is constructed. In the business scenario domain, highly correlated device attributes with identification information are screened. By mining continuously associated entities, similar device attributes and device continuity attack business risks are identified, and problems are proactively discovered in a targeted manner, thereby aligning the layout of risk control strategies for prevention. The prerequisite for the implementation of the anomaly identification method of the embodiment of the present application is a business buried point device fingerprint product to ensure that the device fingerprint related information can be stably collected and obtained.
[0133] like Figure 3 As shown, the anomaly identification method of the embodiment of the present application is implemented based on the following components: basic data, fingerprint item screening, high-frequency item filtering, support data, risk association rule output, rule organization and rule application.
[0134] Specifically, component (I): basic data (i.e. historical abnormal data): the basic data of different business scenarios are located differently. The basic data of the traffic field is mainly separate browsing and access information, and the main data of the order submission field is order purchase information, including order number, product information, delivery province information, etc. However, this solution can switch subsequent functions on the basis of ensuring the basic data. Take the order scenario as an example to explain the functions. For the order scenario, the main function of this component is to process the order scenario data in a standardized format. The main contents of the data are as follows, but not limited to the following data:
[0135]
[0136] The sample data of the device attributes when the order is submitted in the above table is as follows:
[0137] Order Number Other columns Equipment attributes when order is submitted No1 ... {brand: AAA, root: 0, ip: xxx: xxx: xx: xx, caid: 123-123} No2 .. {brand: BBB, p_name: fd, channel: os, col: 123} No3 .. {brand:CCC,p_name:bavol,col:12}
[0138] The conventional way to process data is ETL (Extract-Transform-Load), which can be implemented through spark, hive, etc. It is particularly important to note here that in order to meet the different device attributes collected by matching different device fingerprint products, the processing of "device attributes when order submission" in the above table can be done through configuration files. The names of device attributes are recorded in the configuration files, and then in the ETL processing logic, the device data values in the basic data are obtained through a loop, and then the device attribute names and device attribute values are stored in the dictionary that stores the results. In addition, this component needs to ensure the uniqueness of the order number as the primary key, that is, there is one record for one order number.
[0139] Component (two): Fingerprint item screening: The main function of this component is to eliminate unusable device attributes based on the basic data produced by component (one), retain useful information, and avoid problems such as waste of computing resources and decreased device accuracy caused by junk information in subsequent functions. Eliminating unusable device attributes means directly removing the key (i.e. attribute name) and the corresponding value (i.e. attribute value) in the dictionary from the basic data "device attributes at order submission" of component (one); so the number of data items finally produced by this component is exactly the same as that of component (one). Because there is instability in the process of collecting device attributes by device fingerprints, and because of issues such as permissions and networks, some device attributes may not be available in the entire domain. Conventional elimination methods include but are not limited to the following:
[0140] a) Data item is empty: In the orders within a certain time window range, if the device attribute has a high proportion of empty values (orders with empty data items / all orders), it means that it is unavailable. Data items with empty values exceeding a certain threshold are directly eliminated;
[0141] b) Data not related to business: Through manual screening, some equipment attributes are completely irrelevant to risks and can be eliminated, such as equipment reporting time, person in charge and other information;
[0142] c) Highly unique attributes: Equipment attribute values that have a large amount of data uniqueness will be eliminated. They can be eliminated using the following example formula: Equipment attribute duplicate quantity / total order quantity>0.95 (the elimination threshold here is manually defined).
[0143] Component (three): High-frequency item filtering: The function of this component is to continue to eliminate equipment items with low information differentiation based on the results of component (two). The elimination operation is similar to component (two), which means directly removing the key and key corresponding value in the dictionary from the basic data "equipment attributes when order is submitted" of component (two); so the number of data items finally output by this component is exactly the same as that of component (two). Because the distribution of some equipment information items may be relatively uniform, there is no distinction between anomalies, which will ineffectively increase the subsequent output rules. Such equipment information items can be eliminated. The elimination methods include but are not limited to the following:
[0144] a) Data purity is too high: Data purity, borrowing the concept from the decision tree algorithm, calculates the data purity in the sample through information entropy. However, the decision tree removes information with low purity, but here it is slightly different, retaining samples with low data purity. The main reason is the difference in p_k. Here p_k means the proportion of data samples. Data purity is too high, which means that the data is too high in frequency and has no dispersion. The definition of information entropy is as follows:
[0145]
[0146] The data set D (which can be obtained from basic data, such as historical abnormal data) is a sample set of a certain device attribute, |y| is the number of types of the device attribute, and p k is the proportion of the kth category of the device attribute:
[0147] p k = Number of samples in the kth class / Number of samples D
[0148] A threshold for manual data analysis can be set to eliminate device attributes with too small Ent(D).
[0149] b) If the variance of a numerical attribute is too small, it means that the data is evenly distributed and does not have much information differentiation. The variance is calculated as follows:
[0150] δ 2 =∑(xu) 2 / N
[0151] Where x is the value of the device attribute, u is the mean of the device attribute, and N is the total number of samples. Here, we also need to set an artificial threshold for elimination.
[0152] Component (IV): Support data: The core function of this component is to calculate the frequency of the corresponding values of device attributes, and then eliminate device attributes with too low or too high frequencies. In conventional association rules algorithms, there is only a single filter of support, but there are certain differences between risk control and conventional algorithms. Risk control focuses on mining a small amount of abnormal data from a large range of data, so here, a maximum support limit is also set for the support data. That is, based on the data produced by component (III), samples in the data set that do not meet the following conditions are eliminated, and the conditions for the retained sample data are as follows:
[0153] Minimum support <= support data <= maximum support.
[0154] For support data, it is the frequency of non-repeated occurrence of device attribute values in the overall sample. For example, support can be the frequency of occurrence orders of device items. The frequency corresponding to the support can be obtained by multiplying the support by the total number of occurrences of the data item in the corresponding data set (such as historical abnormal data). The minimum support data can be set directly according to the business and manual experience of risk control. The smaller the data, the more recalls will be brought, but the accuracy of the recalled risk control sample abnormalities will be problematic. The larger the data, the better the model accuracy will be, but the number of recalled samples will be reduced, and it can be adjusted iteratively step by step; the maximum support can be set simultaneously through the sample proportion and absolute value. The absolute value represents a fixed manual experience value, and the sample proportion is the threshold that the value cannot exceed in the total sample. When the value of the device attribute appears too many times and frequently, it is likely to be a regular value, such as AAA brand mobile phones. This attribute value itself is a high-frequency value, which is not very helpful for subsequent risk rules. In addition, there are some differences between the filtering here and component (two) and component (three). Components (two) and (three) directly remove the device attributes from all samples. After removal, all samples will not have the device attributes. This component removes the value of the device attribute. After the sample is removed, the attribute data of the device will still exist in other samples.
[0155] Component (V): Risk association rule output: The main function of this component is to output specific risk association rules (i.e., output preset frequent item sets) through the frequent pattern growth FP-grouth algorithm. In the embodiment of the present application, the output risk association rules are preset frequent item sets. The frequent pattern tree FP-tree algorithm is described as follows:
[0156]
[0157] By traversing the FP-tree as described above, this component can generate corresponding risk association rules (i.e., frequent item sets in this application) and store them in a dictionary. An example of risk association rules (i.e., frequent item sets in this application) in the dictionary storing the results is as follows:
[0158] ['null', {'brand:redmin':8}, {'os:android 4.2':4}, {'changing:1':3}, {'channel:jd':2}, {'bissid:xxxxx':2}]
[0159] ['null', {'brand:redmin':8}, {'os:android 4.2':4}, {'changing:1':3}, {'timezose:us':1}]
[0160] ['null', {'brand:redmin':8}, {'os:android 4.2':4}, {'dip:1999':1}, {'nfc:1':1}
[0161] Among them, null is the root node of the tree and has no practical meaning. The key in the dictionary is "device attribute name: device attribute value", such as brand:redmin in the above example. The value of the dictionary is the frequency of occurrence of the corresponding "device attribute name: device attribute value", such as 8 after 'brand:redmin' in the above example.
[0162] Component (six): rule arrangement (in the embodiment of the present application, the rule is the risk association rule, that is, the frequent item set, and the rule arrangement is the arrangement of the frequent item set). The main function of this component is to screen the output risk association rules (that is, the output frequent item sets) in combination with actual scenarios and applications, eliminate invalid risk association rules (that is, eliminate invalid frequent item sets), and retain valid risk association rules (that is, retain valid frequent item sets as preset frequent item sets) to ensure the accuracy of the model. The parameter thresholds involved in this process can be combined with component (four) for joint adjustment; and there is a certain correlation between the two. If the threshold in component (four) is adjusted to a lower level, the accuracy can be improved accordingly by the length of the rules in this component, and vice versa. The rule arrangement in this component mainly includes the following aspects:
[0163] a) Rule length (i.e., frequent item set length): Eliminate rules that are too short. A short rule length is likely caused by coincidence in the data. By increasing the rule length, the accuracy of the final model recall can be improved.
[0164] b) Frequency value of device attribute values in rules: In component (IV), the data frequency has been basically guaranteed through support data, but support data is a global value. The attribute frequency in the rule is more detailed and guarantees the frequency of the joint device attributes and attribute values in the rule. Because the rule attribute values produced by component (V) have certain particularities, the attribute frequency at the front of the rule list is higher than that at the back, so the rule of rule frequency needs to start from the end of the rule list.
[0165] After a) and b) above, invalid rules (i.e. invalid frequent item sets) have been eliminated, and the final output is all the abnormal identification attribute value rules (i.e. preset frequent item sets) produced by the current model. It should also be noted that the rule length in a) and the minimum frequency value in b) need to be adjusted based on the magnitude of data in the scenario application and human experience, and continuously iterated.
[0166] Component (VII): Rule application. The main function of this component is to compare the basic data in component (I) with the abnormal rules in component (VI) (i.e., the preset frequent item sets). If the rule attributes (i.e., the attributes of the preset frequent item sets) and attribute values (i.e., the attribute values of the preset frequent item sets) in component (VI) completely match the "device attributes when order is submitted" in component (I), i.e., the device attribute data (including attribute name and attribute value) when the order is submitted, then the order is considered to be an abnormal order. After passing through this component, all abnormal order sets are finally output.
[0167] The embodiment of the present application can realize the automatic generation of risk association rules (i.e., automatically generate preset frequent item sets), reduce costs and increase efficiency, realize automatic identification and confrontation of potential risks, better protect business security, and expand the scope of risk recall through model iteration. The embodiment of the present application takes the device attributes of the device fingerprint product as a prerequisite to solve similar continuous attacks in risk control scenarios, but it can also be generalized in the following directions:
[0168] a) Device attributes can be replaced by other behavioral data, such as the user's click sequence;
[0169] b) In addition to risk control, the scenario expansion can be generalized to user similarity judgment, etc. There are certain related behaviors and users have certain similarities;
[0170] c) Generalization of the association algorithm, focusing on seeking certain behavioral sequences with similar associations. The fp-grouth algorithm is only one of the algorithms, and it can also be implemented through methods such as rule learning.
[0171] Figure 4 Schematic diagram of the main units of the abnormality identification device according to the embodiment of the present application. Figure 4As shown, the abnormality identification device 400 includes an acquisition unit 401 , a frequent item set generation unit 402 , a marking unit 403 and an output unit 404 .
[0172] The acquisition unit 401 is configured to acquire a business scenario identifier according to the received abnormality identification request, and acquire corresponding fingerprint item data according to the business scenario identifier.
[0173] The frequent item set generation unit 402 is configured to obtain the corresponding historical abnormal data in real time according to the business scenario identifier, determine the support of the abnormal items in the historical abnormal data, build a frequent pattern tree based on the abnormal items and the support, and generate a preset frequent item set based on the frequent pattern tree.
[0174] The marking unit 403 is configured to mark an abnormality mark for the fingerprint item data when there is one or more preset frequent item sets whose attribute data completely matches the attribute data in the acquired fingerprint item data.
[0175] The output unit 404 is configured to output the fingerprint item data corresponding to all the abnormal identifications and the business scenario identification corresponding to the fingerprint item data marked with the abnormal identification.
[0176] In some embodiments, the frequent item set generating unit 402 is further configured to: generate a frequent item header table and a sorted data set based on the abnormal items, support and a preset sorting method; and construct a frequent pattern tree based on the sorted data set.
[0177] In some embodiments, the frequent item set generation unit 402 is further configured to: generate a node linked list based on the frequent item header table and the frequent pattern tree; determine a conditional pattern base of each frequent item in the frequent item header table based on the node linked list; and generate a preset frequent item set based on each conditional pattern base.
[0178] In some embodiments, the frequent itemset generating unit 402 is further configured to: generate candidate frequent itemsets based on each conditional pattern base; and screen the candidate frequent itemsets to obtain preset frequent itemsets based on a preset itemset length and a preset support threshold.
[0179] In some embodiments, the abnormality identification device further includes Figure 4 The data processing unit not shown in the figure is configured to: calculate the information entropy of each abnormal item in the historical abnormal data, and remove the abnormal items corresponding to the information entropy less than a preset information entropy threshold from the historical abnormal data.
[0180] In some embodiments, the abnormality identification device further includes Figure 4 The data processing unit not shown in the figure is configured to: remove abnormal items corresponding to support values that are not within a preset support value range from the historical abnormal data.
[0181] It should be noted that the abnormality identification method and the abnormality identification device of the present application have a corresponding relationship in terms of specific implementation content, so the repeated content will not be described again.
[0182] Figure 5 An exemplary system architecture 500 to which the abnormality identification method or abnormality identification device according to the embodiment of the present application can be applied is shown.
[0183] like Figure 5 As shown, system architecture 500 may include terminal devices 501, 502, 503, a network 504 and a server 505. Network 504 is used to provide a medium for communication links between terminal devices 501, 502, 503 and server 505. Network 504 may include various connection types, such as wired, wireless communication links or optical fiber cables, etc.
[0184] Users can use terminal devices 501, 502, 503 to interact with server 505 through network 504 to receive or send messages, etc. Various communication client applications can be installed on terminal devices 501, 502, 503, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc. (only as examples).
[0185] The terminal devices 501, 502, 503 may be various electronic devices having an abnormality identification processing screen and supporting web browsing, including but not limited to smart phones, tablet computers, laptop computers, desktop computers, and the like.
[0186] Server 505 can be a server that provides various services, such as a background management server that supports the abnormal identification request submitted by the user using terminal devices 501, 502, and 503 (only as an example). The background management server can obtain the business scenario identifier according to the received abnormal identification request, and obtain the corresponding fingerprint item data according to the business scenario identifier; obtain the corresponding historical abnormal data in real time according to the business scenario identifier, determine the support of the abnormal items in the historical abnormal data, build a frequent pattern tree based on the abnormal items and the support, and generate a preset frequent item set based on the frequent pattern tree; when there is one or more preset frequent item sets whose attribute data completely matches the attribute data in the acquired fingerprint item data, mark the fingerprint item data with an abnormal identifier; output the fingerprint item data corresponding to all abnormal identifiers and the business scenario identifier corresponding to the fingerprint item data marked with the abnormal identifier. Accurately and timely identify and determine the risks behind the business, and improve the efficiency, accuracy and timeliness of abnormal identification.
[0187] It should be noted that the anomaly identification method provided in the embodiment of the present application is generally executed by the server 505, and accordingly, the anomaly identification device is generally set in the server 505.
[0188] It should be understood that Figure 5 The number of terminal devices, networks and servers in the embodiment is only for illustration. Any number of terminal devices, networks and servers may be provided according to implementation requirements.
[0189] Reference below Figure 6 , which shows a schematic diagram of the structure of a computer system 600 of a terminal device suitable for implementing an embodiment of the present application. Figure 6 The terminal device shown is merely an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.
[0190] like Figure 6 As shown, the computer system 600 includes a central processing unit (CPU) 601, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 602 or a program loaded from a storage part 608 into a random access memory (RAM) 603. In the RAM 603, various programs and data required for the operation of the computer system 600 are also stored. The CPU 601, the ROM 602, and the RAM 603 are connected to each other via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.
[0191] The following components are connected to the I / O interface 605: an input section 606 including a keyboard, a mouse, etc.; an output section 607 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 608 including a hard disk, etc.; and a communication section 609 including a network interface card such as a LAN card, a modem, etc. The communication section 609 performs communication processing via a network such as the Internet. A drive 610 is also connected to the I / O interface 605 as needed. A removable medium 611, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 610 as needed, so that a computer program read therefrom is installed into the storage section 608 as needed.
[0192] In particular, according to the embodiments disclosed in the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, the embodiments disclosed in the present application include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 609, and / or installed from the removable medium 611. When the computer program is executed by the central processing unit (CPU) 601, the above-mentioned functions defined in the system of the present application are executed.
[0193] It should be noted that the computer-readable medium shown in the present application may be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. Computer-readable storage media may include, for example, but are not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices or devices, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in combination with an instruction execution system, device or device. In the present application, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, in which a computer-readable program code is carried. This propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. Computer-readable signal media may also be any computer-readable medium other than computer-readable storage media, which may send, propagate or transmit a program for use by or in conjunction with an instruction execution system, apparatus or device. The program code contained on the computer-readable medium may be transmitted using any appropriate medium, including but not limited to: wireless, wire, optical cable, RF, etc., or any suitable combination of the above.
[0194] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present application. In this regard, each box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the above-mentioned module, program segment or a part of a code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flow chart, and the combination of the boxes in the block diagram or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0195] The units involved in the embodiments described in the present application may be implemented by software or hardware. The units described may also be set in a processor, for example, it may be described as: a processor includes an acquisition unit, a frequent item set generation unit, a marking unit, and an output unit. The names of these units do not constitute limitations on the units themselves in some cases.
[0196] As another aspect, the present application also provides a computer-readable medium, which may be included in the device described in the above embodiment; or it may exist independently and not be assembled into the device. The above computer-readable medium carries one or more programs. When the above one or more programs are executed by a device, the device obtains a business scenario identifier according to the received abnormal identification request, and obtains corresponding fingerprint item data according to the business scenario identifier; obtains corresponding historical abnormal data in real time according to the business scenario identifier, determines the support of abnormal items in the historical abnormal data, builds a frequent pattern tree based on the abnormal items and the support, and generates a preset frequent item set based on the frequent pattern tree; when there is one or more preset frequent item sets whose attribute data completely matches the attribute data in the acquired fingerprint item data, the fingerprint item data is marked with an abnormal identifier; and outputs the fingerprint item data corresponding to all abnormal identifiers and the business scenario identifier corresponding to the fingerprint item data marked with the abnormal identifier.
[0197] The computer program product of the present application includes a computer program, and when the computer program is executed by a processor, the abnormality identification method in the embodiment of the present application is implemented.
[0198] According to the technical solution of the embodiment of the present application, it is possible to accurately and timely identify and determine the risks behind the business, and improve the efficiency, accuracy and timeliness of anomaly identification.
[0199] The above specific implementations do not constitute a limitation on the protection scope of this application. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions may occur depending on design requirements and other factors. Any modifications, equivalent substitutions and improvements made within the spirit and principles of this application should be included in the protection scope of this application.
Claims
1. A method for identifying anomalies, characterized in that: include: Obtaining a business scenario identifier according to the received exception identification request, and obtaining corresponding fingerprint item data according to the business scenario identifier; Acquire corresponding historical abnormal data in real time according to the business scenario identifier, determine the support of abnormal items in the historical abnormal data, build a frequent pattern tree based on the abnormal items and the support, and generate a preset frequent item set based on the frequent pattern tree; When there is one or more preset frequent item sets whose attribute data completely matches the attribute data in the acquired fingerprint item data, marking the fingerprint item data with an abnormality mark; Output the fingerprint item data corresponding to all abnormal identifications and the business scenario identification corresponding to the fingerprint item data marked with abnormal identifications.
2. The method according to claim 1, characterized in that The constructing a frequent pattern tree based on the abnormal items and the support comprises: Based on the abnormal items, the support and the preset sorting method, generating a frequent item header table and a sorted data set; A frequent pattern tree is constructed based on the sorted data set.
3. The method according to claim 2, characterized in that The generating of a preset frequent item set includes: Based on the frequent item header table and the frequent pattern tree, generating a node linked list; Based on the node linked list, determining a conditional pattern base of each frequent item in the frequent item header table; Based on each of the conditional pattern bases, a preset frequent item set is generated.
4. The method according to claim 3, characterized in that The generating of preset frequent item sets based on each of the conditional pattern bases includes: Generate candidate frequent item sets based on each conditional pattern base; Based on a preset item set length and a preset support threshold, the candidate frequent item sets are screened to obtain a preset frequent item set.
5. The method according to claim 1, characterized in that Before determining the support degree of the abnormal item in the historical abnormal data, the method further includes: The information entropy of each abnormal item in the historical abnormal data is calculated, and the abnormal items corresponding to the information entropy less than a preset information entropy threshold are eliminated from the historical abnormal data.
6. The method according to claim 1, characterized in that Before constructing the frequent pattern tree based on the abnormal items and the support, the method further includes: Abnormal items corresponding to support degrees that are not within a preset support degree range are removed from the historical abnormal data.
7. An abnormality identification device, characterized in that: include: an acquisition unit, configured to acquire a business scenario identifier according to the received abnormal identification request, and acquire corresponding fingerprint item data according to the business scenario identifier; A frequent item set generation unit is configured to obtain corresponding historical abnormal data in real time according to the business scenario identifier, determine the support of abnormal items in the historical abnormal data, build a frequent pattern tree based on the abnormal items and the support, and generate a preset frequent item set based on the frequent pattern tree; a marking unit configured to mark an abnormality mark for the fingerprint item data when there is one or more attribute data of a preset frequent item set that completely matches the attribute data in the acquired fingerprint item data; The output unit is configured to output the fingerprint item data corresponding to all abnormal identifications and the business scenario identification corresponding to the fingerprint item data marked with the abnormal identification.
8. An abnormality recognition electronic device, characterized in that: include: one or more processors; a storage device for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 6.
9. A computer readable medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.