Abnormal data identification method and device, storage medium and electronic equipment
By using the exception database and abnormal affinity to identify potential abnormal data, and determining the identification threshold through cluster evaluation coefficients, the problem of low accuracy in the identification of abnormal data in the prior art is solved, and efficient identification of different abnormal feature data is achieved.
Patent Information
- Application Number
- CN202311567764.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-22
- Publication Date
- 2025-05-23
AI Technical Summary
In the prior art, when identifying abnormal data, the recognition accuracy is low and it is impossible to effectively identify abnormal data with different abnormal characteristics.
通过获取目标业务场景下的业务数据,利用异常数据库从中识别出潜在异常数据。 异常数据库中的异常数据是通过异常亲和度对基础异常数据进行部分选择变异后得到的。 Based on the abnormal affinity of the potential abnormal data, the reference recognition threshold and the cluster evaluation coefficient are determined, and the target recognition threshold is determined through the cluster evaluation coefficient, and data with an abnormal affinity greater than the target recognition threshold are identified as the target abnormal data.
The identification accuracy of abnormal data is improved, and abnormal data with different abnormal characteristics can be effectively identified, solving the problem of low recognition accuracy in the prior art.
Smart Images

Figure CN120030457A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computers, and in particular, to a method and device for identifying abnormal data, a storage medium, and an electronic device. Background Art
[0002] In recent years, with the vigorous development of customer-oriented (also known as ToC) services in the Internet, a large amount of user service data has been generated. However, a lot of abnormal data is often mixed in the large-scale service data.
[0003] In order to identify the above abnormal data from the large-scale data, the currently related technologies mainly adopt the following methods: identifying abnormal data based on statistical methods. The statistical methods adopted here usually rely on a preset fixed threshold, such as marking the data exceeding the fixed threshold as abnormal data.
[0004] However, the fixed threshold set by the above statistical method is often a single value set manually according to the experience of professionals, and the threshold condition composed of such a single value has certain limitations. It can only identify a part of the abnormal data and cannot identify the abnormal data with different abnormal characteristics. In other words, the method for identifying abnormal data provided by the related technologies has the problem of low identification accuracy.
[0005] In response to the above problems, no effective solution has been proposed yet. Summary of the Invention
[0006] The embodiments of this application provide a method and device for identifying abnormal data, a storage medium, and an electronic device, so as to at least solve the technical problem of low identification accuracy existing in the method for identifying abnormal data provided by the related technologies.
[0007] According to one aspect of the embodiments of this application, a method for identifying abnormal data is provided, including: obtaining service data to be identified in a target service scenario; identifying potential abnormal data from the service data by using an abnormal database, where the abnormal data in the abnormal database is obtained by partially selecting and mutating basic abnormal data by using an abnormal affinity, and the abnormal affinity is used to indicate the affinity degree between abnormal data and normal data; determining at least one reference identification threshold and a clustering evaluation coefficient respectively matching each reference identification threshold based on the abnormal affinity corresponding to each potential abnormal data, where the clustering evaluation coefficient is used to evaluate the result after clustering the potential abnormal data by using the reference identification threshold; and when determining a target identification threshold from at least one reference identification threshold based on the clustering evaluation coefficient, identifying the potential abnormal data with an abnormal affinity greater than the target identification threshold as target abnormal data.
[0008] According to another aspect of an embodiment of the present application, a device for identifying abnormal data is also provided, including: an acquisition unit, used to acquire business data to be identified in a target business scenario; a first identification unit, used to identify potential abnormal data from the business data using an abnormal database, wherein the abnormal data in the abnormal database is obtained after partial selection and mutation of basic abnormal data using abnormal affinity, and the abnormal affinity is used to indicate the degree of affinity between the abnormal data and normal data; a determination unit, used to determine at least one reference identification threshold and a clustering evaluation coefficient respectively matched with each reference identification threshold based on the abnormal affinity corresponding to each potential abnormal data, wherein the clustering evaluation coefficient is used to evaluate the result of clustering the potential abnormal data using the reference identification threshold; a second identification unit, used to identify the potential abnormal data whose abnormal affinity is greater than the target identification threshold as target abnormal data when a target identification threshold is determined from at least one reference identification threshold based on the clustering evaluation coefficient.
[0009] Optionally, in this embodiment, the above-mentioned determination unit includes: a first acquisition module, which is used to obtain the maximum abnormal affinity and the minimum abnormal affinity from the abnormal affinities corresponding to the potential abnormal data; an extraction module, which is used to extract a numerical value as a reference recognition threshold in the numerical interval formed by the minimum abnormal affinity and the maximum abnormal affinity according to the target step size; a clustering processing module, which is used to cluster the potential abnormal data in turn using each reference recognition threshold to obtain at least two abnormal data clusters matching the reference recognition threshold; a first determination module, which is used to determine the clustering evaluation coefficient matching the reference recognition threshold using the clustering distance between each abnormal data in at least two abnormal data clusters.
[0010] Optionally, in this embodiment, the above-mentioned first clustering processing module is also used to take each reference identification threshold as the current reference identification threshold in turn, and perform the following operations: compare the abnormal affinity corresponding to each potential abnormal data with the current reference identification threshold; cluster the potential abnormal data with an abnormal affinity greater than the current reference identification threshold into a first abnormal data cluster, and cluster the potential abnormal data with an abnormal affinity less than or equal to the current reference identification threshold into a second abnormal data cluster; determine the first abnormal data cluster and the second abnormal data cluster as abnormal data clusters that match the current reference identification threshold.
[0011] Optionally, in this embodiment, the above-mentioned first determination module is also used to take each abnormal data in at least two abnormal data clusters as the current abnormal data in turn, and perform the following operations: determine the current abnormal data cluster where the current abnormal data is located; obtain the first average clustering distance between the current abnormal data and each first reference abnormal data contained in the current abnormal data cluster; obtain the second average clustering distance between the current abnormal data and each second reference abnormal data contained in other abnormal data clusters outside the current abnormal data cluster; use the first average clustering distance and the second average clustering distance to determine the current object clustering evaluation coefficient matching the current abnormal data; when the object clustering evaluation coefficient corresponding to each abnormal data in at least two abnormal data clusters is obtained, perform weighted sum calculation on all object clustering evaluation coefficients to obtain a clustering evaluation coefficient matching the reference recognition threshold.
[0012] Optionally, in this embodiment, the above-mentioned device also includes: a sorting unit, used to sort the clustering evaluation coefficients matching each reference recognition threshold; a first determination unit, used to determine the reference recognition threshold corresponding to the largest clustering evaluation coefficient as the target recognition threshold.
[0013] Optionally, in this embodiment, the above-mentioned device also includes: a first acquisition unit, used to acquire basic abnormality data; a second acquisition unit, used to acquire extended abnormality data other than the basic abnormality data through at least one immune selection mode, wherein the immune selection mode is used to select abnormality data similar to the basic abnormality data from the data using an immune algorithm; and an updating unit, used to update the abnormality database using the extended abnormality data.
[0014] Optionally, in this embodiment, the second acquisition unit includes: a second acquisition module, used to acquire the candidate extended abnormality data determined by each immune selection mode; a second determination module, used to perform one of the following operations on the candidate extended abnormality data to determine the extended abnormality data: averaging the candidate extended abnormality data determined by each immune selection mode to obtain the extended abnormality data; performing weighted sum calculation on the candidate extended abnormality data determined by each immune selection mode to obtain the extended abnormality data; voting processing on the candidate extended abnormality data determined by each immune selection mode to obtain the extended abnormality data.
[0015] Optionally, in this embodiment, the above-mentioned device also includes: a comparison unit, which is used to compare the abnormal affinity and the clone threshold corresponding to each basic abnormal data in the first immune selection mode; a clone copy unit, which is used to clone and copy the basic abnormal data whose abnormal affinity is greater than the clone threshold to obtain a copy of the abnormal data that matches the basic abnormal data; a mutation processing unit, which is used to perform mutation processing on the abnormal data copy to obtain mutated abnormal data; a second determination unit, which is used to use the mutated abnormal data as the candidate extended abnormal data determined in the first immune selection mode when the abnormal affinity of the mutated abnormal data is greater than the abnormal affinity of the abnormal data copy.
[0016] Optionally, in this embodiment, the above-mentioned device also includes: a first training unit, used to expose a group of detectors to normal data for training in the second immune selection mode to obtain target detectors for detecting abnormal data, wherein, during the training process, detectors that match the normal data will be discarded, and detectors that do not match the normal data will be retained; a third determination unit, used to use the data identified by using the target detector as candidate extended abnormal data determined in the second immune selection mode.
[0017] Optionally, in this embodiment, the above-mentioned first identification unit includes: a cleaning processing module, which is used to clean the business data to obtain processed business data; an encoding module, which is used to perform feature encoding on the processed business data to obtain business features; a normalization processing module, which is used to normalize the business features to obtain business features of a unified scale; a third acquisition module, which is used to obtain feature similarity between the business features and data features of abnormal data in the abnormal database; and an identification module, which is used to identify the business data corresponding to the business features whose feature similarity is greater than the target threshold as potential abnormal data.
[0018] Optionally, in this embodiment, the above-mentioned device also includes: a third acquisition unit, used to obtain neighboring normal data associated with the target abnormal data from the business data, wherein the neighboring normal data is normal data whose distance from the target abnormal data is less than a target distance threshold; and a repair unit, used to repair the target abnormal data using the neighboring normal data.
[0019] According to another aspect of the embodiments of the present application, a computer-readable storage medium is provided, in which a computer program is stored, wherein the computer program is configured to execute the above-mentioned abnormal data identification method when running.
[0020] According to another aspect of the embodiments of the present application, a computer program product or a computer program is provided, the computer program product or the computer program including computer instructions, the computer instructions being stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device performs the above abnormal data identification method.
[0021] According to another aspect of the embodiments of the present application, there is also provided an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to execute the above-mentioned abnormal data identification method through the computer program.
[0022] In an embodiment of the present application, business data to be identified in a target business scenario is obtained. Then, potential abnormal data is identified from the business data using an abnormal database, wherein the abnormal data in the abnormal database is obtained after partial selection and variation of the basic abnormal data using abnormal affinity, and the abnormal affinity is used to indicate the affinity between the abnormal data and the normal data. Further, based on the abnormal affinity corresponding to each of the potential abnormal data, at least one reference recognition threshold is determined, and a clustering evaluation coefficient that matches each reference recognition threshold is determined, wherein the clustering evaluation coefficient is used to evaluate the result of clustering the potential abnormal data using the reference recognition threshold. Then, in the case where a target recognition threshold is determined from at least one reference recognition threshold based on the clustering evaluation coefficient, the potential abnormal data whose abnormal affinity is greater than the target recognition threshold is identified as the target abnormal data. In other words, using an embodiment of the present application, the potential abnormal data after being screened by the abnormal database is compared using an adaptive threshold to determine the target abnormal data from the business data. The purpose of effectively and accurately identifying large-scale data abnormalities is achieved, and it is no longer limited to marking data exceeding a fixed threshold as abnormal data. Thereby, the technical effect of improving the recognition accuracy of abnormal data is achieved, and the problem of low recognition accuracy in the recognition method of abnormal data provided by related technologies is solved. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:
[0024] Figure 1 is a schematic diagram of an application environment of an optional abnormal data identification method according to an embodiment of the present application;
[0025] Figure 2 is a flow chart of an optional abnormal data identification according to an embodiment of the present application;
[0026] Figure 3 is a flow chart of an optional abnormal data identification according to an embodiment of the present application;
[0027] Figure 4 is a flow chart of an optional abnormal data identification according to an embodiment of the present application;
[0028] Figure 5 is a schematic diagram of an optional abnormal data identification according to an embodiment of the present application;
[0029] Figure 6 is a schematic diagram of an optional abnormal data identification according to an embodiment of the present application;
[0030] Figure 7 is a flow chart of an optional abnormal data identification according to an embodiment of the present application;
[0031] Figure 8 is a flow chart of an optional abnormal data identification according to an embodiment of the present application;
[0032] Fig. 9 is a schematic structural diagram of an optional abnormal data identification device according to an embodiment of the present application;
[0033] Fig.10 It is a schematic diagram of the structure of an optional electronic device according to an embodiment of the present application. DETAILED DESCRIPTION
[0034] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present application.
[0035] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0036] According to one aspect of the embodiments of the present application, a method for identifying abnormal data is provided. Optionally, as an optional implementation, the above-mentioned identification of abnormal data can be but is not limited to being applied to: Figure 1 In the environment shown. Figure 1 As shown, the terminal device 102 includes a memory 104 for storing various data generated during the operation of the terminal device 102, a processor 106 for processing and calculating the above data, and a display 108. The terminal device 102 can exchange data with a server 112 through a network 110. The server 112 is connected to a database 114, and the database 114 is used to store various data. The terminal device 102 can run an application for identifying business data.
[0037] Furthermore, the above method Figure 1 The corresponding specific application process in the environment shown is shown in the following steps:
[0038] Execute step S102, the terminal device 102 obtains the business data to be identified in the target business scenario.
[0039] Then, steps S104-S110 are performed, and the terminal device 102 sends the business data to the server 112 through the network 110. The server 112 uses the abnormal database to identify potential abnormal data from the business data, wherein the abnormal data in the abnormal database is obtained by partially selecting and mutating the basic abnormal data using abnormal affinity, and the abnormal affinity is used to indicate the affinity between the abnormal data and the normal data. Based on the abnormal affinity corresponding to each potential abnormal data, the server 112 determines at least one reference recognition threshold and a clustering evaluation coefficient that matches each reference recognition threshold, wherein the clustering evaluation coefficient is used to evaluate the result of clustering the potential abnormal data using the reference recognition threshold. When the server 112 determines the target recognition threshold from at least one reference recognition threshold based on the clustering evaluation coefficient, the potential abnormal data with an abnormal affinity greater than the target recognition threshold is identified as the target abnormal data.
[0040] Then, step S112 is executed, and the server 112 sends the target abnormal data to the terminal device 102 through the network 110 .
[0041] In an embodiment of the present application, business data to be identified in a target business scenario is obtained. Then, potential abnormal data is identified from the business data using an abnormal database, wherein the abnormal data in the abnormal database is obtained after partial selection and variation of the basic abnormal data using abnormal affinity, and the abnormal affinity is used to indicate the affinity between the abnormal data and the normal data. Further, based on the abnormal affinity corresponding to each of the potential abnormal data, at least one reference recognition threshold is determined, and a clustering evaluation coefficient that matches each reference recognition threshold is determined, wherein the clustering evaluation coefficient is used to evaluate the result of clustering the potential abnormal data using the reference recognition threshold. Then, in the case where a target recognition threshold is determined from at least one reference recognition threshold based on the clustering evaluation coefficient, the potential abnormal data whose abnormal affinity is greater than the target recognition threshold is identified as the target abnormal data. In other words, using an embodiment of the present application, the potential abnormal data after being screened by the abnormal database is compared using an adaptive threshold to determine the target abnormal data from the business data. The purpose of effectively and accurately identifying large-scale data abnormalities is achieved, and it is no longer limited to marking data exceeding a fixed threshold as abnormal data. Thereby, the technical effect of improving the recognition accuracy of abnormal data is achieved, and the problem of low recognition accuracy in the recognition method of abnormal data provided by related technologies is solved.
[0042] Optionally, in this embodiment, the terminal device may be a terminal device configured with a target client, which may include but is not limited to at least one of the following: a mobile phone (such as an Android phone, an iOS phone, etc.), a laptop, a tablet computer, a PDA, a MID (Mobile Internet Devices), a PAD, a desktop computer, a smart TV, etc. The target client may be a video client, an instant messaging client, a browser client, an education client, etc. The network may include but is not limited to: a wired network, a wireless network, wherein the wired network includes: a local area network, a metropolitan area network and a wide area network, and the wireless network includes: Bluetooth, WIFI and other networks that realize wireless communication. The server may be a single server, or a server cluster consisting of multiple servers, or a cloud server. The above is only an example, and this embodiment does not impose any limitation on this.
[0043] Alternatively, as an alternative, Figure 2 As shown, the above-mentioned method for identifying abnormal data includes:
[0044] S202, obtaining business data to be identified in a target business scenario.
[0045] It should be noted that the above abnormal data identification method can be applied at least but not limited to the following scenarios: in the financial field, the identification scenario of transaction data with human intervention included in large-scale financial transaction data. In the game field, the identification scenario of plug-in game operation data included in the game operation data of large-scale user accounts.
[0046] Furthermore, assuming that the above-mentioned abnormal data identification method is applied to the identification scenario of transaction data that is considered to be interfered with in large-scale financial transaction data, the above-mentioned target business scenario is used to indicate the financial transaction scenario, and the above-mentioned business data can be used, but not limited to, to indicate the large-scale transaction record data involved in the above-mentioned financial transaction scenario. Among them, the transaction data may be mixed with transaction data that has been interfered with manually, and in this embodiment, the above-mentioned transaction data that has been interfered with manually needs to be identified.
[0047] Assuming that the above abnormal data identification method is applied in the field of gaming, in the identification scenario of external game operation data included in the game operation data of a large number of user accounts, the above target business scenario is used to indicate the game settlement scenario, and the above business data can be used, but not limited to, to indicate the large-scale game operation data obtained in the game settlement scenario. Among them, external game operation data may be mixed in the game operation data, and in this embodiment, the above external game operation data needs to be identified.
[0048] S204, identifying potential abnormal data from the business data using an abnormal database, wherein the abnormal data in the abnormal database is obtained by partially selecting and mutating the basic abnormal data using an abnormal affinity, and the abnormal affinity is used to indicate the affinity between the abnormal data and the normal data.
[0049] It should be noted that the above-mentioned abnormal database is a pre-created database for storing abnormal data. Specifically, the abnormal database includes abnormal data in the target business scenario. For example, assuming that the above-mentioned abnormal data identification method is applied to the field of games, in the identification scenario of plug-in game operation data included in the game operation data of large-scale user accounts, then the above-mentioned abnormal data is plug-in game operation data. Assuming that the above-mentioned abnormal data identification method is applied to the financial field, in the identification scenario of transaction data that has been manually intervened included in large-scale financial transaction data, then the above-mentioned abnormal data is transaction data that has been manually intervened.
[0050] Furthermore, the above-mentioned basic abnormal data can be obtained by, but is not limited to, at least one of the following methods: randomly obtaining data that has been determined to be abnormal data as basic abnormal data; generating basic abnormal data by a heuristic method.
[0051] Optionally, in this embodiment, the above abnormal affinity can be, but is not limited to, used to indicate the probability of data deviating from normal data. The higher the abnormal affinity, the higher the probability that the data is abnormal data. Specifically, the abnormal affinity of the above basic abnormal data can be, but is not limited to, obtained based on the following steps:
[0052] A=exp(-D)(1)
[0053] Among them, A is the affinity of the basic abnormal data, and D is the distance between the basic abnormal data and the normal data (such as Euclidean distance, cosine similarity, with a value range of [0, 1]). exp is an exponential function used to convert the distance to the anomaly affinity. It can be seen that the larger the distance D, the lower the anomaly affinity.
[0054] It should be noted that the above-mentioned abnormal database can be, but is not limited to, generated using a negative selection algorithm (NSA) and / or an immune network algorithm (INA). Among them, the processing logic of NSA includes: S1, generate detectors: create a set of random detectors. These detectors are designed to detect abnormal data. S2, define self: determine normal data (self). This is usually done by providing a set of normal data samples. S3, training: expose the detector to autologous samples. Any detector that matches the autologous sample will be eliminated. The remaining detectors are used to identify non-self (i.e., abnormal data). S4, detection: after the training phase is completed, use the detector to detect anomalies. If the detector matches the new sample, then the sample is marked as non-self (i.e., abnormal data). S5, update: by repeating the training phase, the detector set is regularly updated to adapt to changes in the definition of the self.
[0055] The processing logic of INA includes: S1, initialization: obtain the initialized abnormal data (i.e., basic abnormal data). S2, abnormal affinity acquisition: abnormal affinity of basic abnormal data. S3, selection processing: select the basic abnormal data to be cloned and mutated according to the abnormal affinity of the basic abnormal data. S4, cloning and mutation processing: clone mutation processing is performed on the basic abnormal data to be cloned and mutated. S5, abnormal affinity evaluation: evaluate the abnormal affinity of the abnormal data obtained after cloning and mutation of the basic abnormal data. The abnormal data with abnormal affinity greater than the predetermined threshold are retained, and the abnormal data with abnormal affinity that does not reach the predetermined threshold are eliminated. S6, network formation: a network is formed by connecting mutually recognized abnormal data to simulate the immune memory system. S7, inhibition and stimulation: suppress similar abnormal data to avoid redundancy, and stimulate diverse abnormal data to maintain diversity. S8, memory cell processing: select some abnormal data with higher abnormal affinity as memory cells. S9, update: regularly introduce new random abnormal data into the abnormal database.
[0056] Furthermore, the above-mentioned partial selective mutation of the basic abnormal data using abnormal affinity may include, but is not limited to, using at least one immune selection mode to perform partial selective mutation of the basic abnormal data using abnormal affinity of the basic abnormal data.
[0057] Optionally, in this embodiment, the potential abnormal data may be, but is not limited to, potential data indicating that it may be abnormal data, and the method of identifying the potential abnormal data from the business data using the abnormal database may include, but is not limited to: preprocessing the business data to obtain business features of a unified scale and represented by numerical values. Then, using the feature similarity between the business features and the abnormal data in the abnormal database, potential abnormal data that may be abnormal data is identified in the business data.
[0058] S206, based on the abnormal affinity corresponding to each potential abnormal data, determine at least one reference recognition threshold and a clustering evaluation coefficient that matches each reference recognition threshold, wherein the clustering evaluation coefficient is used to evaluate the result of clustering the potential abnormal data using the reference recognition threshold.
[0059] It should be noted that, in this embodiment, the method for obtaining the abnormal affinity corresponding to the potential abnormal data can refer to the method for obtaining the abnormal affinity of the basic abnormal data mentioned above, which will not be described in detail in this embodiment.
[0060] Furthermore, the clustering evaluation coefficient may be, but is not limited to, a coefficient for evaluating the quality of clustering processing, such as a silhouette coefficient. Furthermore, the above-mentioned determination of at least one reference identification threshold based on the abnormal affinity corresponding to each potential abnormal data, and the clustering evaluation coefficient that matches each reference identification threshold may include, but is not limited to: obtaining an affinity value interval based on the potential abnormal data, and then extracting a value in the value interval in turn as a reference threshold. Then, clustering the potential abnormal data in turn using each reference identification threshold to obtain an abnormal data cluster that matches the reference identification threshold. Thus, using at least an abnormal data cluster, a clustering evaluation coefficient that matches the reference identification threshold is determined.
[0061] S208, when a target recognition threshold is determined from at least one reference recognition threshold based on the clustering evaluation coefficient, potential abnormal data having an abnormal affinity greater than the target recognition threshold is identified as target abnormal data.
[0062] Optionally, in this embodiment, determining the target recognition threshold from at least one reference recognition threshold based on the clustering evaluation coefficient may include, but is not limited to: determining the reference recognition threshold with the largest corresponding clustering evaluation coefficient among the reference recognition thresholds as the above target recognition threshold.
[0063] Furthermore, after the potential abnormal data whose abnormal affinity is greater than the target identification threshold is identified as the target abnormal data, the above method may, but is not limited to, also include: obtaining neighboring normal data associated with the target abnormal data from the business data, wherein the neighboring normal data is normal data whose distance to the target abnormal data is less than the target distance threshold; and repairing the target abnormal data using the neighboring normal data.
[0064] As an optional implementation, assume that the above abnormal data identification method is applied in the financial field to identify transaction data that has been manually intervened in large-scale financial transaction data, so as to calculate the transaction reputation score of the user account by using the transaction data that is not considered to have been intervened. Figure 3 The following steps are shown to illustrate the above method by way of example:
[0065] Step S302, obtaining historical transaction data of the bank user account for which the transaction credit score is to be calculated.
[0066] Step S304, using the abnormal database storing abnormal transaction data that has been manually intervened, identifying potential abnormal transaction data from historical transaction business data;
[0067] Step S306, based on the abnormal affinity corresponding to each of the potential abnormal transaction data, determine at least one reference recognition threshold and a clustering evaluation coefficient that matches each reference recognition threshold, wherein the clustering evaluation coefficient is used to evaluate the result of clustering the potential abnormal transaction data using the reference recognition threshold.
[0068] Step S308, when a target identification threshold is determined from at least one reference identification threshold based on the clustering evaluation coefficient, potential abnormal transaction data having an abnormal affinity greater than the target identification threshold is identified as target abnormal transaction data.
[0069] Step S310, recovering the target abnormal transaction data included in the historical transaction business data, and obtaining the historical transaction data that has not been subject to any human intervention.
[0070] Step S312, using historical transaction data that has not been subject to human intervention, determine the corresponding transaction credit scores for the bank user accounts whose transaction credit scores are to be calculated.
[0071] As another optional implementation, assume that the above abnormal data identification method is applied in the field of gaming to identify the external game operation data included in the game operation data of a large number of user accounts, so as to calculate the game operation level of the user account using the game operation data without external game operation data. Figure 4 The following steps are shown to illustrate the above method by way of example:
[0072] Step S402, obtaining historical game operation data of the game user account whose game operation level is to be calculated.
[0073] Step S404, using the stored external game operation data executed by the external game, identifying potential external game operation data from the historical game operation data;
[0074] Step S406, based on the abnormal affinity corresponding to each potential cheating game operation data, determine at least one reference recognition threshold and a clustering evaluation coefficient that matches each reference recognition threshold, wherein the clustering evaluation coefficient is used to evaluate the result of clustering the potential cheating game operation data using the reference recognition threshold.
[0075] Step S408, when a target identification threshold is determined from at least one reference identification threshold based on the clustering evaluation coefficient, potential cheating game operation data having an abnormal affinity greater than the target identification threshold is identified as target cheating game operation data.
[0076] Step S410, clearing the target external game operation data included in the historical transaction business data, and acquiring the historical game operation data without external game operation data.
[0077] Step S412, using the historical game operation data that has not been subject to any human intervention, determine the corresponding game operation level scores for the game user accounts whose game operation levels are to be calculated.
[0078] In an embodiment of the present application, business data to be identified in a target business scenario is obtained. Then, potential abnormal data is identified from the business data using an abnormal database, wherein the abnormal data in the abnormal database is obtained after partial selection and variation of the basic abnormal data using abnormal affinity, and the abnormal affinity is used to indicate the affinity between the abnormal data and the normal data. Further, based on the abnormal affinity corresponding to each of the potential abnormal data, at least one reference recognition threshold is determined, and a clustering evaluation coefficient that matches each reference recognition threshold is determined, wherein the clustering evaluation coefficient is used to evaluate the result of clustering the potential abnormal data using the reference recognition threshold. Then, in the case where a target recognition threshold is determined from at least one reference recognition threshold based on the clustering evaluation coefficient, the potential abnormal data whose abnormal affinity is greater than the target recognition threshold is identified as the target abnormal data. In other words, using an embodiment of the present application, the potential abnormal data after being screened by the abnormal database is compared using an adaptive threshold to determine the target abnormal data from the business data. The purpose of effectively and accurately identifying large-scale data abnormalities is achieved, and it is no longer limited to marking data exceeding a fixed threshold as abnormal data. Thereby, the technical effect of improving the recognition accuracy of abnormal data is achieved, and the problem of low recognition accuracy in the recognition method of abnormal data provided by related technologies is solved.
[0079] Optionally, as an optional solution, based on the abnormal affinity corresponding to each potential abnormal data, at least one reference recognition threshold is determined, and the clustering evaluation coefficients respectively matched with each reference recognition threshold include:
[0080] S1, obtaining the maximum abnormal affinity and the minimum abnormal affinity from the abnormal affinities corresponding to the potential abnormal data.
[0081] S2, in the numerical range formed by the minimum abnormal affinity and the maximum abnormal affinity, a numerical value is sequentially extracted according to the target step size as a reference recognition threshold.
[0082] Optionally, in this embodiment, the target step length may be determined as the result of calculating the abnormal affinity corresponding to each potential abnormal data, but is not limited to the above. For example, half of the standard deviation of the abnormal affinity corresponding to each potential abnormal data is determined as the above target step length.
[0083] S3. Use each reference recognition threshold to perform clustering processing on the potential abnormal data in turn, and obtain at least two abnormal data clusters that match the reference recognition threshold.
[0084] It should be noted that using each of the above reference recognition thresholds to perform clustering processing on the potential abnormal data in turn to obtain at least two abnormal data clusters that match the reference recognition threshold may include but is not limited to: taking each of the reference recognition thresholds as the current reference recognition threshold in turn, and performing the following operations: comparing the abnormal affinity corresponding to each potential abnormal data with the current reference recognition threshold; clustering the potential abnormal data with an abnormal affinity greater than the current reference recognition threshold into a first abnormal data cluster, and clustering the potential abnormal data with an abnormal affinity less than or equal to the current reference recognition threshold into a second abnormal data cluster; determining the first abnormal data cluster and the second abnormal data cluster as the abnormal data clusters that match the current reference recognition threshold.
[0085] S4. Use the clustering distance between each abnormal data in at least two abnormal data clusters to determine a clustering evaluation coefficient that matches the reference recognition threshold.
[0086] Optionally, the clustering distance between each abnormal data in the above at least two abnormal data clusters may include but is not limited to the average distance between each abnormal data in the abnormal data cluster and other abnormal data in the abnormal data cluster.
[0087] Furthermore, using the clustering distance between each abnormal data in at least two abnormal data clusters to determine a clustering evaluation coefficient that matches the reference recognition threshold may include but is not limited to: in the case of obtaining the clustering distance between each abnormal data in the above at least two abnormal data clusters, performing a weighted sum calculation on the clustering distance between each abnormal data in the above at least two abnormal data clusters to obtain the clustering evaluation coefficient corresponding to the reference recognition threshold of the above at least two abnormal data clusters.
[0088] As an optional implementation manner, taking the above potential abnormal data including potential abnormal data A, potential abnormal data B, potential abnormal data C, potential abnormal data D, and potential abnormal data E as an example, assuming that the abnormal affinity corresponding to potential abnormal data A is x + c, assuming that the abnormal affinity corresponding to potential abnormal data B is x + 2c, assuming that the abnormal affinity corresponding to potential abnormal data C is x + 3c, assuming that the abnormal affinity corresponding to potential abnormal data D is x + 4c, and assuming that the abnormal affinity corresponding to potential abnormal data E is x + 5c. Assuming that the above target step size is 2c, where both x and z are natural numbers. The above method is explained by the following steps:
[0089] Such as Figure 5As shown in (a), from the abnormal affinities x+c, x+2c, x+3c, x+4c, and x+5c corresponding to potential abnormal data A, potential abnormal data B, potential abnormal data C, potential abnormal data D, and potential abnormal data E, the smallest abnormal affinity is determined to be the abnormal affinity x+c corresponding to abnormal data A, and the largest abnormal affinity is determined to be the abnormal affinity x+5c corresponding to abnormal data E.
[0090] Then, if Figure 5 As shown in (b), the numerical interval [x+c, x+2c, x+3c, x+4c, x+5c] is determined based on the abnormal affinity x+c corresponding to the abnormal data A and the abnormal affinity x+5c corresponding to the abnormal data E, so that x+c, x+3c, and x+5c are determined as reference recognition thresholds in turn according to the target step size 2c.
[0091] Then, the potential abnormal data are clustered in turn using each reference recognition threshold to obtain at least two abnormal data clusters that match the reference recognition threshold. Furthermore, the average distance between each abnormal data in the at least two abnormal data clusters and other abnormal data in the abnormal data cluster other than itself is used to determine the clustering evaluation coefficient that matches the reference recognition threshold.
[0092] It should be noted that in this embodiment, in order to facilitate understanding, it is assumed that the potential abnormal data includes: abnormal data A, potential abnormal data B, potential abnormal data C, potential abnormal data D, and potential abnormal data E, but in fact the number of the potential abnormal data is not limited to this, and no limitation is made to this in this embodiment.
[0093] In the embodiment of the present application, the maximum abnormal affinity and the minimum abnormal affinity are obtained from the abnormal affinity corresponding to each potential abnormal data. Then, in the numerical interval formed by the minimum abnormal affinity and the maximum abnormal affinity, a numerical value is sequentially extracted according to the target step length as a reference recognition threshold. Then, the potential abnormal data is clustered in turn using each reference recognition threshold to obtain at least two abnormal data clusters that match the reference recognition threshold. Thus, the clustering distance between each abnormal data in at least two abnormal data clusters is used to determine the clustering evaluation coefficient that matches the reference recognition threshold. In other words, using the embodiment of the present application, the abnormal affinity corresponding to each potential abnormal data is used to determine the numerical interval formed by the minimum abnormal affinity and the maximum abnormal affinity. Then, the reference recognition threshold is determined using the numerical interval, and then the matching clustering evaluation coefficient is determined for each reference recognition threshold. Then, the target recognition threshold adapted to the recognition abnormal data is determined using the clustering evaluation coefficients matched by each reference recognition threshold, and the target recognition threshold is used to determine the target abnormal data from the potential abnormal data. Thereby, the technical effect of improving the recognition accuracy of abnormal data is achieved, and the problem of low recognition accuracy in the recognition method of abnormal data provided by related technologies is solved.
[0094] Optionally, as an optional solution, clustering is performed on the potential abnormal data in sequence using each reference recognition threshold, and obtaining at least two abnormal data clusters matching the reference recognition threshold includes:
[0095] Take each reference recognition threshold as the current reference recognition threshold in turn, and perform the following operations:
[0096] The anomaly affinity corresponding to each potential anomaly data is compared with the current reference recognition threshold.
[0097] The potential abnormal data with an abnormal affinity greater than the current reference recognition threshold are clustered into a first abnormal data cluster, and the potential abnormal data with an abnormal affinity less than or equal to the current reference recognition threshold are clustered into a second abnormal data cluster.
[0098] The first abnormal data cluster and the second abnormal data cluster are determined as abnormal data clusters that match the current reference recognition threshold.
[0099] As an optional implementation, still taking the above potential abnormal data including potential abnormal data A, potential abnormal data B, potential abnormal data C, potential abnormal data D, and potential abnormal data E as an example, it is assumed that the abnormal affinity corresponding to the potential abnormal data A is x+c, the abnormal affinity corresponding to the potential abnormal data B is x+2c, the abnormal affinity corresponding to the potential abnormal data C is x+3c, the abnormal affinity corresponding to the potential abnormal data D is x+4c, and the abnormal affinity corresponding to the potential abnormal data E is x+5c. Assume that the above target step size is 2c, where x and z are both natural numbers. The above method is explained by the following steps:
[0100] From the abnormal affinities x+c, x+2c, x+3c, x+4c, and x+5c corresponding to potential abnormal data A, potential abnormal data B, potential abnormal data C, potential abnormal data D, and potential abnormal data E, the smallest abnormal affinity is determined to be the abnormal affinity x+c corresponding to abnormal data A, and the largest abnormal affinity is determined to be the abnormal affinity x+5c corresponding to abnormal data E.
[0101] Next, based on the abnormal affinity x+c corresponding to the abnormal data A and the abnormal affinity x+5c corresponding to the abnormal data E, the numerical interval [x+c, x+2c, x+3c, x+4c, x+5c] is determined, so that x+c, x+3c, and x+5c are determined as reference recognition thresholds in turn according to the target step size 2c.
[0102] Further, x+c in the reference recognition threshold is used as the current reference recognition threshold, and the following operations are performed: the abnormal affinity x+c corresponding to the abnormal data A, the abnormal affinity x+2c corresponding to the potential abnormal data B, the abnormal affinity x+3c corresponding to the potential abnormal data C, the abnormal affinity x+4c corresponding to the potential abnormal data D, and the abnormal affinity x+5c corresponding to the potential abnormal data E are compared with the current reference recognition threshold. The potential abnormal data (i.e., potential abnormal data B, potential abnormal data C, potential abnormal data D, potential abnormal data E) corresponding to the abnormal affinity greater than x+c (i.e., x+2c, x+3c, x+4c, x+5c) are clustered into a first abnormal data cluster. Then, the first abnormal data cluster is determined as the abnormal data cluster that matches the current reference recognition threshold x+c.
[0103] Then, x+3c in the reference recognition threshold is used as the current reference recognition threshold, and the following operations are performed: the abnormal affinity x+c corresponding to the abnormal data A, the abnormal affinity x+2c corresponding to the potential abnormal data B, the abnormal affinity x+3c corresponding to the potential abnormal data C, the abnormal affinity x+4c corresponding to the potential abnormal data D, and the abnormal affinity x+5c corresponding to the potential abnormal data E are compared with the current reference recognition threshold. The potential abnormal data (i.e., potential abnormal data D, potential abnormal data E) corresponding to the abnormal affinity greater than x+3c (i.e., x+4c, x+5c) are clustered into the first abnormal data cluster. The potential abnormal data (i.e., potential abnormal data A, potential abnormal data B) corresponding to the abnormal affinity less than x+3c (i.e., x+c, x+2c) are clustered into the second abnormal data cluster. Then, the first abnormal data cluster and the second type of data cluster are determined as abnormal data clusters that match the current reference recognition threshold x+3c.
[0104] Further, x+5c in the reference recognition threshold is used as the current reference recognition threshold, and the abnormal data cluster matching the reference recognition threshold x+5c is determined. For the specific determination steps, please refer to the processing steps of other reference recognition thresholds above, which will not be repeated in this embodiment.
[0105] It should be noted that in this embodiment, in order to facilitate understanding, it is assumed that the potential abnormal data includes: abnormal data A, potential abnormal data B, potential abnormal data C, potential abnormal data D, and potential abnormal data E, but in fact the number of the potential abnormal data is not limited to this, and no limitation is made to this in this embodiment.
[0106] In an embodiment of the present application, each reference recognition threshold is used as the current reference recognition threshold in turn, and the following operations are performed: the abnormal affinity corresponding to each potential abnormal data is compared with the current reference recognition threshold. Then, the potential abnormal data with an abnormal affinity greater than the current reference recognition threshold are clustered into a first abnormal data cluster, and the potential abnormal data with an abnormal affinity less than or equal to the current reference recognition threshold are clustered into a second abnormal data cluster. Furthermore, the first abnormal data cluster and the second abnormal data cluster are determined as abnormal data clusters that match the current reference recognition threshold. In other words, by adopting an embodiment of the present application, by determining a target recognition threshold that is adaptive to the potential abnormal data from multiple reference recognition thresholds, the target abnormal data determined from the potential abnormal data using the target recognition threshold is more accurate. In other words, by adopting an embodiment of the present application, the technical effect of improving the recognition accuracy of abnormal data is achieved, and the problem of low recognition accuracy in the abnormal data recognition method provided by the related art is solved.
[0107] Optionally, as an optional solution, using the clustering distance between each abnormal data in at least two abnormal data clusters to determine the clustering evaluation coefficient matching the reference identification threshold includes:
[0108] Each abnormal data in at least two abnormal data clusters is used as the current abnormal data in turn, and the following operations are performed:
[0109] Determine the current abnormal data cluster where the current abnormal data is located.
[0110] A first average clustering distance between the current abnormal data and each first reference abnormal data included in the current abnormal data cluster is obtained.
[0111] A second average clustering distance between the current abnormal data and each second reference abnormal data included in other abnormal data clusters outside the current abnormal data cluster is obtained.
[0112] The first average clustering distance and the second average clustering distance are used to determine a current object clustering evaluation coefficient that matches the current abnormal data.
[0113] When the object clustering evaluation coefficient corresponding to each abnormal data in at least two abnormal data clusters is obtained, a weighted sum calculation is performed on all the object clustering evaluation coefficients to obtain a clustering evaluation coefficient that matches the reference recognition threshold.
[0114] It should be noted that, in the embodiment of the present application, the first reference abnormal data may be used, but not limited to, to indicate all other potential abnormal data contained in the current abnormal data cluster except the current abnormal data. Further, the first average cluster distance may be used, but not limited to, to indicate that the current abnormal data is divided into the average distance between all other potential abnormal data contained in the previous abnormal data cluster except the current abnormal data. Wherein, the above distance may be, but not limited to, the Euclidean distance between the current abnormal data and all other potential abnormal data contained in the previous abnormal data cluster except the current abnormal data, or the cosine similarity, and the value range is [0, 1].
[0115] Further, the second average cluster distance may be, but is not limited to, used to indicate the current abnormal data, and is divided into the average distance between each second reference abnormal data contained in other abnormal data clusters outside the current abnormal data cluster. Wherein, the above distance may be, but is not limited to, the current abnormal data, and is divided into the Euclidean distance between each second reference abnormal data contained in other abnormal data clusters outside the current abnormal data cluster, or the cosine similarity, and the value range is [0, 1].
[0116] Optionally, in this embodiment, the following steps may be used, but are not limited to, to determine a current object clustering evaluation coefficient that matches the current abnormal data using the first average clustering distance and the second average clustering distance:
[0117] S(i)=(b(i)-a(i)) / max(a(i), b(i))(2)
[0118] Wherein, i is the current abnormal data, a(i) is the above-mentioned first average clustering distance, b(i) is the above-mentioned second average clustering distance, and S(i) is the current object clustering evaluation coefficient.
[0119] As an optional implementation, Figure 6 As shown, it is assumed that the current reference identification threshold corresponds to two abnormal data clusters (i.e., a first abnormal data cluster and a second abnormal data cluster), wherein the abnormal data included in the first abnormal data cluster are: abnormal data D, abnormal data E, abnormal data F, and the abnormal data included in the second abnormal data cluster are: abnormal data A, abnormal data B, abnormal data C. The above method is explained by the following steps:
[0120] The abnormal data D is determined as the current abnormal data, and the first abnormal data cluster where the current abnormal data is located is determined as the current abnormal data cluster.
[0121] Then, still Figure 6 As shown, the Euclidean distance d1 between the abnormal data D and the abnormal data E is obtained, and the Euclidean distance d2 between the abnormal data D and the abnormal data F is obtained. Then, the average value between d1 and d2 is determined as the first average cluster distance.
[0122] Then, still Figure 6 As shown, the Euclidean distance d3 between the abnormal data D and the abnormal data A included in the second abnormal data cluster is obtained, the Euclidean distance d4 between the abnormal data D and the abnormal data B included in the second abnormal data cluster is obtained, and the Euclidean distance d5 between the abnormal data D and the abnormal data C included in the second abnormal data cluster is obtained. Then, the average value among d3, d4, and d5 is determined as the second average cluster distance.
[0123] Next, the clustering evaluation coefficient of the current object matching the abnormal data D is determined by using the first average clustering distance and the second average clustering distance.
[0124] Next, referring to the method of determining the clustering evaluation coefficient for the abnormal data D, the corresponding clustering evaluation coefficients are determined for the abnormal data E, the abnormal data F, the abnormal data A, the abnormal data B, and the abnormal data C respectively.
[0125] When the clustering evaluation coefficients corresponding to abnormal data D, abnormal data E, abnormal data F, abnormal data A, abnormal data B, and abnormal data C are obtained, a weighted sum calculation is performed on the clustering evaluation coefficients corresponding to abnormal data D, abnormal data E, abnormal data F, abnormal data A, abnormal data B, and abnormal data C to obtain a clustering evaluation coefficient that matches the reference recognition threshold.
[0126] In an embodiment of the present application, the current abnormal data cluster where the current abnormal data is located is determined. Then, the first average cluster distance between the current abnormal data and each first reference abnormal data contained in the current abnormal data cluster is obtained. Next, the second average cluster distance between the current abnormal data and each second reference abnormal data contained in other abnormal data clusters outside the current abnormal data cluster is obtained. Further, the current object cluster evaluation coefficient matching the current abnormal data is determined using the first average cluster distance and the second average cluster distance. Furthermore, in the case where the object cluster evaluation coefficient corresponding to each abnormal data in at least two abnormal data clusters is obtained, the weighted sum calculation is performed on all the object cluster evaluation coefficients to obtain the cluster evaluation coefficient matching the reference recognition threshold. Thereby, the purpose of improving the reliability of the cluster evaluation coefficient is achieved, so as to further improve the accuracy of the target recognition threshold determined based on the cluster evaluation coefficient. Thereby, the technical effect of improving the recognition accuracy of abnormal data is achieved, and the problem of low recognition accuracy in the recognition method of abnormal data provided by the related art is solved.
[0127] Optionally, as an optional solution, after determining at least one reference recognition threshold and a clustering evaluation coefficient respectively matching each reference recognition threshold based on the abnormal affinity corresponding to each potential abnormal data, the method further includes:
[0128] The clustering evaluation coefficients that match each reference recognition threshold are ranked.
[0129] The reference recognition threshold corresponding to the maximum clustering evaluation coefficient is determined as the target recognition threshold.
[0130] As an alternative embodiment, assume that the above-mentioned respective reference recognition thresholds include: reference recognition threshold A, reference recognition threshold B, and reference recognition threshold C. The above method is illustrated and explained by the following steps: Determine that the clustering evaluation coefficient corresponding to the reference recognition threshold A is y, the clustering evaluation coefficient corresponding to the reference recognition threshold B is y + d, and the clustering evaluation coefficient corresponding to the reference recognition threshold C is y + 2d, where both y and d are natural numbers, and the value ranges of y, y + d, and y + 2d are [0, 1]. Sort the clustering evaluation coefficient y corresponding to the reference recognition threshold A, the clustering evaluation coefficient y + d corresponding to the reference recognition threshold B, and the clustering evaluation coefficient y + 2d corresponding to the reference recognition threshold C, so as to determine the reference recognition threshold C corresponding to the maximum clustering evaluation coefficient y + 2d as the target recognition threshold.
[0131] In the embodiments of the present application, the clustering evaluation coefficients respectively matched with the respective reference recognition thresholds are sorted. Then, the reference recognition threshold corresponding to the maximum clustering evaluation coefficient is determined as the target recognition threshold. In other words, by adopting the embodiments of the present application, the target recognition threshold adapted to the potential abnormal data is determined from multiple reference recognition thresholds, so that the target abnormal data determined from the potential abnormal data using the target recognition threshold is more accurate. In other words, by adopting the embodiments of the present application, the technical effect of improving the recognition accuracy of abnormal data is achieved, and the problem that the recognition method of abnormal data provided by the related art has low recognition accuracy is solved.
[0132] Optionally, as an alternative solution, before obtaining the service data to be recognized in the target service scenario, it further includes:
[0133] S1, obtain basic abnormal data.
[0134] Optionally, in this embodiment, the above basic abnormal data can be generated in a heuristic manner, but not limited to this. Specifically, the generation of the above basic abnormal data in a heuristic manner can include, but not limited to, at least one of the following:
[0135] 1) Random generation method: Randomly generate basic abnormal data in the feature space of the data.
[0136] 2) Clustering analysis generation method: Perform clustering analysis on the data, and then select the samples that are far from the cluster center from each cluster as the basic abnormal data.
[0137] 3) Rule-based generation method: Define rules according to business logic or statistical characteristics, and select the samples that meet these rules from the data set as the basic abnormal data.
[0138] 4) Sampling generation method: Use specific sampling methods, such as Synthetic Minority Over-sampling Technique (SMOTE for short), Adaptive Synthetic Sampling Approach for Imbalanced Learning (ADASYN for short), etc., to generate basic abnormal data. These methods are usually used to process imbalanced data sets and can enhance the prediction ability of the model by generating synthetic minority class samples.
[0139] 5) Abnormal detection generation method: Use its abnormal detection algorithms, such as (e.g., Isolation Forest, Local Outlier Factor, etc.) as a preprocessing step to identify basic abnormal data.
[0140] S2. Obtain extended abnormal data other than the basic abnormal data through at least one immune selection mode, where the immune selection mode is used to select abnormal data similar to the basic abnormal data from the data by using an immune algorithm.
[0141] It should be noted that the above at least one immune selection mode can but is not limited to including the first immune selection mode and the second immune selection mode. Among them, in the first immune selection mode, the abnormal data obtained by cloning and replicating the basic abnormal data and then performing mutation processing can be but is not limited to being determined as the candidate extended abnormal data in the first immune selection mode. In the second immune selection mode, the abnormal data identified by using the target detector for detecting abnormal data obtained by training a group of detectors exposed to normal data can be but is not limited to being determined as the candidate extended abnormal data in the second immune selection mode.
[0142] Furthermore, the obtaining of extended abnormal data other than the basic abnormal data through at least one immune selection mode includes: obtaining the candidate extended abnormal data determined by each immune selection mode respectively. Then, perform the following operation on the candidate extended abnormal data to determine the extended abnormal data: perform an average calculation on the candidate extended abnormal data determined by each immune selection mode respectively to obtain the extended abnormal data; or perform a weighted summation calculation on the candidate extended abnormal data determined by each immune selection mode respectively to obtain the extended abnormal data; or perform a voting process on the candidate extended abnormal data determined by each immune selection mode respectively to obtain the extended abnormal data.
[0143] S3. Update the abnormal database by using the extended abnormal data.
[0144] Optionally, the above-mentioned updating of the anomaly database using the extended anomaly data may include, but is not limited to: continuously iteratively updating the anomaly database using the extended anomaly data until the average value between the anomaly affinities corresponding to the anomaly data included in the anomaly database reaches a preset threshold value N times in a row, and within the above-mentioned N consecutive times, the difference between the above-mentioned average values does not exceed a second preset threshold value, wherein the above-mentioned N is a positive integer greater than 2.
[0145] In an embodiment of the present application, basic abnormal data is obtained. Then, extended abnormal data other than the basic abnormal data is obtained through at least one immune selection mode, wherein the immune selection mode is used to select abnormal data similar to the basic abnormal data from the data using an immune algorithm. Furthermore, the abnormal database is updated using the extended abnormal data. That is to say, by adopting an embodiment of the present application, the richness of the abnormal database is improved by adopting multiple immune selection modes to expand the basic abnormal data included in the abnormal database. Furthermore, it makes it more reliable to identify potential abnormal data from business data using the abnormal database, thereby improving the accuracy of the target abnormal data identified in the potential abnormal data. In other words, by adopting an embodiment of the present application, the technical effect of improving the recognition accuracy of abnormal data is achieved, and the problem of low recognition accuracy in the abnormal data recognition method provided by the related art is solved.
[0146] Optionally, as an optional solution, obtaining extended abnormal data in addition to basic abnormal data through at least one immune selection mode includes:
[0147] Obtaining candidate expansion anomaly data determined by each immune selection mode;
[0148] Perform one of the following operations on the candidate extended abnormal data to determine the extended abnormal data:
[0149] The candidate extended abnormal data determined by each immune selection mode are averaged to obtain the extended abnormal data; or
[0150] Performing weighted sum calculation on the candidate extended abnormal data determined by each immune selection mode to obtain extended abnormal data; or
[0151] The candidate extended abnormal data determined by each immune selection mode are voted to obtain the extended abnormal data.
[0152] It should be noted that the above-mentioned voting process for the candidate extended abnormal data determined by each immune selection mode may include but is not limited to: inputting the candidate extended abnormal data determined by each immune selection mode into different algorithms or models for processing, and obtaining the results of each algorithm or model. Then, according to the pre-set rules, the results of different algorithms or models are voted, and majority voting, weighted voting, etc. can be used. The voting results will determine the final output. Thus, based on the voting results, the extended abnormal data is obtained.
[0153] As an optional implementation, it is assumed that the above-mentioned immune selection mode includes a first immune selection mode and a second immune selection mode, wherein the candidate extended abnormal data corresponding to the first immune selection mode is the candidate extended abnormal data A, and the candidate extended abnormal data corresponding to the second immune selection mode is the candidate extended abnormal data B. The above-mentioned method is explained by the following steps:
[0154] The candidate extended abnormal data A corresponding to the first immune selection mode and the candidate extended abnormal data B corresponding to the second immune selection mode are obtained. Then, one of the following operations is performed on the candidate extended abnormal data A and the candidate extended abnormal data B to determine the extended abnormal data: average calculation is performed on the candidate extended abnormal data A and the candidate extended abnormal data B to obtain the extended abnormal data; or weighted sum calculation is performed on the candidate extended abnormal data A and the candidate extended abnormal data B to obtain the extended abnormal data; or voting processing is performed on the candidate extended abnormal data A and the candidate extended abnormal data B to obtain the extended abnormal data.
[0155] In an embodiment of the present application, the candidate extended abnormal data determined by each immune selection mode is obtained. Then, one of the following operations is performed on the candidate extended abnormal data to determine the extended abnormal data: average calculation is performed on the candidate extended abnormal data determined by each immune selection mode to obtain the extended abnormal data; or weighted sum calculation is performed on the candidate extended abnormal data determined by each immune selection mode to obtain the extended abnormal data; or voting is performed on the candidate extended abnormal data determined by each immune selection mode to obtain the extended abnormal data. That is to say, by adopting the embodiment of the present application, the method of expanding the abnormal data included in the abnormal database by using the candidate extended abnormal number determined by multiple immune selection modes is used to improve the richness of the abnormal database. Furthermore, it makes it more reliable to identify potential abnormal data from business data using the abnormal database, thereby improving the accuracy of the target abnormal data identified in the potential abnormal data. In other words, by adopting the embodiment of the present application, the technical effect of improving the recognition accuracy of abnormal data is achieved, and the problem of low recognition accuracy in the abnormal data recognition method provided by the related technology is solved.
[0156] Optionally, as an optional solution, before obtaining extended abnormal data other than basic abnormal data through at least one immune selection mode, the method further includes:
[0157] In the first immune selection mode, the abnormal affinity and clonal threshold corresponding to each of the basic abnormal data are compared.
[0158] The basic abnormal data whose abnormal affinity is greater than the cloning threshold is cloned to obtain a copy of the abnormal data that matches the basic abnormal data.
[0159] When the abnormal affinity of the mutated abnormal data is greater than the abnormal affinity of the abnormal data copy, the mutated abnormal data is used as the candidate extended abnormal data determined in the first immune selection mode.
[0160] It should be noted that in this embodiment, the clonal selection algorithm can be used, but is not limited to, to clone and copy the basic abnormal data whose abnormal affinity is greater than the cloning threshold to obtain a copy of the abnormal data that matches the basic abnormal data. Specifically, the clonal selection algorithm is a basic immune algorithm that simulates the clonal selection mechanism in the biological immune system.
[0161] The abnormal data copy is mutated to obtain mutated abnormal data.
[0162] Optionally, the above-mentioned mutation processing of the abnormal data copy to obtain the mutated abnormal data may include, but is not limited to: performing Gaussian mutation on the abnormal data copy to obtain the mutated abnormal data. Specifically, as shown in the following formula:
[0163] x′=x+σ*N(0,1) (3)
[0164] Among them, x is the copy of the abnormal data mentioned above, x ′ is the abnormal data of mutation, σ is the standard deviation of the abnormal data copy, and N(0,1) is a random number that obeys the standard normal distribution. It should be noted that in this embodiment, σ is an important parameter that determines the strength of the mutation. The larger the σ value, the greater the strength of the mutation, and the smaller the σ value, the smaller the strength of the mutation.
[0165] As an optional implementation, Figure 7 The following steps are shown to illustrate the above method by way of example:
[0166] Steps S702 to S704 are executed to classify each basic abnormal data included in the basic abnormal data as current basic abnormal data. It is determined whether the abnormal affinity corresponding to the current basic abnormal data is greater than the cloning threshold.
[0167] Furthermore, when the abnormal affinity corresponding to the current basic abnormal data is less than or equal to the cloning threshold, step S702 is executed again to determine the next basic abnormal data as the current basic abnormal data. When the abnormal affinity corresponding to the current basic abnormal data is greater than the cloning threshold, step S706 is executed to clone the current basic abnormal data to obtain a copy of the current abnormal data that matches the basic abnormal data. Furthermore, the abnormal data copy is mutated to obtain the current mutated abnormal data.
[0168] Then, step S708 is performed to determine whether the abnormal affinity of the current variant abnormal data is greater than the abnormal affinity of the current abnormal data copy. If the abnormal affinity of the current variant abnormal data is less than or equal to the abnormal affinity of the current abnormal data copy, step S702 is performed again. If the abnormal affinity of the current variant abnormal data is greater than the abnormal affinity of the current abnormal data copy, step S710 is performed to add the current variant abnormal data to the candidate extended abnormal data set under the first immune selection mode.
[0169] Then, step S712 is executed to determine whether the current basic abnormal data is the last basic abnormal data. If the current basic abnormal data is not the last basic abnormal data, step S702 is executed again. If the current basic abnormal data is the last basic abnormal data, step S714 is executed to determine the abnormal data included in the candidate extended abnormal data set as the candidate extended abnormal data in the first immune selection mode.
[0170] In an embodiment of the present application, in the first immune selection mode, the abnormal affinity and clone threshold corresponding to each of the basic abnormal data are compared. Then, the basic abnormal data whose abnormal affinity is greater than the clone threshold are cloned and copied to obtain a copy of the abnormal data that matches the basic abnormal data. Next, the copy of the abnormal data is mutated to obtain the mutated abnormal data. Furthermore, in the case where the abnormal affinity of the mutated abnormal data is greater than the abnormal affinity of the copy of the abnormal data, the mutated abnormal data is used as the candidate extended abnormal data determined under the first immune selection mode. In other words, in this embodiment, the candidate extended abnormal data under the first immune selection mode is determined by cloning and copying the basic abnormal data and mutating it, and then the extended data is used to expand the abnormal data included in the abnormal database, thereby improving the richness of the abnormal database.
[0171] Optionally, as an optional solution, before obtaining extended abnormal data other than basic abnormal data through at least one immune selection mode, the method further includes:
[0172] In the second immune selection mode, a group of detectors are exposed to normal data for training to obtain target detectors for detecting abnormal data, wherein the detectors that match the normal data during the training process will be discarded, while the detectors that do not match the normal data will be retained.
[0173] The data identified by the target detector is used as candidate extended abnormal data determined in the second immune selection mode.
[0174] It should be noted that the above group of detectors may be, but is not limited to, used to indicate a group of random detectors. The above detectors that match normal data may be, but are not limited to, used to indicate detectors used to identify normal data. The above detectors that do not match normal data may be, but are not limited to, used to indicate detectors used to identify abnormal data.
[0175] Optionally, in this embodiment, before the data identified by the target detector is used as the candidate extended abnormal data determined in the second immune selection mode, the method may, but is not limited to, further include: comparing the abnormal affinity and the cloning threshold corresponding to each of the basic abnormal data. The basic abnormal data with an abnormal affinity greater than the cloning threshold are cloned to obtain a copy of the abnormal data that matches the basic abnormal data.
[0176] Furthermore, the above-mentioned use of the data identified by the target detector as candidate extended abnormal data determined in the second immune selection mode may include but is not limited to: using the data identified by the target detector in the abnormal data copy as candidate extended abnormal data determined in the second immune selection mode.
[0177] In an embodiment of the present application, in the second immune selection mode, a group of detectors are exposed to normal data for training to obtain a target detector for detecting abnormal data, wherein the detectors that match the normal data during the training process will be discarded, while the detectors that do not match the normal data will be retained. Then, the data identified by the target detector is used as the candidate extended abnormal data determined in the second immune selection mode. In other words, in an embodiment of the present application, the candidate extended abnormal data in the second immune selection mode is identified by the target detector for detecting abnormal data, thereby using the candidate extended abnormal data pairs to expand the abnormal data included in the abnormal database, thereby improving the richness of the abnormal database.
[0178] Optionally, as an optional solution, it is characterized in that identifying potential abnormal data from business data using an abnormal database includes:
[0179] The business data is cleaned to obtain processed business data.
[0180] The processed business data is feature encoded to obtain business features.
[0181] The business characteristics are normalized to obtain business characteristics of a unified scale.
[0182] Obtain feature similarity between business features and data features of abnormal data in an abnormal database.
[0183] The business data corresponding to the business features whose feature similarity is greater than the target threshold is identified as potential abnormal data.
[0184] It should be noted that the above-mentioned cleaning process of the business data may include, but is not limited to: removing duplicates and removing missing values from the above-mentioned business data. Specifically, the above-mentioned removal of duplicates may include, but is not limited to: removing duplicates from the repeated business data included in the business data. The above-mentioned removal of missing values may include, but is not limited to: filling the missing values included in the business data with the average or median, using the interpolation method for estimation, or using a machine learning model to predict the filling of missing values.
[0185] Optionally, in this embodiment, the processed business data may be feature encoded based on, but not limited to, at least one of the following methods: 1) One-hot encoding: One-hot encoding refers to splitting a discrete feature into multiple binary features. 2) Label encoding: Label encoding refers to representing a discrete feature with a set of integers.
[0186] Furthermore, in this embodiment, the service characteristics may be normalized based on, but not limited to, at least one of the following methods:
[0187] Min-max Normalization: Min-max normalization refers to scaling the features to the range [0, 1]. The specific formula is as follows:
[0188]
[0189] Among them, x is the business characteristic, and x′ is the business characteristic of the unified scale.
[0190] Z-score normalization: Z-score normalization refers to scaling the features to a normal distribution interval with a mean of 0 and a variance of 1. The specific formula is as follows:
[0191]
[0192] Among them, x is the business characteristic, μ and σ are the mean and standard deviation of the business characteristic respectively, and x′ is the business characteristic of unified scale.
[0193] In an embodiment of the present application, the business data is cleaned to obtain processed business data. Then, the processed business data is feature encoded to obtain business features. Then, the business features are normalized to obtain business features of a unified scale. Next, the feature similarity between the business features and the data features of the abnormal data in the abnormal database is obtained. Then, the business data corresponding to the business features whose feature similarity is greater than the target threshold is identified as potential abnormal data. In other words, using an embodiment of the present application, by performing data preprocessing methods such as cleaning, feature encoding, and normalization on the business data, the business data can be directly identified and processed in the subsequent processing process. Furthermore, the data features of the abnormal data in the normal database are used to identify the preprocessed business data to preliminarily obtain potential abnormal data. Thereby, the subsequent identification of target abnormal data is more efficient, thereby achieving the technical effect of improving the identification efficiency of abnormal data.
[0194] Optionally, as an optional solution, it is characterized in that after identifying the potential abnormal data with an abnormal affinity greater than the target identification threshold as the target abnormal data, it also includes:
[0195] The neighboring normal data associated with the target abnormal data is obtained from the business data, wherein the neighboring normal data is normal data whose distance to the target abnormal data is less than a target distance threshold.
[0196] Use the neighboring normal data to repair the target abnormal data.
[0197] It should be noted that the above-mentioned neighboring normal data associated with the target abnormal data obtained from the business data can be but is not limited to being used to indicate normal data whose distance (such as Euclidean distance, Manhattan distance, etc.) to the target abnormal data is less than the target distance threshold.
[0198] Furthermore, the above-mentioned repairing of target abnormal data by using neighboring normal data may include but is not limited to at least one of the following:
[0199] The average value of K neighbor normal data is obtained, and the average value of the neighbor normal data is determined as the target repair value. Then, the target abnormal data is replaced with the target repair value; or the median of the neighbor normal data is obtained, and the median of the neighbor normal data is determined as the target repair value. Then, the target abnormal data is replaced with the target repair value; wherein the above K is an integer greater than 2.
[0200] In an embodiment of the present application, neighboring normal data associated with target abnormal data is obtained from business data, wherein the neighboring normal data is normal data whose distance from the target abnormal data is less than the target distance threshold. Then, the target abnormal data is repaired using the neighboring normal data. Thus, the purpose of repairing the target abnormal data is achieved, thereby achieving the technical effect of improving the availability of business data.
[0201] Optionally, as an optional implementation, by Figure 8 The following steps are used to illustrate the above abnormal data identification method as a whole:
[0202] Step S802, constructing an abnormality database;
[0203] Specifically, in this embodiment, the abnormality database can be constructed by using the following steps but is not limited to:
[0204] S802-1, generate basic abnormal data in a heuristic manner. The above-mentioned generation of the above-mentioned basic abnormal data in a heuristic manner may include but is not limited to at least one of the following: 1) Random generation method: randomly generate basic abnormal data in the feature space of the data; 2) Cluster analysis generation method: cluster analysis is performed on the data, and then samples that are far away from the cluster center are selected from each cluster as basic abnormal data. 3) Rule-based generation method: define rules according to business logic or statistical characteristics, and select samples that meet these rules from the data set as basic abnormal data. 4) Sampling generation method: use specific sampling methods, such as Synthetic Minority Over-sampling Technique (SMOTE), Adaptive Synthetic Sampling Approach for Imbalanced Learning (ADASYN), etc., to generate basic abnormal data. These methods are usually used to process unbalanced data sets, and can enhance the prediction ability of the model by generating synthetic minority class samples.
[0205] S802-2, obtaining the candidate extended abnormal data determined by each immune selection mode. Then, performing one of the following operations on the candidate extended abnormal data to determine the extended abnormal data: averaging the candidate extended abnormal data determined by each immune selection mode to obtain the extended abnormal data; or performing weighted sum calculation on the candidate extended abnormal data determined by each immune selection mode to obtain the extended abnormal data; or performing voting processing on the candidate extended abnormal data determined by each immune selection mode to obtain the extended abnormal data.
[0206] S802-3, continuously iteratively update the anomaly database using the extended anomaly data until the average value of the anomaly affinities corresponding to the anomaly data included in the anomaly database reaches a preset threshold value N times in a row, and within the above N consecutive times, the difference between the above average values does not exceed a second preset threshold value, wherein the above N is a positive integer greater than 2.
[0207] Step S804, obtaining the business data to be identified in the target business scenario;
[0208] Step S806, performing pre-processing operations on the business data;
[0209] Specifically, the business data is preprocessed through the following steps:
[0210] S806-1, perform deduplication processing and missing value processing on the above business data. Specifically, the above deduplication processing may include, but is not limited to: deduplication processing of repeated business data included in the business data. The above missing value processing may include, but is not limited to: filling the missing values included in the business data with the average value or the median, using the interpolation method for estimation. Or using a machine learning model to predict the filling of missing values.
[0211] S806-2, performing one-hot encoding or label encoding on the processed business data to obtain business features, wherein one-hot encoding refers to splitting a discrete feature into multiple binary features. Label encoding refers to representing a discrete feature with a set of integers.
[0212] S806-3, perform minimum-maximum normalization processing or Z-score normalization processing on the service features to obtain service features of a unified scale, wherein the minimum-maximum normalization (Min-max Normalization): the minimum-maximum normalization refers to scaling the features to the interval [0, 1]. The specific formula is as follows:
[0213]
[0214] Among them, x is the business characteristic, and x′ is the business characteristic of the unified scale.
[0215] Z-score normalization: Z-score normalization refers to scaling the features to a normal distribution interval with a mean of 0 and a variance of 1. The specific formula is as follows:
[0216]
[0217] Among them, x is the business characteristic, μ and σ are the mean and standard deviation of the business characteristic respectively, and x′ is the business characteristic of unified scale.
[0218] Step S808, identifying potential abnormal data from the business data using the abnormal database;
[0219] Specifically, the following steps are used to identify potential abnormal data from business data using the abnormal database:
[0220] S808-1, obtaining feature similarity between the business feature and the data feature of the abnormal data in the abnormal database.
[0221] S808-2, identifying the business data corresponding to the business features whose feature similarity is greater than the target threshold as potential abnormal data.
[0222] Step S810, determining at least one reference identification threshold based on potential abnormal data;
[0223] Specifically, the following steps are used to determine at least one reference identification threshold based on potential abnormal data:
[0224] S810-1, obtaining a maximum abnormal affinity and a minimum abnormal affinity from the abnormal affinities corresponding to the potential abnormal data.
[0225] S810-2, half of the standard deviation of the abnormal affinity corresponding to each potential abnormal data is determined as the target step length.
[0226] S810-3, in the numerical range formed by the minimum abnormal affinity and the maximum abnormal affinity, a numerical value is sequentially extracted according to the target step length as a reference recognition threshold.
[0227] Step S812, determining a clustering evaluation coefficient of a reference recognition threshold;
[0228] Specifically, the following steps are used to determine the clustering evaluation coefficient of the reference identification threshold based on the potential abnormal data:
[0229] S812-1, clustering the potential abnormal data in turn using each reference recognition threshold to obtain at least two abnormal data clusters that match the reference recognition threshold.
[0230] S812-2, when the clustering distances between the respective abnormal data in at least two abnormal data clusters are obtained, a weighted sum calculation is performed on the clustering distances between the respective abnormal data in at least two abnormal data clusters to obtain clustering evaluation coefficients of reference identification thresholds corresponding to the at least two abnormal data clusters.
[0231] Step S814, determining a target recognition threshold from at least one reference recognition threshold based on the clustering evaluation coefficient;
[0232] Specifically, the following steps are used to determine the target recognition threshold from at least one reference recognition threshold based on the clustering evaluation coefficient:
[0233] S814-1, sorting the clustering evaluation coefficients that match the respective reference recognition thresholds.
[0234] S814-2, determining the reference recognition threshold corresponding to the maximum clustering evaluation coefficient as the target recognition threshold.
[0235] Step S816, identifying potential abnormal data whose abnormal affinity is greater than the target identification threshold as target abnormal data.
[0236] In an embodiment of the present application, business data to be identified in a target business scenario is obtained. Then, potential abnormal data is identified from the business data using an abnormal database, wherein the abnormal data in the abnormal database is obtained after partial selection and variation of the basic abnormal data using abnormal affinity, and the abnormal affinity is used to indicate the affinity between the abnormal data and the normal data. Further, based on the abnormal affinity corresponding to each of the potential abnormal data, at least one reference recognition threshold is determined, and a clustering evaluation coefficient that matches each reference recognition threshold is determined, wherein the clustering evaluation coefficient is used to evaluate the result of clustering the potential abnormal data using the reference recognition threshold. Then, in the case where a target recognition threshold is determined from at least one reference recognition threshold based on the clustering evaluation coefficient, the potential abnormal data whose abnormal affinity is greater than the target recognition threshold is identified as the target abnormal data. In other words, using an embodiment of the present application, the potential abnormal data after being screened by the abnormal database is compared using an adaptive threshold to determine the target abnormal data from the business data. The purpose of effectively and accurately identifying large-scale data abnormalities is achieved, and it is no longer limited to marking data exceeding a fixed threshold as abnormal data. Thereby, the technical effect of improving the recognition accuracy of abnormal data is achieved, and the problem of low recognition accuracy in the recognition method of abnormal data provided by related technologies is solved.
[0237] It should be noted that, for the aforementioned method embodiments, for the sake of simplicity, they are all expressed as a series of action combinations, but those skilled in the art should be aware that the present application is not limited by the described order of actions, because according to the present application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the present application.
[0238] According to another aspect of the embodiments of the present application, a device for identifying abnormal data for implementing the above-mentioned method for identifying abnormal data is also provided. Fig. 9 As shown, the device comprises:
[0239] An acquisition unit 902 is used to acquire business data to be identified in a target business scenario;
[0240] A first identification unit 904 is used to identify potential abnormal data from the business data using an abnormal database, wherein the abnormal data in the abnormal database is obtained by partially selecting and mutating the basic abnormal data using an abnormal affinity, and the abnormal affinity is used to indicate the affinity between the abnormal data and the normal data;
[0241] A determination unit 906 is used to determine at least one reference recognition threshold and a clustering evaluation coefficient respectively matched with each reference recognition threshold based on the abnormal affinity corresponding to each potential abnormal data, wherein the clustering evaluation coefficient is used to evaluate the result of clustering the potential abnormal data using the reference recognition threshold;
[0242] The second identification unit 908 is used to identify potential abnormal data with abnormal affinity greater than the target identification threshold as target abnormal data when a target identification threshold is determined from at least one reference identification threshold based on the clustering evaluation coefficient.
[0243] Optionally, in this embodiment, the above-mentioned determination unit includes: a first acquisition module, which is used to obtain the maximum abnormal affinity and the minimum abnormal affinity from the abnormal affinities corresponding to the potential abnormal data; an extraction module, which is used to extract a numerical value as a reference recognition threshold in the numerical interval formed by the minimum abnormal affinity and the maximum abnormal affinity according to the target step size; a clustering processing module, which is used to cluster the potential abnormal data in turn using each reference recognition threshold to obtain at least two abnormal data clusters matching the reference recognition threshold; a first determination module, which is used to determine the clustering evaluation coefficient matching the reference recognition threshold using the clustering distance between each abnormal data in at least two abnormal data clusters.
[0244] Optionally, in this embodiment, the above-mentioned first clustering processing module is also used to take each reference identification threshold as the current reference identification threshold in turn, and perform the following operations: compare the abnormal affinity corresponding to each potential abnormal data with the current reference identification threshold; cluster the potential abnormal data with an abnormal affinity greater than the current reference identification threshold into a first abnormal data cluster, and cluster the potential abnormal data with an abnormal affinity less than or equal to the current reference identification threshold into a second abnormal data cluster; determine the first abnormal data cluster and the second abnormal data cluster as abnormal data clusters that match the current reference identification threshold.
[0245] Optionally, in this embodiment, the above-mentioned first determination module is also used to take each abnormal data in at least two abnormal data clusters as the current abnormal data in turn, and perform the following operations: determine the current abnormal data cluster where the current abnormal data is located; obtain the first average clustering distance between the current abnormal data and each first reference abnormal data contained in the current abnormal data cluster; obtain the second average clustering distance between the current abnormal data and each second reference abnormal data contained in other abnormal data clusters outside the current abnormal data cluster; use the first average clustering distance and the second average clustering distance to determine the current object clustering evaluation coefficient matching the current abnormal data; when the object clustering evaluation coefficient corresponding to each abnormal data in at least two abnormal data clusters is obtained, perform weighted sum calculation on all object clustering evaluation coefficients to obtain a clustering evaluation coefficient matching the reference recognition threshold.
[0246] Optionally, in this embodiment, the above-mentioned device also includes: a sorting unit, used to sort the clustering evaluation coefficients matching each reference recognition threshold; a first determination unit, used to determine the reference recognition threshold corresponding to the largest clustering evaluation coefficient as the target recognition threshold.
[0247] Optionally, in this embodiment, the above-mentioned device also includes: a first acquisition unit, used to acquire basic abnormality data; a second acquisition unit, used to acquire extended abnormality data other than the basic abnormality data through at least one immune selection mode, wherein the immune selection mode is used to select abnormality data similar to the basic abnormality data from the data using an immune algorithm; and an updating unit, used to update the abnormality database using the extended abnormality data.
[0248] Optionally, in this embodiment, the second acquisition unit includes: a second acquisition module, used to acquire the candidate extended abnormality data determined by each immune selection mode; a second determination module, used to perform one of the following operations on the candidate extended abnormality data to determine the extended abnormality data: averaging the candidate extended abnormality data determined by each immune selection mode to obtain the extended abnormality data; performing weighted sum calculation on the candidate extended abnormality data determined by each immune selection mode to obtain the extended abnormality data; voting processing on the candidate extended abnormality data determined by each immune selection mode to obtain the extended abnormality data.
[0249] Optionally, in this embodiment, the above-mentioned device also includes: a comparison unit, which is used to compare the abnormal affinity and the clone threshold corresponding to each basic abnormal data in the first immune selection mode; a clone copy unit, which is used to clone and copy the basic abnormal data whose abnormal affinity is greater than the clone threshold to obtain a copy of the abnormal data that matches the basic abnormal data; a mutation processing unit, which is used to perform mutation processing on the abnormal data copy to obtain mutated abnormal data; a second determination unit, which is used to use the mutated abnormal data as the candidate extended abnormal data determined in the first immune selection mode when the abnormal affinity of the mutated abnormal data is greater than the abnormal affinity of the abnormal data copy.
[0250] Optionally, in this embodiment, the above-mentioned device also includes: a first training unit, used to expose a group of detectors to normal data for training in the second immune selection mode to obtain target detectors for detecting abnormal data, wherein, during the training process, detectors that match the normal data will be discarded, and detectors that do not match the normal data will be retained; a third determination unit, used to use the data identified by using the target detector as candidate extended abnormal data determined in the second immune selection mode.
[0251] Optionally, in this embodiment, the above-mentioned first identification unit includes: a cleaning processing module, which is used to clean the business data to obtain processed business data; an encoding module, which is used to perform feature encoding on the processed business data to obtain business features; a normalization processing module, which is used to normalize the business features to obtain business features of a unified scale; a third acquisition module, which is used to obtain feature similarity between the business features and data features of abnormal data in the abnormal database; and an identification module, which is used to identify the business data corresponding to the business features whose feature similarity is greater than the target threshold as potential abnormal data.
[0252] Optionally, in this embodiment, the above-mentioned device also includes: a third acquisition unit, used to obtain neighboring normal data associated with the target abnormal data from the business data, wherein the neighboring normal data is normal data whose distance from the target abnormal data is less than a target distance threshold; and a repair unit, used to repair the target abnormal data using the neighboring normal data.
[0253] For a specific embodiment, reference may be made to the example shown in the above-mentioned method for identifying abnormal data, and this embodiment will not be described in detail here.
[0254] According to another aspect of the embodiment of the present application, an electronic device for implementing the above-mentioned abnormal data identification method is also provided. This embodiment is described by taking the electronic device as a terminal as an example. Fig.10As shown, the electronic device includes a memory 1002 and a processor 1004. The memory 1002 stores a computer program, and the processor 1004 is configured to execute the steps in any of the above method embodiments through the computer program.
[0255] Optionally, in this embodiment, the electronic device may be located in at least one network device among a plurality of network devices of a computer network.
[0256] Optionally, in this embodiment, the processor may be configured to perform the following steps through a computer program:
[0257] S1, obtaining the business data to be identified in the target business scenario;
[0258] S2, identifying potential abnormal data from the business data using an abnormal database, wherein the abnormal data in the abnormal database is obtained by partially selecting and mutating the basic abnormal data using an abnormal affinity, and the abnormal affinity is used to indicate the affinity between the abnormal data and the normal data;
[0259] S3, based on the abnormal affinity corresponding to each of the potential abnormal data, determining at least one reference recognition threshold and a clustering evaluation coefficient respectively matched with each reference recognition threshold, wherein the clustering evaluation coefficient is used to evaluate the result of clustering the potential abnormal data using the reference recognition threshold;
[0260] S4, when a target recognition threshold is determined from at least one reference recognition threshold based on the clustering evaluation coefficient, potential abnormal data having an abnormal affinity greater than the target recognition threshold is identified as target abnormal data.
[0261] Alternatively, a person skilled in the art may understand that: Fig.10 The structure shown is for illustration only, and the electronic device may also be a smart phone (such as an Android phone, an iOS phone, etc.), a tablet computer, a PDA, a mobile Internet device (Mobile Internet Devices, MID), a PAD, or other terminal devices. Fig.10 The structure of the electronic device is not limited. Fig.10 More or fewer components (such as network interfaces, etc.) as shown in, or with Fig.10 Different configurations are shown.
[0262] Among them, the memory 1002 can be used to store software programs and modules, such as the program instructions / modules corresponding to the abnormal data identification method and device in the embodiment of the present application. The processor 1004 executes various functional applications and data processing by running the software programs and modules stored in the memory 1002, that is, realizing the above-mentioned abnormal data identification method. The memory 1002 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 1002 may further include a memory remotely arranged relative to the processor 1004, and these remote memories may be connected to the terminal via a network. Examples of the above-mentioned networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof. Among them, the memory 1002 can be specifically, but not limited to, used to store information such as business data. As an example, such as Fig.10 As shown, the memory 1002 may include, but is not limited to, the acquisition unit 902, the first identification unit 904, the determination unit 906, and the second identification unit 908 in the abnormal data identification device. In addition, other module units in the abnormal data identification device may also be included but are not limited to, which will not be repeated in this example.
[0263] Optionally, the transmission device 1006 is used to receive or send data via a network. Specific examples of the above-mentioned network may include a wired network and a wireless network. In one example, the transmission device 1006 includes a network adapter (Network Interface Controller, NIC), which can be connected to other network devices and routers via a network cable so as to communicate with the Internet or a local area network. In one example, the transmission device 1006 is a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0264] In addition, the electronic device mentioned above further includes: a connection bus 1008, which is used to connect various module components in the electronic device mentioned above.
[0265] In other embodiments, the terminal device or server may be a node in a distributed system, wherein the distributed system may be a blockchain system, and the blockchain system may be a distributed system formed by connecting the multiple nodes through network communication. The nodes may form a point-to-point network, and any form of computing device, such as a server, terminal or other electronic device, may become a node in the blockchain system by joining the point-to-point network.
[0266] According to one aspect of the present application, a computer program product is provided, the computer program product comprising a computer program / instruction, the computer program / instruction comprising a program code for executing the above method. In such an embodiment, the computer program can be downloaded and installed from a network through a communication part, and / or installed from a removable medium. When the computer program is executed by a central processing unit, various functions provided by the embodiments of the present application are executed.
[0267] According to one aspect of the present application, a computer-readable storage medium is provided, and a processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the above method.
[0268] Optionally, in this embodiment, the computer-readable storage medium may be configured to store a computer program for performing the following steps:
[0269] S1, obtaining the business data to be identified in the target business scenario;
[0270] S2, identifying potential abnormal data from the business data using an abnormal database, wherein the abnormal data in the abnormal database is obtained by partially selecting and mutating the basic abnormal data using an abnormal affinity, and the abnormal affinity is used to indicate the affinity between the abnormal data and the normal data;
[0271] S3, based on the abnormal affinity corresponding to each of the potential abnormal data, determining at least one reference recognition threshold and a clustering evaluation coefficient respectively matched with each reference recognition threshold, wherein the clustering evaluation coefficient is used to evaluate the result of clustering the potential abnormal data using the reference recognition threshold;
[0272] S4, when a target recognition threshold is determined from at least one reference recognition threshold based on the clustering evaluation coefficient, potential abnormal data having an abnormal affinity greater than the target recognition threshold is identified as target abnormal data.
[0273] Optionally, in the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program with a predetermined function, and works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as processing circuits or memories) or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be part of an overall module or unit that includes the function of the module or unit.
[0274] Optionally, in this embodiment, a person of ordinary skill in the art may understand that all or part of the steps in the various methods of the above embodiments may be completed by instructing hardware related to the terminal device through a program, and the program may be stored in a computer-readable storage medium, and the storage medium may include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a disk or an optical disk, etc.
[0275] If the integrated units in the above embodiments are implemented in the form of software functional units and sold or used as independent products, they can be stored in the above computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for enabling one or more computer devices (which may be personal computers, servers, or network devices, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present application.
[0276] In the above embodiments of the present application, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, please refer to the relevant description of other embodiments.
[0277] In the several embodiments provided in the present application, it should be understood that the disclosed client can be implemented in other ways. Among them, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.
[0278] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0279] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0280] The above is only a preferred implementation of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.
Claims
1. A method for identifying abnormal data, It is characterized in that include: Obtain the business data to be identified in the target business scenario; Identifying potential abnormal data from the business data using an abnormal database, wherein the abnormal data in the abnormal database is obtained by partially selecting and mutating basic abnormal data using an abnormal affinity, and the abnormal affinity is used to indicate the affinity between the abnormal data and normal data; Based on the abnormal affinity corresponding to each of the potential abnormal data, at least one reference recognition threshold and a clustering evaluation coefficient respectively matched with each of the reference recognition thresholds are determined, wherein the clustering evaluation coefficient is used to evaluate the result of clustering the potential abnormal data using the reference recognition threshold; When a target recognition threshold is determined from the at least one reference recognition threshold based on the clustering evaluation coefficient, the potential abnormal data having the abnormal affinity greater than the target recognition threshold is identified as target abnormal data.
2. The method according to claim 1, It is characterized in that The determining of at least one reference identification threshold based on the abnormal affinity corresponding to each of the potential abnormal data, and the clustering evaluation coefficient respectively matching each of the reference identification thresholds comprises: Obtaining a maximum abnormal affinity and a minimum abnormal affinity from the abnormal affinities corresponding to the potential abnormal data; In the numerical range formed by the minimum abnormal affinity and the maximum abnormal affinity, a numerical value is sequentially extracted according to the target step length as the reference recognition threshold; Using each of the reference identification thresholds to perform clustering processing on the potential abnormal data in turn, to obtain at least two abnormal data clusters that match the reference identification threshold; The clustering evaluation coefficient matching the reference identification threshold is determined by using the clustering distance between each abnormal data in the at least two abnormal data clusters.
3. The method according to claim 2, It is characterized in that The clustering process of the potential abnormal data in sequence by using each of the reference identification thresholds to obtain at least two abnormal data clusters matching the reference identification thresholds includes: Each of the reference recognition thresholds is used as the current reference recognition threshold in turn, and the following operations are performed: Comparing the anomaly affinity corresponding to each of the potential anomaly data with the current reference recognition threshold; Clustering the potential abnormal data whose abnormal affinity is greater than the current reference recognition threshold into a first abnormal data cluster, and clustering the potential abnormal data whose abnormal affinity is less than or equal to the current reference recognition threshold into a second abnormal data cluster; The first abnormal data cluster and the second abnormal data cluster are determined as the abnormal data clusters matching the current reference recognition threshold.
4. The method according to claim 2, It is characterized in that The determining the clustering evaluation coefficient matching the reference identification threshold by using the clustering distance between each abnormal data in the at least two abnormal data clusters includes: Each abnormal data in the at least two abnormal data clusters is sequentially used as current abnormal data, and the following operations are performed: Determine the current abnormal data cluster where the current abnormal data is located; Acquire a first average cluster distance between the current abnormal data and each first reference abnormal data included in the current abnormal data cluster; Acquire a second average clustering distance between the current abnormal data and each second reference abnormal data included in other abnormal data clusters outside the current abnormal data cluster; Determine a current object clustering evaluation coefficient matching the current abnormal data by using the first average clustering distance and the second average clustering distance; When the object clustering evaluation coefficient corresponding to each abnormal data in the at least two abnormal data clusters is obtained, a weighted sum calculation is performed on all the object clustering evaluation coefficients to obtain the clustering evaluation coefficient that matches the reference recognition threshold.
5. The method according to claim 4, It is characterized in that After determining at least one reference recognition threshold and a clustering evaluation coefficient respectively matching each reference recognition threshold based on the abnormal affinity corresponding to each of the potential abnormal data, the method further includes: Sorting the clustering evaluation coefficients that match the reference recognition thresholds; The reference recognition threshold corresponding to the maximum clustering evaluation coefficient is determined as the target recognition threshold.
6. The method according to claim 1, It is characterized in that Before obtaining the business data to be identified in the target business scenario, the method further includes: Acquiring the basic abnormal data; Acquire extended abnormal data other than the basic abnormal data through at least one immune selection mode, wherein the immune selection mode is used to select abnormal data similar to the basic abnormal data from the data using an immune algorithm; The anomaly database is updated using the extended anomaly data.
7. The method according to claim 6, It is characterized in that The obtaining of extended abnormal data other than the basic abnormal data by at least one immune selection mode includes: Obtaining candidate extended abnormality data determined by each of the immune selection modes; Perform one of the following operations on the candidate extended abnormal data to determine the extended abnormal data: averaging the candidate extended abnormal data determined by each of the immune selection modes to obtain the extended abnormal data; or Performing weighted sum calculation on the candidate extended abnormal data determined by each of the immune selection modes to obtain the extended abnormal data; or Voting is performed on the candidate extended abnormal data determined by each of the immune selection modes to obtain the extended abnormal data.
8. The method according to claim 6, It is characterized in that Before obtaining the extended abnormal data other than the basic abnormal data through at least one immune selection mode, the method further includes: In a first immune selection mode, the abnormal affinity and clone threshold corresponding to each of the basic abnormal data are compared; Cloning and duplicating the basic abnormal data whose abnormal affinity is greater than the cloning threshold to obtain an abnormal data copy matching the basic abnormal data; Performing mutation processing on the abnormal data copy to obtain mutated abnormal data; When the abnormal affinity of the mutated abnormal data is greater than the abnormal affinity of the abnormal data copy, the mutated abnormal data is used as the candidate extended abnormal data determined in the first immune selection mode.
9. The method according to claim 6, It is characterized in that Before obtaining the extended abnormal data other than the basic abnormal data through at least one immune selection mode, the method further includes: In the second immune selection mode, a group of detectors are exposed to normal data for training to obtain target detectors for detecting abnormal data, wherein during the training process, detectors matching the normal data will be discarded, while detectors not matching the normal data will be retained; The data identified by the target detector is used as candidate extended abnormal data determined in the second immune selection mode.
10. The method according to any one of claims 1 to 9, It is characterized in that The step of identifying potential abnormal data from the business data using the abnormality database includes: Cleaning the business data to obtain processed business data; Performing feature encoding on the processed service data to obtain service features; Normalizing the service characteristics to obtain the service characteristics of a unified scale; Acquire feature similarity between the business feature and data features of abnormal data in the abnormal database; The business data corresponding to the business features whose feature similarity is greater than a target threshold are identified as the potential abnormal data.
11. The method according to any one of claims 1 to 9, It is characterized in that After identifying the potential abnormal data having the abnormal affinity greater than the target identification threshold as target abnormal data, the method further includes: Acquire neighboring normal data associated with the target abnormal data from the business data, wherein the neighboring normal data is normal data whose distance from the target abnormal data is less than a target distance threshold; The target abnormal data is repaired using the neighboring normal data.
12. A device for identifying abnormal data, It is characterized in that include: An acquisition unit, used to acquire business data to be identified in a target business scenario; A first identification unit is used to identify potential abnormal data from the business data using an abnormal database, wherein the abnormal data in the abnormal database is obtained by partially selecting and mutating basic abnormal data using abnormal affinity, and the abnormal affinity is used to indicate the affinity between the abnormal data and the normal data; a determination unit, configured to determine, based on the abnormal affinity corresponding to each of the potential abnormal data, at least one reference recognition threshold and a clustering evaluation coefficient respectively matched with each of the reference recognition thresholds, wherein the clustering evaluation coefficient is used to evaluate the result of clustering the potential abnormal data using the reference recognition threshold; The second identification unit is used to identify the potential abnormal data whose abnormal affinity is greater than the target identification threshold as target abnormal data when a target identification threshold is determined from the at least one reference identification threshold based on the clustering evaluation coefficient.
13. A computer-readable storage medium, It is characterized in that The computer-readable storage medium includes a stored program, wherein the program is executed by a processor to perform the method described in any one of claims 1 to 11.
14. A computer program product comprising a computer program / instructions, It is characterized in that When the computer program / instructions are executed by a processor, the steps of the method according to any one of claims 1 to 11 are implemented.
15. An electronic device comprising a memory and a processor, It is characterized in that A computer program is stored in the memory, and the processor is configured to execute the method according to any one of claims 1 to 11 through the computer program.