Abnormal data detection method and device, storage medium and electronic equipment

By clustering and merging the target business data, the problem of low accuracy of abnormal data detection in the prior art is solved, and high-precision detection under complex data distribution is achieved.

CN120449038APending Publication Date: 2025-08-08DUXIAOMAN TECH (BEIJING) CO LTD

Patent Information

Application Number
CN202510538769.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-27
Publication Date
2025-08-08

AI Technical Summary

Technical Problem

In the prior art, the accuracy of abnormal data detection is low, especially when the data distribution is complex and changeable, it is difficult to effectively improve the detection accuracy.

Method used

By obtaining the target business data set, calling the target clustering model for clustering, and merging the clustering results, and finally performing abnormal data detection to avoid assuming that the data obeys a specific distribution.

Benefits of technology

It improves the accuracy of abnormal data detection, can flexibly adapt to different data distributions, and provides higher-precision abnormal data detection results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120449038A_ABST
    Figure CN120449038A_ABST
Patent Text Reader

Abstract

The invention provides an abnormal data detection method and device, a storage medium and electronic equipment, and the method comprises the steps: obtaining a target business data set, the target business data set comprises multiple pieces of target business data, one piece of target business data comprises the field information of each target field in M target fields, and M is a positive integer; calling a target clustering model, and performing clustering processing on the target business data set to obtain a clustering result; carrying out merging processing on the clustering result to obtain a clustering result after merging processing; and performing abnormal data detection on the clustering result after the merging processing to obtain an abnormal data detection result. According to the embodiment of the invention, the accuracy of abnormal data detection can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular to an abnormal data detection method, device, storage medium and electronic equipment. Background Art

[0002] Currently, with the increasing informatization of enterprises, the continuous expansion and complexity of business operations are leading to a dramatic increase in data volumes. Inadequate or improperly implemented concurrency control mechanisms, system failures or errors, and data integration and migration processes can all lead to the generation and accumulation of abnormal data in online production environments. Modern enterprises increasingly rely on data to guide decision-making and formulate strategies. Data outliers can lead to decision-making errors and strategic deviations, making data accuracy maintenance particularly important. Relevant technologies typically detect abnormal data by being sensitive to data assumptions and distributions. However, data distributions are often complex and variable, resulting in datasets that often do not conform to the assumed distribution, resulting in low anomaly detection accuracy. Therefore, there is currently no effective solution for improving the accuracy of anomaly data detection. Summary of the Invention

[0003] In view of this, the embodiments of the present invention provide an abnormal data detection method, device, storage medium and electronic device to solve the problem of low accuracy of abnormal data detection caused by related technologies; that is, the embodiments of the present invention can perform abnormal data detection by merging the clustering results after processing, without assuming that the data obeys a certain specific distribution (such as a normal distribution, etc.), thereby performing abnormal data detection more flexibly, effectively improving the accuracy of abnormal data detection, and obtaining abnormal data detection results with higher accuracy.

[0004] According to one aspect of an embodiment of the present invention, a method for detecting abnormal data is provided, the method comprising:

[0005] Acquire a target business data set, where the target business data set includes multiple target business data, and one target business data includes field information of each target field in M target fields, where M is a positive integer;

[0006] Calling the target clustering model to perform clustering processing on the target business data set to obtain a clustering result;

[0007] Merging the clustering results to obtain a merged clustering result;

[0008] Anomaly data detection is performed on the clustering results after the merging process to obtain an abnormal data detection result.

[0009] According to another aspect of an embodiment of the present invention, there is provided an abnormal data detection device, the device comprising:

[0010] an acquiring unit, configured to acquire a target business data set, wherein the target business data set includes a plurality of target business data, and one target business data includes field information of each target field in M target fields, where M is a positive integer;

[0011] A processing unit, configured to call a target clustering model, perform clustering processing on the target business data set, and obtain a clustering result;

[0012] The processing unit is further configured to merge the clustering results to obtain a merged clustering result;

[0013] The processing unit is further configured to perform abnormal data detection on the clustering result after the merging process to obtain an abnormal data detection result.

[0014] According to another aspect of an embodiment of the present invention, an electronic device is provided, comprising a processor and a memory storing a program, wherein the program comprises instructions, which, when executed by the processor, enable the processor to perform the above-mentioned method.

[0015] According to another aspect of an embodiment of the present invention, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable a computer to execute the above-mentioned method.

[0016] The embodiment of the present invention can obtain a target business data set, which includes multiple target business data, and one target business data includes field information of each target field in M target fields, where M is a positive integer; and call the target clustering model to perform clustering processing on the target business data set to obtain a clustering result. Based on this, the clustering results can be merged to obtain a clustering result after the merge processing; and abnormal data detection can be performed on the clustering result after the merge processing to obtain an abnormal data detection result. It can be seen that the embodiment of the present invention can perform abnormal data detection through the clustering result after the merge processing, without assuming that the data obeys a certain specific distribution (such as a normal distribution, etc.), thereby performing abnormal data detection more flexibly, so as to effectively improve the accuracy of abnormal data detection and obtain an abnormal data detection result with higher accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Further details, features and advantages of the present invention are disclosed in the following description of exemplary embodiments in conjunction with the accompanying drawings, in which:

[0018] Figure 1 A schematic flow chart of a method for detecting abnormal data according to an exemplary embodiment of the present invention is shown;

[0019] Figure 2A schematic flow chart of another abnormal data detection method according to an exemplary embodiment of the present invention is shown;

[0020] Figure 3 A schematic flow chart of another abnormal data detection method according to an exemplary embodiment of the present invention is shown;

[0021] Figure 4 A schematic block diagram of an abnormal data detection device according to an exemplary embodiment of the present invention is shown;

[0022] Figure 5 A block diagram of an exemplary electronic device capable of implementing the embodiments of the present invention is shown. DETAILED DESCRIPTION

[0023] Embodiments of the present invention will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present invention are shown in the accompanying drawings, it should be understood that the present invention can be implemented in various forms and should not be construed as limited to the embodiments described herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the present invention. It should be understood that the drawings and embodiments of the present invention are for illustrative purposes only and are not intended to limit the scope of protection of the present invention.

[0024] It should be understood that the various steps described in the method embodiments of the present invention may be performed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present invention is not limited in this respect.

[0025] The term "including" and its variations used in this document are open inclusions, that is, "including but not limited to". The term "based on" means "based at least in part on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one other embodiment"; the term "some embodiments" means "at least some embodiments". The relevant definitions of other terms will be given in the following description. It should be noted that the concepts of "first", "second", etc. mentioned in the present invention are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.

[0026] It should be noted that the modifications of "one" and "multiple" mentioned in the present invention are illustrative rather than restrictive. Those skilled in the art should understand that unless otherwise clearly indicated in the context, it should be understood as "one or more".

[0027] The names of the messages or information exchanged between multiple devices in the embodiments of the present invention are only used for illustrative purposes and are not used to limit the scope of these messages or information.

[0028] It should be noted that the execution subject of the abnormal data detection method provided in the embodiment of the present invention can be one or more electronic devices, and the embodiment of the present invention does not limit this; wherein, the electronic device can be a terminal (i.e., a client) or a server. Then, when the execution subject includes multiple electronic devices, and the multiple electronic devices include at least one terminal and at least one server, the abnormal data detection method provided in the embodiment of the present invention can be jointly executed by the terminal and the server. Accordingly, the terminals mentioned here can include but are not limited to: smart phones, tablet computers, laptops, desktop computers, intelligent voice interaction devices, etc. The server mentioned here can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides cloud services, cloud databases, cloud computing (cloud computing), cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and basic cloud computing services such as big data and artificial intelligence platforms, etc.

[0029] Based on the above description, an embodiment of the present invention proposes an abnormal data detection method, which can be executed by the electronic device (terminal or server) mentioned above; or, the abnormal data detection method can be executed by the terminal and the server together. For the sake of convenience, the following description will be based on the example of an electronic device executing the abnormal data detection method; Figure 1 As shown, the abnormal data detection method may include the following steps S101-S104:

[0030] S101 , obtaining a target business data set, where the target business data set includes a plurality of target business data, and one target business data set includes field information of each target field in M target fields, where M is a positive integer.

[0031] Optionally, the M target fields may include one or more fields in the target data table, which is not limited in this embodiment of the present invention. Optionally, the target data table may be any data table, which is not limited in this embodiment of the present invention. Based on this, the target business data set may include field information of each of the M target fields in the target data table. Optionally, the M target fields may be set based on experience or actual needs, etc., which is not limited in this embodiment of the present invention. For example, all fields in a segment of JSON (JavaScript Object Notation, a lightweight data exchange format) data (such as data amount, order information, etc.) may be used as target fields (in which case the M target fields may include all fields in the segment of JSON data), or each field in the target data table may be sequentially formed into M target fields (i.e., the M target fields may include any field in the target data table), in which case the value of M may be 1, thereby performing abnormal data detection (also known as consistency analysis) on the field information under each field in the target data table (i.e., on each column of data in the target data table), etc. Based on this, the embodiment of the present invention can perform abnormal data detection on the field information under each field in the target data table respectively (such as M target fields can include a field in the target data table in turn, and at this time, each column of data in the target data table can be analyzed for consistency, such as analyzing the consistency between multiple business data composed of any column of data (at this time, one business data may include a value in any column of data, that is, each value in any column of data is a business data)), and can also perform abnormal data detection on the field information under multiple fields in the target data table at one time (such as M target fields can include multiple fields in the target data table, and at this time, multiple columns of data in the target data table can be analyzed for consistency at one time, that is, analyzing the consistency between multiple business data composed of multiple columns of data (at this time, one business data may include a value in each column of data in multiple columns of data)), etc.; the embodiment of the present invention is not limited to this. Optionally, a target business data including the field information of each target field in the M target fields can also be expressed as a target business data including the field information of the corresponding target business data under each target field. Optionally, the target business data set may include target business data of each target field within a target time range; optionally, the target time range may be set based on experience or actual needs, and the embodiment of the present invention is not limited to this.

[0032] In the embodiment of the present invention, the target service data set may be obtained in the following ways, but is not limited to:

[0033] The first acquisition method: the electronic device can obtain an initial business data set, which includes multiple initial business data, and one initial business data can include initial field information of each target field; then, a data preprocessing strategy can be determined, and the initial business data set can be preprocessed according to the data preprocessing strategy to obtain the initial business data set after data preprocessing; wherein, the data preprocessing strategy can include a text-to-value vector strategy, and the text-to-value vector strategy can be used to indicate that text-type data is converted into a value vector; thus, the initial business data set after data preprocessing can be used as the target business data set to obtain the target business data set. Optionally, an initial business data including the initial field information of each target field can also be expressed as an initial business data including the initial field information of the corresponding initial business data under each target field.

[0034] Optionally, the data preprocessing strategy may also include but is not limited to at least one of the following: a data cleaning strategy and a data normalization strategy, etc., which is not limited in the embodiment of the present invention; that is, the data preprocessing strategy may include one or more processing strategies, such as a data cleaning strategy, a data normalization strategy, and a text-to-value vector strategy, etc. Then, accordingly, the electronic device may perform data preprocessing on the initial business data set according to each processing strategy in the data preprocessing strategy, so as to realize data preprocessing of the initial business data set according to the data preprocessing strategy. In other words, when performing data preprocessing on the initial business data set according to the data preprocessing strategy, the initial business data set may be preprocessed according to the data cleaning strategy, and the initial business data set may be preprocessed according to the data normalization strategy, and the initial business data set may be preprocessed according to the text-to-value vector strategy, etc.; based on this, data preprocessing of the initial business data set according to the data preprocessing strategy can be realized. It should be noted that the embodiment of the present invention does not limit the execution order of each processing strategy in the data preprocessing strategy. For example, data cleaning can be performed first (that is, data preprocessing can be performed according to the data cleaning strategy), then data preprocessing can be performed according to the data normalization strategy, and finally data preprocessing can be performed according to the text-to-value vector strategy. Data preprocessing can also be performed according to the data normalization strategy and the text-to-value vector strategy at the same time, and so on. Optionally, the acquired initial business data set can be all the collected business data, that is, it can be all the business data collected within the target time range. For example, assuming that the number of data collection moments within the target time range is N, and N is a positive integer, then the number of initial business data in the acquired initial business data set can be N, that is, the number of initial field information under each target field can be N, that is, N can represent the number of data items corresponding to each target field, then the initial field information under each target field can be represented in sequence as [A 11 ,A 12,...,A 1N ],[A 21 ,A 22 ,...,A 2N ],…,[A M1 ,A M2 ,...,A MN ]; In this case, the number of initial service data in the initial service data set may be N, and the nth initial service data may be expressed as [A 1n ,A 2n ,...,A Mn ], which can include the nth initial field information under each target field, n∈[1,N]; for example, when the value of M is 1, the nth initial business data can be A 1n , which can be the nth initial field information under the target field, and so on.

[0035] Based on this, when the initial business data set is preprocessed according to the data cleaning strategy, the electronic device can perform data cleaning on the initial business data set to identify the null values (such as NaN, None, and empty strings, etc.) in the initial business data set; optionally, after identifying the null values in the initial business data set, for any null value in the initial business data set, the electronic device can delete the initial business data where any null value is located from the initial business data set, and / or mark the initial business data where any null value is located as abnormal data and issue an abnormal data warning in advance, and / or fill in the data type of the field where any null value is located (such as if the data type is a numerical type, it can be filled with a numerical value such as 0), etc.; the embodiment of the present invention does not limit this. Based on this, the electronic device can obtain the initial business data set after data preprocessing by the data cleaning strategy. At this time, the current initial business data set can also be called the initial business data set after data cleaning, etc.

[0036] Accordingly, when the initial business data set (i.e., the current initial business data set, such as the acquired initial business data set or the initial business data set after the above data cleaning, etc.) is pre-processed according to the text-to-value vector strategy, the electronic device can determine whether there is a text field to be converted in the M target fields. The text field to be converted can refer to a field whose initial field information is text-type data; if there is a text field to be converted in the M target fields, the initial field information of the text field to be converted in each initial business data included in the initial business data set can be converted (the initial field information of the text field to be converted in an initial business data can also be referred to as the initial field information of an initial business data in the text field to be converted). The initial field information under this field is used for text conversion to obtain the text conversion numerical vector of the text field to be converted in each initial business data, and the text conversion numerical vector of the text field to be converted in each initial business data is used to update the initial field information of the text field to be converted in the corresponding initial business data, so as to realize data preprocessing of the initial business data set according to the text conversion numerical vector strategy, thereby realizing the update of the initial business data set, then the current initial business data set can be the initial business data set after data preprocessing by the text conversion numerical vector strategy; optionally, the current initial business data set can also be called the initial business data set after text processing. Optionally, the number of text fields to be converted can be Q, where Q is a non-negative integer. Optionally, for any initial business data in the initial business data set and any text field to be converted in the M target fields, the electronic device may perform text preprocessing (such as word segmentation, stop word removal, stemming / lemmatization, etc.) on the initial field information of any text field to be converted in any initial business data to obtain the text preprocessing field information of any text field to be converted in any initial business data, and may convert the text preprocessing field information of any text field to be converted in any initial business data into a numerical vector through text representation technologies such as TF-IDF (Term Frequency-Inverse Document Frequency), Word2Vec (word to vector, a related model for generating word vectors), and BERT (Bidirectional Encoder Representation from Transformers, a bidirectional language model), thereby obtaining a text conversion numerical vector of any text field to be converted in any initial business data, so as to implement text conversion of the initial field information of any text field to be converted in any initial business data, etc. Based on this, the embodiment of the present invention does not limit the specific implementation method of text conversion.

[0037] Accordingly, when performing data preprocessing on the initial business data set (i.e., the current initial business data set, such as the initial business data set after the above-mentioned data cleaning or the initial business data set after text processing, etc.) according to the data normalization strategy, the electronic device can determine whether there is a numerical data field in the M target fields, and the numerical data field can refer to a field whose initial field information is numerical data; if there is a numerical data field in the M target fields, then data normalization processing can be performed on the initial field information of the numerical data field in each initial business data to obtain the normalized field information of the numerical data field in each initial business data, and the normalized field information of the numerical data field in each initial business data is used to update the initial field information of the numerical data field in the corresponding initial business data, so as to realize data preprocessing of the initial business data set according to the data normalization strategy, thereby realizing the update of the initial business data set, then the current initial business data set can be the initial business data set after data preprocessing by the data normalization strategy; optionally, at this time, the current initial business data set can also be called the initial business data set after data normalization. Optionally, the number of numerical data fields in the M target fields can be P, where P is a non-negative integer. Optionally, for any numerical data field in the M target fields, the electronic device can perform data normalization on the initial field information of any numerical data field in each initial business data to obtain normalized field information of any numerical data field in each initial business data, and so on. In this case, the electronic device can scale the eigenvalues to the same range (such as 0 to 1, or -1 to 1, etc.), thereby improving data quality, eliminating the impact of different dimensions on model training, and optimizing the convergence speed and clustering effect of subsequent clustering algorithms.

[0038] Based on this, the data preprocessing process can be completed after the initial business data set is preprocessed according to each processing strategy in the data preprocessing strategy; then accordingly, the initial business data set after the data preprocessing is completed (that is, the current initial business data set after the initial business data set is preprocessed according to each processing strategy in the data preprocessing strategy, such as the initial business data set after the text processing or the initial business data set after data normalization, etc.) can be used as the initial business data set after data preprocessing, thereby obtaining the initial business data set after data preprocessing, that is, the current initial business data set after the data preprocessing is completed can be used as the initial business data set after data preprocessing. Exemplarily, when data preprocessing is implemented in accordance with the data cleaning strategy, the text-to-value vector strategy (also referred to as the text processing strategy) and the data normalization strategy in sequence, the data preprocessing process can be completed after the data preprocessing is performed according to the data normalization strategy, thereby using the current initial business data set (which can be the initial business data set after data normalization) as the initial business data set after data preprocessing, and so on.

[0039] The second acquisition method: the target service data set is stored in the storage space of the electronic device itself. In this case, the target service data set can be acquired from the storage space itself.

[0040] The third acquisition method: the electronic device may obtain a download link of the target business data set, and collect the business data downloaded based on the target business data download link as the target business data set, and so on.

[0041] S102: Call the target clustering model to perform clustering processing on the target business data set to obtain a clustering result.

[0042] Optionally, the target clustering model can be a DBSCAN (Density-Based Spatial Clustering of Applications with Noise) clustering model, or a hierarchical clustering model, etc.; this is not limited in the present embodiment. Based on this, the target clustering model can be an unsupervised machine learning model. In this case, the present embodiment can perform clustering processing on the target business data set using the unsupervised machine learning algorithm to obtain a clustering result.

[0043] Optionally, the M target fields may correspond to one target clustering model; optionally, when the M target fields are different, the target clustering models corresponding to the different M target fields may be the same or different, and the embodiment of the present invention does not limit this. For ease of explanation, the target clustering model is subsequently obtained by tuning parameters through the field information under the M target fields, that is, the target clustering models corresponding to different M target fields can be obtained by tuning parameters through the field information under the corresponding M target fields. At this time, the model parameters in the target clustering models corresponding to different M target fields can be different, that is, the corresponding target clustering models can be different.

[0044] Optionally, the model parameters of the target clustering model may be obtained through pre-parameter tuning, or may be obtained through real-time parameter tuning of the target business data set, etc.; this is not limited in the embodiment of the present invention.

[0045] S103: merging the clustering results to obtain a merged clustering result.

[0046] In an embodiment of the present invention, the electronic device may merge data clusters with a high degree of similarity into the same data cluster, that is, may merge similar data clusters to implement merging processing of clustering results (also referred to as similarity merging processing).

[0047] S104: performing abnormal data detection on the clustering results after the merging process to obtain abnormal data detection results.

[0048] Optionally, the clustering result after the merge process can indicate the consistency of the data, thereby realizing abnormal data detection (i.e., consistency analysis); illustratively, when the target business data set is clustered into a single class (i.e., the clustering result after the merge process includes only one target data cluster), it can indicate the high consistency of the target business data set, that is, it can indicate the high consistency between the various target business data. At this time, the abnormal data detection result may be empty, or the no abnormal data indication information may be used as the abnormal data detection result, etc.; correspondingly, when abnormal data is detected, the abnormal data detection result may not be empty, which indicates that there is a data inconsistency problem. Optionally, the no abnormal data indication information can be set according to experience or according to actual needs, and the embodiment of the present invention does not limit this.

[0049] Optionally, when abnormal data is detected (i.e., the abnormal data detection result includes abnormal data), the electronic device may also trigger a real-time feedback mechanism, that is, it may also issue an alarm notification for the abnormal data according to a preset alarm strategy, that is, it may send an alarm message to the target object corresponding to the abnormal data (such as a management personnel, etc.) according to the preset alarm strategy to implement an alarm notification for the abnormal data. Optionally, the preset alarm strategy may be set based on experience or according to actual needs, and the embodiment of the present invention is not limited to this. Exemplarily, the alarm message may be notified to the target object via email, text message, or internal system. Optionally, the embodiment of the present invention does not limit the specific content of the alarm message. Exemplarily, the alarm message may include but is not limited to at least one of the following: the abnormal type, time, abnormal data, and abnormal data indication information corresponding to the abnormal data, etc., and the embodiment of the present invention is not limited to this. Optionally, the abnormal data indication information may be set based on experience or according to actual needs, and the embodiment of the present invention is not limited to this.

[0050] Optionally, the electronic device may further determine the abnormality analysis data to be recorded and record the data in a database. Optionally, the abnormality analysis data to be recorded may include, but is not limited to, at least one of the following: abnormal data in the abnormal data detection result, the time of abnormal data detection, the time of abnormal data collection, the field in which the abnormal data is located, etc., which are not limited in the embodiment of the present invention. Based on this, the target object can subsequently conduct in-depth analysis and correction of the abnormal data.

[0051] Optionally, after the target object completes the data correction, it can perform a data correction confirmation operation, then the electronic device can detect the data correction confirmation instruction, and thus iteratively execute the acquisition of the target business data set, thereby performing abnormal data detection on the target business data set (that is, the currently acquired target business data set can be clustered and merged again, thereby performing abnormal data detection on the clustering result after the current merge); based on this, the embodiment of the present invention can re-run the batch test to evaluate the data status to ensure that the anomaly has been properly resolved, forming a detection closed loop. Optionally, the electronic device can also, after detecting abnormal data, preset a waiting correction time period to iteratively execute the acquisition of the target business data set, thereby performing abnormal data detection on the currently acquired target business data set, until there is no abnormality in the data within the target time range. That is to say, when it is determined that there are no abnormal values in the current target business data set (that is, there is no abnormal data in the current target business data set within the target time range), the abnormal data detection of the target business data within the target time range can be terminated, such as Figure 2 shown, etc.

[0052] The embodiment of the present invention can obtain a target business data set, which includes multiple target business data, and one target business data includes field information of each target field in M target fields, where M is a positive integer; and call the target clustering model to perform clustering processing on the target business data set to obtain a clustering result. Based on this, the clustering results can be merged to obtain a clustering result after the merge processing; and abnormal data detection can be performed on the clustering result after the merge processing to obtain an abnormal data detection result. It can be seen that the embodiment of the present invention can perform abnormal data detection through the clustering result after the merge processing, without assuming that the data obeys a certain specific distribution (such as a normal distribution, etc.), thereby performing abnormal data detection more flexibly, so as to effectively improve the accuracy of abnormal data detection and obtain an abnormal data detection result with higher accuracy.

[0053] Based on the above description, the embodiment of the present invention also proposes another abnormal data detection method. Accordingly, the abnormal data detection method can be executed by the electronic device (terminal or server) mentioned above; or, the abnormal data detection method can be executed by the terminal and the server together. For the sake of convenience, the following description will be based on the example of an electronic device executing the abnormal data detection method; please refer to Figure 3 The abnormal data detection method may include the following steps S301-S305:

[0054] S301 , obtaining a target business data set, where the target business data set includes multiple target business data, and one target business data set includes field information of each target field in M target fields, where M is a positive integer.

[0055] S302: Call the target clustering model to perform clustering processing on the target business data set to obtain a clustering result.

[0056] Optionally, the electronic device may also determine a training business data set, and the training business data set may include multiple training business data, and one training business data may include field information of each target field, that is, one training business data may include field information of the corresponding training business data under each target field. Optionally, the training business data set may include multiple historical business data; or, the target business data set may be added to the training business data set to achieve the determination of the training business data set (that is, each target business data may be used as training business data) to determine the training business data set. In this case, the training business data set may include multiple target business data and / or multiple historical business data in the target business data set, and so on; the embodiment of the present invention is not limited to this. Among them, when the training business data set only includes multiple target business data, the target business data set may be used as the training business data set to achieve the addition of the target business data set to the training business data set, thereby determining the training business data set. Correspondingly, when the training business data set includes multiple target business data, the electronic device can perform parameter tuning (i.e., calculate the optimal parameters) through real-time business data, so that the target clustering model can continuously learn the changes in business data, thereby more accurately identifying data patterns, thereby improving the performance of the target clustering model, and continuously improving the accuracy of abnormal data detection.

[0057] Based on this, the electronic device can verify the clustering effect of each clustering model in the multiple clustering models based on the training business data set, and obtain clustering effect indication information of each clustering model; wherein the clustering effect indication information of a clustering model is used to indicate the clustering effect of the corresponding clustering model. Optionally, multiple clustering models can be set according to experience or according to actual needs, and the embodiment of the present invention does not limit this; optionally, different clustering models can be different clustering algorithms, or the same clustering algorithm but different parameter combinations (i.e., model parameters), etc., and the embodiment of the present invention does not limit this. Optionally, the clustering effect indication information may include but is not limited to at least one of the following: silhouette coefficient (the value range of the silhouette coefficient is [-1, 1], and the larger the value, the better the clustering effect) and Calinski-Harabasz index (the Calinski-Harabasz index evaluates the clustering effect by calculating the ratio of the inter-cluster dispersion (the sum of the weighted square distances from all cluster centroids to the global centroid) to the intra-cluster dispersion (the sum of the square distances of all samples to their centroids). The larger the index, the better the clustering effect), etc. The embodiment of the present invention is not limited to this.

[0058] In one embodiment, the clustering effect verification can be a K-fold cross-validation, where K is a positive integer. In this case, for any cross-validation in the K-fold cross-validation and any clustering model among multiple clustering models, the electronic device can determine the training set and test set under any cross-validation from the training business data set, and call any clustering model to perform clustering processing on the training set under any cross-validation to obtain the clustering processing result of any clustering model under any cross-validation, thereby determining the clustering cluster division result of the test set of any clustering model under any cross-validation through the clustering processing result of any clustering model under any cross-validation, and the clustering cluster division result can be used to indicate the category (i.e., cluster cluster) to which each data point (i.e., training business data) in the test set under any cross-validation belongs, and then determine the clustering effect indication information of any clustering model under any cross-validation through the clustering cluster division result of the test set of any clustering model under any cross-validation, and use the mean between the clustering effect indication information of any clustering model under K-fold cross-validation as the clustering effect indication information of any clustering model. In another embodiment, for any cross-validation in the K-fold cross-validation and any clustering model among multiple clustering models, the electronic device can determine the training set and test set under any cross-validation from the training business data set, and call any clustering model to perform clustering processing on the training set under any cross-validation, and obtain the clustering processing result of any clustering model under any cross-validation, thereby determining the clustering effect indication information of any clustering model under any cross-validation through the clustering processing result of any clustering model under any cross-validation, and then using the mean between the clustering effect indication information of any clustering model under K-fold cross-validation as the clustering effect indication information of any clustering model. In another embodiment, for any clustering model among multiple clustering models, the electronic device can call any clustering model to perform clustering processing on the training business data set, obtain the clustering result of the training business data set, and determine the clustering effect indication information of any clustering model through the clustering result of the training business data, thereby realizing clustering effect verification of each clustering model, and so on.

[0059] Furthermore, the electronic device may select a clustering model with the best clustering effect from multiple clustering models based on the clustering effect indication information of each clustering model, and use the selected clustering model as the target clustering model. For example, when the greater the clustering effect indication information, the better the clustering effect, the clustering model with the largest clustering effect indication information may be selected from the multiple clustering models, and the clustering model with the largest clustering effect indication information may be used as the target clustering model, and so on. Based on this, embodiments of the present invention may select a clustering model composed of a parameter combination that achieves the best clustering effect to obtain a target clustering model.

[0060] Optionally, the electronic device may trigger the execution of adding the target business data set to the training business data set every preset tuning interval to determine the training business data set, and then perform parameter tuning to obtain the current target clustering model, and then trigger the execution of calling the target clustering model (that is, the current target clustering model), clustering the target business data set, and obtaining the clustering result; that is, the target clustering model can be continuously updated. Optionally, the electronic device may also use the target business data set to determine the training business data set each time before calling the target clustering model and clustering the target business data set to obtain the clustering result, so as to perform parameter tuning on the clustering model, thereby determining the current target clustering model, and so on; the embodiment of the present invention is not limited to this. Optionally, the preset tuning interval duration may be set according to experience or according to actual needs, and the embodiment of the present invention is not limited to this.

[0061] Optionally, the clustering results may include one or more clusters (also referred to as data clusters), which is not limited in this embodiment of the present invention. Optionally, the clustering results may also include W outliers (i.e., abnormal business data), where W is a non-negative integer. For example, when the target clustering algorithm is DBSCAN clustering, data points in the data set that are not included in any cluster constitute outliers, and so on. This is not limited in this embodiment of the present invention. Based on this, when the clustering results include outliers, the electronic device may also add the outliers in the clustering results to the abnormal data detection results as abnormal data in the abnormal data detection results.

[0062] S303: When the clustering result includes multiple clusters, respectively calculate cluster indication information of each of the multiple clusters.

[0063] In an embodiment of the present invention, when the clustering result includes multiple clusters, the electronic device may trigger the execution of respectively calculating cluster indication information of each of the multiple clusters to perform merging processing, etc. Optionally, when the clustering result includes one cluster, the cluster in the clustering result may be used as the target cluster to implement merging processing on the clustering results to obtain a clustering result after merging processing, that is, the clustering result may be used as the clustering result after merging processing, etc.

[0064] Optionally, the cluster indication information of a cluster can be the intra-cluster data variance of the corresponding cluster, or the intra-cluster data density of the corresponding cluster, etc., which is not limited in the embodiment of the present invention. Among them, the intra-cluster data variance can be used to measure the degree of dispersion of data points in a cluster (i.e., cluster cluster), which reflects the degree of deviation of each data point in the cluster from the center of the cluster (i.e., the mean point). The smaller the variance, the more closely the data points in the cluster are connected. The higher the homogeneity of the data, the larger the variance, which means that the data points in the cluster are more dispersed. The intra-cluster data density can be an indicator used to measure the degree of distribution density of data points in a cluster. A high density means that the data points are relatively tightly clustered in the spatial area occupied by the cluster, while a low density means that the data points are relatively dispersed, which reflects the degree of concentration of the data in the cluster. Optionally, when calculating the intra-cluster data density of any cluster, the local density of any data point in any cluster may be calculated. The local density of any data point may be the number of data points whose distance from any data point to all data points in any cluster is less than a specified neighborhood radius. The mean of the local densities of the data points in any cluster is used as the intra-cluster data density of any cluster, and so on. Optionally, the specified neighborhood radius may be set based on experience or actual needs, and this is not limited in the embodiments of the present invention.

[0065] S304: Based on the cluster indication information of each cluster, the clustering results are merged to obtain a merged clustering result, where the merged clustering result includes at least one target cluster.

[0066] In one embodiment, the electronic device can determine the similarity indication information between each cluster in multiple clusters based on the cluster indication information of each cluster, and merge the clustering results based on the similarity indication information between each cluster in multiple clusters to obtain the merged clustering results. Optionally, if the minimum value in the similarity indication information between the two clusters is less than a preset similarity threshold, the two clusters with the smallest similarity indication information can be merged into the same cluster, thereby obtaining the current cluster. If the minimum value in the similarity indication information between the two clusters is greater than or equal to the preset similarity threshold, the merging process can be stopped. If the current number of clusters is greater than 1, the above-mentioned determination of the similarity indication information between the two clusters (i.e., the two clusters in the current cluster) is iteratively executed to perform the merging process until a merging stop condition is met (e.g., the current number of clusters is 1 or the minimum value in the similarity indication information between the two clusters in the current cluster is greater than or equal to the preset similarity threshold, etc.), thereby obtaining a clustering result after the merging process. Each current cluster can be used as a target cluster, so that the clustering result after the merging process includes the current cluster, i.e., the cluster when the merging stop condition is met. Optionally, the preset similarity threshold can be set based on experience or actual needs, and this is not limited in this embodiment of the present invention. Optionally, the similarity indication information between two clusters may be the absolute value of the difference between the cluster indication information of the two clusters, and the similarity indication information between any two clusters may include the similarity indication information between any cluster in all clusters (such as multiple clusters or the current cluster) and each cluster except any cluster (that is, it may include the similarity indication information between any two paired clusters in all clusters), and so on.

[0067] In another embodiment, the electronic device may determine a merge indication threshold and, based on the merge indication threshold and the cluster indication information of each cluster cluster, merge the cluster results to obtain a merged cluster result. Optionally, when a cluster indication information is a variance of data within a cluster, the merge indication threshold may be a variance merge indication threshold. In this case, cluster clusters with cluster indication information less than the variance merge indication threshold in multiple cluster clusters may be merged into the same cluster cluster to achieve merging the cluster results and obtain a merged cluster result. Alternatively, when a cluster indication information is a density of data within a cluster, the merge indication threshold may be a density merge indication threshold. In this case, cluster clusters with cluster indication information greater than the density merge indication threshold in multiple cluster clusters may be merged into the same cluster cluster to achieve merging the cluster results and obtain a merged cluster result, etc. This embodiment of the present invention is not limited to this. Optionally, both the variance merge indication threshold and the density merge indication threshold may be set based on experience or actual needs, which is not limited in this embodiment of the present invention.

[0068] Based on this, the embodiment of the present invention can improve the accuracy and interpretability of clustering results through merging processing.

[0069] S305: Perform abnormal data detection on the clustering results after the merging process to obtain abnormal data detection results.

[0070] Optionally, the electronic device can separately determine the number of cluster samples of each target cluster in the clustering result after the merge processing, that is, the number of data points in the cluster; based on this, based on the number of cluster samples of each target cluster, an abnormal cluster can be determined from the clustering result after the merge processing, and the abnormal cluster can be added to the abnormal data detection result. The number of cluster samples of the abnormal cluster is less than the number of cluster samples of any target cluster other than the abnormal cluster in the clustering result after the merge processing. Optionally, when the number of target clusters in the clustering result after the merge process is 1, the abnormal cluster cluster can be determined to be empty; or, when the number of target clusters in the clustering result after the merge process is greater than 1, the target cluster result with the smallest number of cluster samples in the clustering result after the merge process can be used as the abnormal cluster cluster; or, when the number of target clusters in the clustering result after the merge process is greater than 1, the target cluster cluster with the number of cluster samples less than a preset cluster sample number threshold in the clustering result after the merge process can also be used as the abnormal cluster cluster, so as to determine the abnormal cluster cluster from the clustering result after the merge process, etc.; the embodiment of the present invention is not limited to this. Based on this, the abnormal data detection result can include all data in the abnormal cluster cluster (i.e., target business data), that is, all data in the abnormal cluster cluster can be used as abnormal data.

[0071] Based on this, the embodiments of the present invention can mine hidden features in the data through unsupervised machine learning clustering algorithms, build a model of data cluster distribution, and then identify possible abnormal clusters (i.e., the above-mentioned abnormal cluster clusters) or abnormal points (i.e., samples in abnormal cluster clusters). Abnormal clusters may contain data points that are inconsistent with or erroneous with most data to achieve abnormal data detection, thereby obtaining abnormal data detection results with higher accuracy.

[0072] From the above, it can be seen that the embodiment of the present invention proposes a business anomaly data detection method based on machine learning, which aims to conduct in-depth analysis of business data through machine learning algorithms, automatically discover potential data anomalies, effectively improve the efficiency and detection range of data anomaly detection, and effectively improve the coverage of detection results, provide strong support for data quality assurance, and ensure the normal operation of business functions. In other words, the embodiment of the present invention can use unsupervised machine learning technology to deeply mine the inherent structure and pattern characteristics of data, and autonomously identify and discover hidden structures, laws or knowledge from massive historical data accumulation or unlabeled data sets generated in real time; this process aims to achieve automatic classification and anomaly detection of data, and intelligently identify data points in the data set that deviate from the norm or anomaly through algorithms; based on this, the embodiment of the present invention can continuously monitor the data stream by implementing an automated batch data detection process, and use a real-time feedback mechanism to quickly return the detection results. This process aims to identify and mark potential abnormal data, thereby ensuring the consistency and high quality of the data.

[0073] An embodiment of the present invention can obtain a target business data set, the target business data set including multiple target business data, each target business data including field information of each target field in M target fields, where M is a positive integer; and call a target clustering model to perform clustering processing on the target business data set to obtain a clustering result; wherein the clustering result includes multiple clusters. Based on this, cluster indication information of each cluster in the multiple clusters can be calculated separately; and based on the cluster indication information of each cluster, the clustering results are merged to obtain a merged clustering result, and the merged clustering result includes at least one target cluster. Furthermore, abnormal data detection can be performed on the merged clustering result to obtain an abnormal data detection result. It can be seen that the embodiment of the present invention can use an unsupervised machine learning algorithm to perform data outlier detection (i.e., abnormal data detection) on the internal data source of the online production environment in an offline environment, aiming to ensure the quality of data stored in the data table and achieve data cleanliness and consistency. In addition, the embodiment of the present invention can effectively solve the problems of low efficiency, limited detection range, and lack of intelligence in abnormal data detection. That is, the embodiment of the present invention can introduce a machine learning algorithm to perform feature extraction and data learning on the business data set to achieve fast and intelligent data anomaly detection. In other words, the embodiment of the present invention can intelligently and flexibly perform abnormal data detection without setting rule matching, etc., to effectively avoid missed detection or false detection due to incomplete rule setting or increased data complexity. The machine learning algorithm can adapt to changes in different data types and business logic, and dynamically adjust the parameters of the machine learning model according to changes in real-time data to cope with complex scenarios. In addition, the unsupervised learning method can deeply explore the potential patterns and structures in the data, and can be flexibly applied in various complex and changing business environments, thereby revealing the business laws and anomalies hidden in the data, discovering potential patterns and anomalies, and thus effectively improving the accuracy of abnormal data detection, providing effective support for business decision-making.

[0074] Based on the description of the related embodiments of the above abnormal data detection method, the embodiment of the present invention further proposes an abnormal data detection device, which can be a computer program (including program code) running in an electronic device; Figure 4 As shown, the abnormal data detection device may include an acquisition unit 401 and a processing unit 402. The abnormal data detection device may execute Figure 1 or Figure 3 The abnormal data detection method shown, that is, the abnormal data detection device can run the above units:

[0075] An acquiring unit 401 is configured to acquire a target business data set, wherein the target business data set includes a plurality of target business data, and one target business data set includes field information of each target field in M target fields, where M is a positive integer.

[0076] The processing unit 402 is configured to call a target clustering model, perform clustering processing on the target business data set, and obtain a clustering result;

[0077] The processing unit 402 is further configured to merge the clustering results to obtain a merged clustering result;

[0078] The processing unit 402 is further configured to perform abnormal data detection on the clustering result after the merging process to obtain an abnormal data detection result.

[0079] In one embodiment, the processing unit 402 may further be configured to:

[0080] Determine a training service data set, where the training service data set includes a plurality of training service data;

[0081] Based on the training business data set, respectively verify the clustering effect of each clustering model in the plurality of clustering models to obtain clustering effect indication information of each clustering model; wherein the clustering effect indication information of a clustering model is used to indicate the clustering effect of the corresponding clustering model;

[0082] Based on the clustering effect indication information of each clustering model, a clustering model with the best clustering effect is selected from the multiple clustering models, and the selected clustering model is used as the target clustering model.

[0083] In another embodiment, when determining the training service data set, the processing unit 402 may be specifically configured to:

[0084] The target business data set is added to the training business data set to determine the training business data set.

[0085] In another embodiment, when the processing unit 402 merges the clustering results to obtain the merged clustering results, it can be specifically used to:

[0086] When the clustering result includes multiple clusters, respectively calculating cluster indication information of each of the multiple clusters;

[0087] Based on the cluster indication information of each cluster, the clustering results are merged to obtain a merged clustering result, where the merged clustering result includes at least one target cluster.

[0088] In another embodiment, when the processing unit 402 performs abnormal data detection on the clustering result after the merge process and obtains the abnormal data detection result, it can be specifically used to:

[0089] Respectively determining the number of cluster samples of each target cluster in the clustering result after the merging process;

[0090] Based on the number of cluster samples of each target cluster, an abnormal cluster is determined from the clustering result after the merge processing, and the abnormal cluster is added to the abnormal data detection result. The number of cluster samples of the abnormal cluster is less than the number of cluster samples of any target cluster other than the abnormal cluster in the clustering result after the merge processing.

[0091] In another embodiment, when acquiring the target business data set, the acquiring unit 401 may be specifically configured to:

[0092] Acquire an initial business data set, wherein the initial business data set includes a plurality of initial business data, and one initial business data includes initial field information of each target field;

[0093] Determine a data preprocessing strategy, and perform data preprocessing on the initial business data set according to the data preprocessing strategy to obtain the initial business data set after data preprocessing; wherein the data preprocessing strategy includes a text-to-numeric vector strategy, and the text-to-numeric vector strategy is used to instruct to convert text data into a numeric vector;

[0094] The initial business data set after the data preprocessing is used as the target business data set to obtain the target business data set.

[0095] According to one embodiment of the present invention, Figure 4 Each unit in the abnormal data detection device shown can be individually or entirely combined into one or several other units to form a unit, or one (or more) of the units can be further divided into multiple smaller units in function to form a unit, which can achieve the same operation without affecting the realization of the technical effects of the embodiments of the present invention. The above-mentioned units are divided based on logical functions. In actual applications, the functions of one unit can also be implemented by multiple units, or the functions of multiple units can be implemented by one unit. In other embodiments of the present invention, any abnormal data detection device may also include other units. In actual applications, these functions can also be implemented with the assistance of other units, and can be implemented by the collaboration of multiple units.

[0096] According to another embodiment of the present invention, the program can be executed by running a program on a general electronic device such as a computer including a central processing unit (CPU), a random access memory (RAM), a read-only memory (ROM), and other processing elements and storage elements. Figure 1 or Figure 3 A computer program (including program code) for each step involved in the corresponding method shown in Figure 4The abnormal data detection device shown in and the abnormal data detection method of the embodiment of the present invention are implemented. The computer program can be recorded on, for example, a computer storage medium, and loaded into the above-mentioned electronic device through the computer storage medium and run therein.

[0097] The embodiment of the present invention can obtain a target business data set, which includes multiple target business data, and one target business data includes field information of each target field in M target fields, where M is a positive integer; and call the target clustering model to perform clustering processing on the target business data set to obtain a clustering result. Based on this, the clustering results can be merged to obtain a clustering result after the merge processing; and abnormal data detection can be performed on the clustering result after the merge processing to obtain an abnormal data detection result. It can be seen that the embodiment of the present invention can perform abnormal data detection through the clustering result after the merge processing, without assuming that the data obeys a certain specific distribution (such as a normal distribution, etc.), thereby performing abnormal data detection more flexibly, so as to effectively improve the accuracy of abnormal data detection and obtain an abnormal data detection result with higher accuracy.

[0098] Based on the description of the above method embodiment and apparatus embodiment, the exemplary embodiments of the present invention further provide an electronic device, comprising: at least one processor; and a memory communicatively connected to the at least one processor. The memory stores a computer program executable by the at least one processor, and when executed by the at least one processor, the computer program causes the electronic device to perform a method according to an embodiment of the present invention.

[0099] Exemplary embodiments of the present invention further provide a non-transitory computer-readable storage medium storing a computer program, wherein the computer program, when executed by a processor of a computer, is used to cause the computer to perform a method according to an embodiment of the present invention.

[0100] An exemplary embodiment of the present invention further provides a computer program product, comprising a computer program, wherein when the computer program is executed by a processor of a computer, the computer is configured to cause the computer to perform a method according to an embodiment of the present invention.

[0101] refer to Figure 5, a block diagram of an electronic device 500 that can serve as a server or client of the present invention will now be described, which is an example of a hardware device that can be applied to various aspects of the present invention. The electronic device is intended to represent various forms of digital electronic computer devices, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or required herein.

[0102] like Figure 5 As shown, the electronic device 500 includes a computing unit 501, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 502 or a computer program loaded from a storage unit 508 into a random access memory (RAM) 503. Various programs and data required for the operation of the electronic device 500 can also be stored in the RAM 503. The computing unit 501, the ROM 502, and the RAM 503 are connected to each other via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.

[0103] Multiple components within electronic device 500 are connected to I / O interface 505, including an input unit 506, an output unit 507, a storage unit 508, and a communication unit 509. Input unit 506 can be any type of device capable of inputting information into electronic device 500. Input unit 506 can receive input numeric or character information and generate key input signals related to user settings and / or function control of the electronic device. Output unit 507 can be any type of device capable of presenting information and may include, but is not limited to, a display, a speaker, a video / audio output terminal, a vibrator, and / or a printer. Storage unit 508 may include, but is not limited to, a magnetic disk or an optical disk. Communication unit 509 allows electronic device 500 to exchange information / data with other devices via computer networks such as the Internet and / or various telecommunication networks and may include, but is not limited to, a modem, a network card, an infrared communication device, a wireless communication transceiver and / or chipset, such as a Bluetooth™ device, a WiFi device, a WiMax device, a cellular communication device, and / or the like.

[0104] The computing unit 501 can be various general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 501 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units that run machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 501 performs the various methods and processes described above. For example, in some embodiments, the abnormal data detection method can be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as a storage unit 508. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 500 via the ROM 502 and / or the communication unit 509. In some embodiments, the computing unit 501 can be configured to perform the abnormal data detection method in any other appropriate manner (e.g., by means of firmware).

[0105] The program code for implementing the method of the present invention can be written in any combination of one or more programming languages. Such program code can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0106] In the context of the present invention, machine-readable medium can be a tangible medium that can contain or store a program for use with an instruction execution system, device or equipment or used in combination with an instruction execution system, device or equipment. Machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. Machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared or semiconductor systems, devices or equipment, or any suitable combination of the foregoing. More specific examples of machine-readable storage media can include electrical connections based on one or more lines, portable computer disks, hard disks, random access memories (RAM), read-only memories (ROM), erasable programmable read-only memories (EPROM or flash memory), optical fibers, portable compact disk read-only memories (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0107] As used herein, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, apparatus, and / or device (e.g., a magnetic disk, an optical disk, a memory, a programmable logic device (PLD)) for providing machine instructions and / or data to a programmable processor, including machine-readable media that receive machine instructions as machine-readable signals. The term "machine-readable signal" refers to any signal used to provide machine instructions and / or data to a programmable processor.

[0108] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0109] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.

[0110] Computer systems may include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The client and server relationship arises through computer programs running on the respective computers and having a client-server relationship to each other.

[0111] Furthermore, it should be understood that the above disclosure is only a preferred embodiment of the present invention and certainly cannot be used to limit the scope of the present invention. Therefore, equivalent changes made according to the claims of the present invention are still within the scope of the present invention.

Claims

1. A method for detecting abnormal data, characterized in that: include: Acquire a target business data set, where the target business data set includes multiple target business data, and one target business data includes field information of each target field in M target fields, where M is a positive integer; Calling the target clustering model to perform clustering processing on the target business data set to obtain a clustering result; Merging the clustering results to obtain a merged clustering result; Anomaly data detection is performed on the clustering results after the merging process to obtain an abnormal data detection result.

2. The method according to claim 1, characterized in that The method further comprises: Determine a training service data set, where the training service data set includes a plurality of training service data; Based on the training business data set, respectively verify the clustering effect of each clustering model in the plurality of clustering models to obtain clustering effect indication information of each clustering model; wherein the clustering effect indication information of a clustering model is used to indicate the clustering effect of the corresponding clustering model; Based on the clustering effect indication information of each clustering model, a clustering model with the best clustering effect is selected from the multiple clustering models, and the selected clustering model is used as the target clustering model.

3. The method according to claim 2, characterized in that The determining of the training service data set includes: The target business data set is added to the training business data set to determine the training business data set.

4. The method according to any one of claims 1 to 3, characterized in that The merging of the clustering results to obtain the merged clustering results includes: When the clustering result includes multiple clusters, respectively calculating cluster indication information of each of the multiple clusters; Based on the cluster indication information of the respective clusters, the clustering results are merged to obtain a merged clustering result, wherein the merged clustering result includes at least one target cluster.

5. The method according to any one of claims 1 to 3, characterized in that The performing abnormal data detection on the clustering result after the merging process to obtain the abnormal data detection result includes: Respectively determining the number of cluster samples of each target cluster in the clustering result after the merging process; Based on the number of cluster samples of each target cluster, an abnormal cluster is determined from the clustering result after the merge processing, and the abnormal cluster is added to the abnormal data detection result. The number of cluster samples of the abnormal cluster is less than the number of cluster samples of any target cluster other than the abnormal cluster in the clustering result after the merge processing.

6. The method according to any one of claims 1 to 3, characterized in that The acquiring of the target business data set includes: Acquire an initial business data set, wherein the initial business data set includes a plurality of initial business data, and one initial business data includes initial field information of each target field; Determine a data preprocessing strategy, and perform data preprocessing on the initial business data set according to the data preprocessing strategy to obtain the initial business data set after data preprocessing; wherein the data preprocessing strategy includes a text-to-numeric vector strategy, and the text-to-numeric vector strategy is used to instruct to convert text data into a numeric vector; The initial business data set after the data preprocessing is used as the target business data set to obtain the target business data set.

7. An abnormal data detection device, characterized in that: The device comprises: an acquiring unit, configured to acquire a target business data set, wherein the target business data set includes a plurality of target business data, and one target business data includes field information of each target field in M target fields, where M is a positive integer; A processing unit, configured to call a target clustering model, perform clustering processing on the target business data set, and obtain a clustering result; The processing unit is further configured to merge the clustering results to obtain a merged clustering result; The processing unit is further configured to perform abnormal data detection on the clustering result after the merging process to obtain an abnormal data detection result.

8. The device according to claim 7, characterized in that The processing unit is further configured to: Determine a training service data set, where the training service data set includes a plurality of training service data; Based on the training business data set, respectively verify the clustering effect of each clustering model in the plurality of clustering models to obtain clustering effect indication information of each clustering model; wherein the clustering effect indication information of a clustering model is used to indicate the clustering effect of the corresponding clustering model; Based on the clustering effect indication information of each clustering model, a clustering model with the best clustering effect is selected from the multiple clustering models, and the selected clustering model is used as the target clustering model.

9. An electronic device, characterized in that: include: processor; as well as Memory for storing programs, The program includes instructions, which, when executed by the processor, cause the processor to perform the method according to any one of claims 1 to 6.

10. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to enable a computer to execute the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Enterprise abnormal business data detection method and device, equipment and storage medium

    CN119293680A

  • Method and apparatus and electronic device for clustering

    US20180129727A1

  • Anomaly detection and anomalous patterns identification

    US20230359706A1

Cited By

  • Intelligent financial risk control optimization method and system for credible multi-source data fusion

    CN121544414A