Data anomaly detection method and device, equipment, storage medium and program product

By using feature mapping and diffusion models to generate abnormal samples, the problems of scarcity and high training cost in structured data abnormal detection are solved, and the abnormal detection effect is improved.

CN120492969APending Publication Date: 2025-08-15BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510571667.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-30
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

The prior art is difficult to effectively detect abnormalities of structured data, mainly due to the lack of abnormal structured sample data and high training costs, resulting in poor detection results.

Method used

By performing feature mapping of normal structured sample data, sample mapping data is generated, and the generation of abnormal structured samples is guided by using the diffusion model, the sample number is expanded, and then classifier training is carried out to generate the target classifier for abnormal detection.

Benefits of technology

It realizes that the classifier is effectively acquired and trained in the absence of abnormal structured samples, improving the accuracy and efficiency of structured data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120492969A_ABST
    Figure CN120492969A_ABST
Patent Text Reader

Abstract

The invention discloses a data anomaly detection method and device, equipment, a storage medium and a program product, and relates to the technical field of data processing.The method comprises the steps that a normal structured sample data set is obtained, and the normal structured sample data set comprises multiple pieces of normal structured sample data; performing feature mapping on the normal structured sample data to obtain sample mapping data corresponding to the normal structured sample data; according to the sample distribution characteristics of each piece of sample mapping data, guiding the generation process of an abnormal structured sample to obtain abnormal structured sample data corresponding to the normal structured sample data; performing classifier training by using the normal structured sample data and the abnormal structured sample data to obtain a target classifier; and performing anomaly detection on the target business data by using the target classifier to generate an anomaly detection result. By implementing the technical scheme of the invention, effective generation of abnormal structured sample data is realized, and the abnormal detection effect of the structured data is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of data processing technology, and in particular to a data anomaly detection method, apparatus, device, storage medium, and program product. Background Art

[0002] Anomaly detection for structured data (such as tabular data) is an important data analysis technique. Its core goal is to identify anomalous data that deviates significantly from normal data within structured data (a two-dimensional table consisting of rows and columns). Because anomalous structured data is difficult to obtain, it is difficult to detect anomalies using pre-trained universal models, which affects the effectiveness of anomaly detection for structured data. Summary of the Invention

[0003] In view of this, the present disclosure provides a data anomaly detection method, apparatus, device, storage medium and program product to solve the problem of poor anomaly detection effect on structured data.

[0004] In a first aspect, the present disclosure provides a data anomaly detection method, comprising: obtaining a normal structured sample data set, the normal structured sample data set including multiple normal structured sample data; performing feature mapping on each normal structured sample data to obtain sample mapping data corresponding to each normal structured sample data; guiding the generation process of abnormal structured samples according to the sample distribution characteristics of each sample mapping data to obtain abnormal structured sample data corresponding to the normal structured sample data; using the normal structured sample data and the abnormal structured sample data to train a classifier to obtain a target classifier; using the target classifier to perform anomaly detection on target business data to generate anomaly detection results.

[0005] In a second aspect, the present disclosure provides a data anomaly detection device, including: an acquisition module for acquiring a normal structured sample data set, the normal structured sample data set including multiple normal structured sample data; a feature mapping module for performing feature mapping on each normal structured sample data to obtain sample mapping data corresponding to each normal structured sample data; an anomaly generation module for guiding the generation process of abnormal structured samples according to the sample distribution characteristics of each sample mapping data to obtain abnormal structured sample data corresponding to the normal structured sample data; a classification training module for performing classifier training using normal structured sample data and abnormal structured sample data to obtain a target classifier; an anomaly detection module for performing anomaly detection on target business data using the target classifier to generate anomaly detection results.

[0006] In a third aspect, the present disclosure provides a computer device comprising: a memory and a processor, the memory and the processor being communicatively connected to each other, the memory storing computer instructions, and the processor executing the computer instructions to thereby execute the data anomaly detection method of the first aspect or any corresponding embodiment thereof.

[0007] In a fourth aspect, the present disclosure provides a computer-readable storage medium having computer instructions stored thereon, the computer instructions being used to enable a computer to execute the data anomaly detection method of the first aspect or any corresponding embodiment thereof.

[0008] In a fifth aspect, the present disclosure provides a computer program product, including computer instructions, which are used to enable a computer to execute the data anomaly detection method of the first aspect or any corresponding embodiment thereof.

[0009] The data anomaly detection method, apparatus, device, storage medium and program product provided by the embodiments of the present disclosure generate corresponding sample mapping data by performing feature mapping on normal structured sample data, and generate abnormal structured sample data according to the sample distribution characteristics of the sample mapping data, thereby expanding the number of samples of abnormal structured sample data and achieving effective acquisition of abnormal structured sample data. Subsequently, the classifier is trained using the normal structured sample data and the abnormal structured sample data to generate a target classifier for anomaly detection. Since the training of the target classifier has sufficient normal structured sample data and abnormal structured sample data, when the target business data is subjected to anomaly detection by the target classifier, the abnormal structured data present in the target business data can be accurately detected, effectively improving the anomaly detection effect of the structured data. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] In order to more clearly illustrate the specific embodiments of the present disclosure or the technical solutions in the related technologies, the following briefly introduces the drawings required for use in the specific embodiments or related technical descriptions. Obviously, the drawings described below are some embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0011] Figure 1 is a schematic diagram of an application scenario according to an embodiment of the present disclosure;

[0012] Figure 2 is a flow chart of a data anomaly detection method according to an embodiment of the present disclosure;

[0013] Figure 3 is a flow chart of another data anomaly detection method according to an embodiment of the present disclosure;

[0014] Figure 4 is a flowchart of another data anomaly detection method according to an embodiment of the present disclosure;

[0015] Figure 5 is a structural block diagram of a data anomaly detection device according to an embodiment of the present disclosure;

[0016] Figure 6 Schematic diagram of the hardware structure of a computer device according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0017] To make the purpose, technical solutions, and advantages of the embodiments of the present disclosure more clear, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the drawings in the embodiments of the present disclosure. Obviously, the described embodiments are part of the embodiments of the present disclosure, not all of the embodiments. Based on the embodiments of the present disclosure, all other embodiments obtained by those skilled in the art without making creative efforts shall fall within the scope of protection of the present disclosure.

[0018] It is understandable that before using the technical solutions disclosed in the various embodiments of this disclosure, the type, scope of use, usage scenarios, etc. of the personal information involved in this disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.

[0019] For example, in response to a user's active request, a prompt message is sent to the user to clearly inform the user that the operation requested will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the computer device, application, server, storage medium, or other software or hardware that performs the operation of the disclosed technical solution based on the prompt message.

[0020] As an optional but non-limiting implementation, in response to receiving a user's active request, the prompt information may be sent to the user in the form of a pop-up window, in which the prompt information may be presented in text form. Furthermore, the pop-up window may also contain a selection control for the user to select "agree" or "disagree" to provide personal information to the computer device.

[0021] It is understandable that the above notification and user authorization process are merely illustrative and do not limit the implementation of the present disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of the present disclosure.

[0022] It is understandable that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) must comply with the requirements of relevant laws, regulations and relevant provisions.

[0023] Anomaly detection in structured data (such as tabular data) is an important data analysis technique widely used in many fields. This type of data typically contains numerical, categorical, or mixed features and has clearly defined fields. Due to its structured, interpretable, and easy-to-process characteristics, structured data (such as tabular data) has become a primary data carrier for businesses, research institutions, and public services.

[0024] Network traffic anomaly detection is a core task in network security, aiming to identify malicious attacks within the network. Traffic attacks are very common in the real world, such as DDoS attacks, port scans, data scraping, and other common methods used by the black market. These attacks pose significant challenges to platform stability and data security, making it crucial to identify these anomalies within daily traffic.

[0025] Traditional threshold- or rule-based detection methods (such as Z-score and boxplot analysis) have been unable to cope with complex data distributions with high dimensions, nonlinearity, and dynamic evolution. There is an urgent need to combine machine learning and deep learning technologies to achieve more intelligent anomaly identification. In actual business applications, compared with other machine learning tasks, anomaly detection of structured data (such as tabular data) usually has the following characteristics: (1) The proportion of abnormal structured samples is usually less than 5%, which is extremely rare, which leads to a significant increase in the cost of manual labeling; (2) Attackers often make adaptive adjustments to the defense system and change their own performance to circumvent the recognition algorithm, making abnormal structured samples more difficult to detect; (3) Structured data such as tabular data are often private data. Due to its confidentiality and column heterogeneity, it is difficult to obtain a large amount of public valid data. Compared with anomaly detection in the direction of text, image, etc., the cost of obtaining general data in the field for learning is too high; (4) It is easy to do upstream work for model pre-training in the direction of text, image, etc., but structured data such as tabular data has the characteristics of heterogeneity, privacy, manual feature design and frequent updates, making it difficult to produce universal detection models through pre-training of structured data, that is, there are strong restrictions on modeling techniques.

[0026] For different types of data carriers, the anomaly detection methods are different. Table 1 shows the differences between various data carrier anomaly types and identification methods.

[0027] Table 1 Comparison of anomaly detection for different types of data

[0028]

[0029] For anomaly detection in structured data, unsupervised anomaly detection methods, supervised anomaly detection methods, and semi-supervised anomaly detection methods are usually used. However, these detection methods are either affected by the scarcity of labeled samples or have relatively high training costs, making it difficult to quickly update and iterate anomaly detection, and the time cost in practical applications is high.

[0030] Based on this, the technical solution disclosed in the present invention can effectively generate abnormal structured sample data through a diffusion model in the absence of abnormal structured samples, and then distinguish between normal samples and abnormal structured samples by training a classifier, thereby improving the anomaly detection effect of structured data.

[0031] As an optional application scenario of the embodiment of the present disclosure, Figure 1 As shown, this optional application scenario mainly includes a feature mapping node 101, an abnormal structured sample generation node 102, and a classifier training node 103. Through the feature mapping node 101, the abnormal structured sample generation node 102, and the classifier training node 103, a normal / abnormal structured sample classifier for this structured sample data is generated, and finally the trained classifier is deployed to the online anomaly detection process of actual business data.

[0032] Specifically, the feature mapping node 101 is used to transform the original normal structured sample data into a linear space, mapping it to a different linear space as much as possible. The abnormal structured sample generation node 102 is used to generate abnormal structured sample data based on the mapping data generated in the linear space. The classifier training node 103 is used to perform classification training based on the normal structured sample data and the generated abnormal structured sample data, producing a classifier network that can be used normally online. The abnormal structured sample generation node 102 and the classifier training node 103 perform interactive training to obtain abnormal structured sample data that better meets business needs and a classifier that meets anomaly detection requirements.

[0033] The feature mapping node 101, the abnormal structured sample generation node 102 and the classifier training node 103 are deployed in a computer device, and each node has computing resources or computing power. The computer device can be a device with computing power, for example, the computer device can be provided with a processor and a memory, etc., or can be equipped with a dedicated accelerator (such as a graphics processing unit (GPU)). In addition, the computer device can store and maintain data. Examples of computers can include supercomputers, personal computers, laptop computers, vehicle-mounted computing devices, mobile devices (such as smart phones, tablet computers, etc.) or a combination of any one or more of the above devices. It should be understood that the computer devices described herein are exemplary and non-restrictive, and for example, other different types of computer devices can also be used.

[0034] According to an embodiment of the present disclosure, an embodiment of a data anomaly detection method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0035] In this embodiment, a data anomaly detection method is provided, which can be used in computer devices such as computers, tablet computers, servers, etc. Figure 2 is a flow chart of a data anomaly detection method according to an embodiment of the present disclosure. Figure 2 As shown, the process includes the following steps:

[0036] Step S201: Acquire a normal structured sample data set, where the normal structured sample data set includes a plurality of normal structured sample data.

[0037] Normal structured sample data is structured data collected during business operations, and a normal structured sample dataset is a sample dataset formed by multiple pieces of normal structured sample data. Specifically, this normal structured sample data is heterogeneous, private, and frequently updated, such as tabular data. This makes it difficult to detect anomalies using pre-trained universal models.

[0038] Specifically, obtain the type of normal structured sample data to be collected, such as numerical data, time series data, etc., monitor the operating status of the business in real time, obtain the operating results generated by the business under various operating states, and collect multiple normal structured sample data of the corresponding type from the operating results to obtain a normal structured sample data set.

[0039] Step S202 : performing feature mapping on each normal structured sample data to obtain sample mapping data corresponding to each normal structured sample data.

[0040] Feature mapping is used to transform normal structured sample data into a linear space. Sample mapping data is the linear transformation data generated in the linear space. Specifically, due to the heterogeneity of normal structured sample data, feature mapping is required for the normal structured sample data. In this case, the features of each normal structured sample data can be transformed into a linear space using a linear transformation method. This allows the features of the normal structured sample data to be mapped into different linear spaces, resulting in sample mapping data generated by the feature mapping of each normal structured sample data.

[0041] Step S203 : guiding the generation process of abnormal structured samples according to the sample distribution characteristics of each sample mapping data, and obtaining abnormal structured sample data corresponding to the normal structured sample data.

[0042] As described above, normal structured sample data is heterogeneous. When generating abnormal structured sample data, the sample distribution must be considered, thus requiring a method that can generate samples with the same distribution. Specifically, since the diffusion model primarily considers sample distribution rather than Euclidean distance, it can be used here as a method for generating samples with the same distribution.

[0043] After completing the feature mapping of normal structured sample data, we can analyze the sample distribution characteristics of the sample mapping data and use the diffusion model to guide the generation of abnormal structured samples according to the sample distribution characteristics, so that the diffusion model can generate samples that are similar to but different from the known sample distribution. These samples are abnormal structured sample data.

[0044] Step S204: Perform classifier training using the normal structured sample data and the abnormal structured sample data to obtain a target classifier.

[0045] The target classifier is used to distinguish between normal structured sample data and abnormal structured sample data. Specifically, after the abnormal structured sample data is generated, the sample features of the normal structured sample data and the sample features of the abnormal structured sample data are extracted, and a classification network (such as an MLP network to implement the network structure of the classification network) is used to perform classification training on the original normal structured sample data and the generated abnormal structured sample data. During the training process, the corresponding loss function is used to treat the label of the normal structured sample data as 0 and the label of the generated abnormal structured sample data as 1 for gradient update to obtain the target classifier discriminator(x).

[0046] Step S205: Utilize the target classifier to perform anomaly detection on the target business data and generate anomaly detection results.

[0047] The trained target classifier discriminator(x) is deployed online to monitor the operating status of each target business online in real time. The target business data generated by each target business during operation is input into the target classifier. The target classifier can then perform anomaly detection on the target business data generated by the target business, identify abnormal business data, and generate corresponding anomaly detection results.

[0048] The data anomaly detection method provided in this embodiment performs feature mapping on normal structured sample data to generate corresponding sample mapping data, and generates abnormal structured sample data according to the sample distribution characteristics of the sample mapping data, thereby expanding the number of samples of abnormal structured sample data and achieving effective acquisition of abnormal structured sample data. Subsequently, the classifier is trained using the normal structured sample data and the abnormal structured sample data to generate a target classifier for anomaly detection. Since the training of the target classifier has sufficient normal structured sample data and abnormal structured sample data, when the target classifier is used to perform anomaly detection on the target business data, the abnormal structured data present in the target business data can be accurately detected, effectively improving the anomaly detection effect of the structured data.

[0049] In this embodiment, a data anomaly detection method is provided, which can be used in computer devices such as computers, tablet computers, servers, etc. Figure 3 is a flow chart of a data anomaly detection method according to an embodiment of the present disclosure. Figure 3 As shown, the process includes the following steps:

[0050] Step S301: Acquire a normal structured sample data set, which includes a plurality of normal structured sample data. For details, please refer to the description of the corresponding steps in the above embodiment, which will not be repeated here.

[0051] Step S302 : performing feature mapping on each normal structured sample data to obtain sample mapping data corresponding to each normal structured sample data.

[0052] Specifically, the above step S302 includes:

[0053] Step S3021: Acquire characteristic attributes corresponding to normal structured sample data.

[0054] Feature attributes are the unique characteristics of normal structured sample data, including uniqueness, discreteness, missing values, numerical features (such as continuous values and discrete values), categorical features (such as hierarchical, sequence, and nominal), data distribution, etc.

[0055] Specifically, normal structured sample data carries a variety of data contents. By parsing the data structure of the normal structured sample data, the various data fields carried by the normal structured sample data are determined, and characteristic attributes that can characterize the normal structured sample data are extracted from the data contents corresponding to each data field.

[0056] Step S3022 , using a preset feature mapping method, performs feature transformation on the normal structured sample data based on feature attributes, maps the normal structured sample data to different linear spaces, and obtains sample mapping data generated in each linear space.

[0057] When converting normal structured sample data into linear space, it is necessary to make the normal structured sample data reach different linear spaces through mapping as much as possible, and the characteristic attributes of each normal structured sample data have the property of being isotropic, so as to effectively maintain the stability of the sample feature gradient in the subsequent abnormal structured sample generation process and improve the generation stability of subsequent abnormal structured samples.

[0058] The preset feature mapping method is a pre-set feature linear transformation method, such as using a multilayer perceptron (MLP) network to perform linear transformation of features. When performing linear transformation of features through the MLP network,

[0059] Specifically, if the normal structured sample data is: X∈R n×d , where n represents the number of samples of normal structured sample data, d represents the data dimension of normal structured sample data, and the preset feature mapping method is encoder(x). The preset feature mapping method can be a feature transformation model generated by MLP network training. The method of feature transformation for normal structured sample data is expressed as follows:

[0060]

[0061] Where W is the transformation matrix for linear transformation, each row of the input matrix W consists of all the feature attributes of a training case; σ(X) is a nonlinear activation function; is the sample data transformed into the linear space, that is, the sample mapping data.

[0062] Thus, according to the above method, the normal structured sample data can be mapped to different linear spaces to obtain the sample mapping data of the normal structured sample data in each linear space. Similarly, the sample mapping data of each normal structured sample data in different linear spaces can be obtained.

[0063] Step S303 : guiding the generation process of abnormal structured samples according to the sample distribution characteristics of each sample mapping data, and obtaining abnormal structured sample data corresponding to the normal structured sample data.

[0064] Specifically, the above step S303 includes:

[0065] Step S3031: Add disturbance data to the normal structured sample data according to the sample distribution characteristics.

[0066] Perturbed data represents randomly changing or erroneous data, including random noise and multiplicative noise. When adding perturbed data, ensure that the perturbed data retains the same distribution and characteristics as the original, normal, structured sample data. Specifically, based on the sample distribution characteristics of the sample mapping data, perturbed data is superimposed on the normal, structured sample data to increase sample data diversity.

[0067] Step S3032: Use the disturbance data to guide the generation process of the abnormal structured sample, and constrain the generation process of the abnormal structured sample according to a preset constraint method to obtain abnormal structured sample data.

[0068] The sample distribution of the abnormal structured sample data is similar to the sample distribution of the normal structured sample data, and the abnormal structured sample data is different from the normal structured sample data.

[0069] The preset constraint method is set to constrain the abnormal structured sample generation process. The preset constraint method can be implemented using a loss function, and the preset constraint method can ensure that the sample distribution characteristics of the abnormal structured sample data are similar to those of the normal structured sample data.

[0070] When mapping data in the sample After superimposing the perturbation data, the perturbed sample data is obtained In the process of generating abnormal structured samples using perturbation data, the generation process of abnormal structured samples is constrained according to a preset constraint method, so that the generated abnormal structured sample data is different from the normal structured sample data, and the sample distribution of the abnormal structured sample data is similar to the sample distribution of the normal structured sample data.

[0071] In some optional implementations, the preset constraint method includes a first constraint and a second constraint. The generation process of the abnormal structured sample is constrained according to the preset constraint method to obtain the abnormal structured sample data, including:

[0072] Step a1: Use the first constraint to perform type constraints on the normal structured sample data and the disturbance data to generate first sample data. The first sample data has the same type as the normal structured sample data.

[0073] Step a2: Using the second constraint to constrain the feature distribution of the first sample data and the disturbance data, to generate second sample data. The second sample data has a sample distribution feature similar to that of the normal structured sample data and is different from the normal structured sample data.

[0074] Step a3: determine the second sample data as abnormal structured sample data.

[0075] The first constraint is the original loss function of the diffusion model. The generation process of abnormal structured samples is constrained by the first constraint so that the abnormal structured sample data generated according to the perturbation data has the same sample type as the normal structured sample data.

[0076] The second constraint is a new loss function added to the original loss function of the diffusion model to constrain the sample distribution. The generation process of abnormal structured samples is constrained according to the second constraint so that the abnormal structured sample data generated according to the perturbation data is similar to the normal structured sample data in sample distribution, thereby ensuring the effectiveness of the generation of abnormal structured sample data and avoiding the generated abnormal structured sample data being meaningless random noise.

[0077] Specifically, when generating abnormal structured sample data using a preset constraint method, a first constraint can be used to constrain the types of normal structured sample data and perturbation data, generating first sample data of the same type as the normal structured sample data. Simultaneously, a second constraint can be used to constrain the feature distribution of the first sample data and perturbation data, generating second sample data with sample distribution characteristics similar to those of the normal structured sample data. Thus, by constraining the generation of abnormal structured samples according to both the original loss function and the new loss function, second sample data is obtained that has a sample distribution similar to that of the normal structured sample data but is different from the normal structured sample data. This second sample data is the abnormal structured sample data determined according to the joint constraints of the original loss function and the new loss function.

[0078] In the above implementation, by designing corresponding constraints to jointly constrain the generation of abnormal structured samples according to the original loss function and the new loss function, the effective generation of abnormal structured sample data is achieved, and the problem of high cost of obtaining abnormal structured sample data is solved.

[0079] In some optional embodiments, the sample distribution features include sample variation features and sample concentration features, wherein the sample variation features can be represented by the covariance matrix of the sample data, and the sample concentration features can be represented by the mean vector of the sample data.

[0080] Specifically, perturbation data is added to the normal structured sample data to generate perturbed sample data According to the disturbed sample data Generate abnormal structured sample data. In the process of generating abnormal structured sample data, the generation of abnormal structured samples is constrained by the original loss function and the new loss function. When the training gradient is updated each time, the abnormal structured sample data generated by the diffusion model will produce one of the following constraint effects. Among them, μ true and σtrue are the sample mean and sample covariance of normal structured sample data, μ pred and σ pred are the sample mean and sample covariance of the abnormal structured sample data generated by the diffusion model. The specific constraint effects are as follows:

[0081] (1) The sample change characteristics of abnormal structured sample data and normal structured sample data are similar, but the characteristics of the sample sets are not similar. The specific constraints are as follows:

[0082]

[0083] (2) The sample change characteristics of abnormal structured sample data and normal structured sample data are different, but the characteristics of the sample sets are similar. The specific constraints are as follows:

[0084]

[0085] (3) The sample change characteristics of abnormal structured sample data and normal structured sample data are similar, and the sample set characteristics are similar. The specific constraints are as follows:

[0086]

[0087] The dissimilar effect is constrained by adding a random noise vector ∈~N(0,1), M mask It is a Bernoulli matrix, that is, some random elements are 0 and some random elements are 1, which means that some dimensions are constrained instead of all dimensions.

[0088] The similarity constraints of the above three effects are all defined as the bi-norm distance between the two matrices (vectors) is as small as possible. Under the joint constraints of the new loss function and the original loss function, the diffusion model can generate samples that are close to but different from the known sample distribution, that is, abnormal structured sample data that are similar to the sample distribution characteristics of normal structured sample data and different from the normal structured sample data.

[0089] In addition, the degree to which the abnormal structured samples generated by the diffusion model deviate from the normal samples can be adjusted by controlling the weight ratio of the original loss function and the new loss function.

[0090] In order to ensure the completeness of the abnormal structured sample data set, the diffusion model can also be controlled to generate abnormal structured sample data that is different from the sample change characteristics and sample concentration characteristics of the normal structured sample data, so as to ensure the abnormal detection effect of the subsequent classifier.

[0091] Step S304: Use the normal structured sample data and the abnormal structured sample data to perform classifier training to obtain a target classifier. For details, please refer to the description of the corresponding steps in the above embodiment, which will not be repeated here.

[0092] Step S305: Detect anomalies in the target service data using the target classifier to generate anomaly detection results. For details, please refer to the description of the corresponding steps in the above embodiment, which will not be repeated here.

[0093] The data anomaly detection method provided in this embodiment maps normal structured sample data to different linear spaces according to the characteristic attributes corresponding to the normal structured sample data. This ensures that the characteristics of each normal structured sample data during the feature mapping process are homogeneous, allowing the subsequent generation process of abnormal structured sample data to effectively maintain gradient stability, thereby improving the stability of the generation of abnormal structured sample data. In the normal structured sample data, corresponding perturbation data is added according to the sample distribution characteristics. The perturbation data is used to guide the generation process of abnormal structured samples, ensuring the similarity in distribution characteristics between abnormal structured samples and normal structured samples, thereby improving the effectiveness of the generation of abnormal structured samples.

[0094] In this embodiment, a data anomaly detection method is provided, which can be used in computer devices such as computers, tablet computers, servers, etc. Figure 4 is a flow chart of a data anomaly detection method according to an embodiment of the present disclosure. Figure 4 As shown, the process includes the following steps:

[0095] Step S401: Acquire a normal structured sample data set, which includes a plurality of normal structured sample data. For details, please refer to the description of the corresponding steps in the above embodiment, which will not be repeated here.

[0096] Step S402: Perform feature mapping on each normal structured sample data to obtain sample mapping data corresponding to each normal structured sample data. Please refer to the description of the corresponding steps in the above embodiment for details, which will not be repeated here.

[0097] Step S403: guide the generation process of abnormal structured samples according to the sample distribution characteristics of each sample mapping data to obtain abnormal structured sample data corresponding to the normal structured sample data. For details, please refer to the description of the corresponding steps in the above embodiment, which will not be repeated here.

[0098] Step S404: Perform classifier training using the normal structured sample data and the abnormal structured sample data to obtain a target classifier.

[0099] Specifically, the above step S404 includes:

[0100] Step S4041 : Perform iterative classification training using normal structured sample data and abnormal structured sample data to obtain an iterative classifier.

[0101] The normal structured sample data and the initially generated abnormal structured sample data are used for classification training to obtain an initial classifier. Subsequently, the initial classifier is iteratively trained for sample classification using the normal structured sample data and the abnormal structured sample data generated by iterative optimization to obtain the optimized iterative classifier in each iterative round.

[0102] Step S4042 : According to the sample detection results generated by the iterative classifier and the generation process of abnormal structured sample data, the generation process of abnormal structured sample data and the classifier training process are interactively trained to obtain a target classifier.

[0103] The iterative classifier generated in each iteration can classify both normal structured sample data and abnormal structured sample data. Based on the sample detection results generated in each iteration, the generation process of abnormal structured sample data can be trained to ensure that the generation of abnormal structured sample data is more consistent with the actual anomalies that may occur during business operations. Subsequently, the classifier is trained using the newly generated abnormal structured sample data and normal structured sample data. This process is repeated repeatedly, interactively training the generation of abnormal structured sample data and the classifier until the classifier can meet the sample classification requirements, at which point iterations are terminated and the corresponding target classifier is obtained.

[0104] In some optional implementations, the above step S4042 includes:

[0105] Step b1, freeze the feature mapping parameters and perturbation parameters for abnormal structured sample generation in the current iteration; wherein, the feature mapping parameters are used to map normal structured sample data to the target linear space; the abnormal structured sample generation parameters are used to control the process of generating corresponding abnormal structured sample data according to normal structured sample data.

[0106] Step b2: using the cross-training constraint parameters to update the classification parameters of the iterative classifier, to obtain the iterative classifier with updated classification parameters.

[0107] Step b3: freeze the classification parameters of the iterative classifier, and use the abnormal structured sample generation constraint parameters and the cross-training constraint parameters to update the feature mapping parameters and the perturbation parameters generated by the abnormal structured sample.

[0108] Since feature mapping converts normal structured sample data into different linear spaces, and then uses the sample mapping data generated in different linear spaces to generate abnormal structured sample data, the sample mapping data generated in different linear spaces can be evaluated based on the classification detection results of subsequent classifiers to determine the linear space with better feature mapping effect.

[0109] The cross-training constraint parameter is used to constrain the classification parameters of the classifier. The cross-training constraint parameter can be characterized by the cross-entropy loss function. Specifically, an interactive training mode can be adopted when performing feature mapping, generating abnormal structured samples, and training the classifier. Specifically, when training the classifier discriminator(x), the update of the feature mapping parameters in the encoder(x) and the perturbation parameter diffusion(x) during the abnormal structured sample generation process is frozen. The cross-training constraint parameter and the optimizer are used to update the classification parameters of the iterative classifier in the current iteration process to generate the iteratively updated classifier.

[0110] The abnormal structured sample generation constraint parameters are used to constrain the generation process of abnormal structured sample data. As described above, the abnormal structured sample generation constraint parameters can be represented by the original loss function of the diffusion model and the designed new loss function.

[0111] Specifically, when training the feature mapping parameters in encoder(x) and the parameters of the perturbation parameter diffusion(x) in the process of generating abnormal structured samples, a multi-task optimization mode can be adopted, and the loss function corresponding to the abnormal structured sample generation constraint parameters and the cross-training constraint parameters can be used to jointly constrain the update of the feature mapping parameters and the perturbation parameters generated by the abnormal structured samples. At the same time, the classification parameters of the iterative classifier in the current iteration round are frozen.

[0112] In the above implementation, the generation training of abnormal structured sample data and the classification training of the classifier are carried out through interactive training, so that the feature mapping of normal structured sample data, the generation of abnormal structured sample data, and the classification training of the classifier approach a state of convergence. As the training rounds deepen, the perturbation data added to the normal structured sample data is gradually reduced, allowing the classifier to focus more on distinguishing the details of the sample data.

[0113] Step S405: Detect anomalies in the target service data using the target classifier to generate anomaly detection results. For details, please refer to the description of the corresponding steps in the above embodiment, which will not be repeated here.

[0114] Step S406: If the target classifier identifies that there is abnormal data in the target business data, an abnormality assessment result and alarm information for the abnormal data are generated.

[0115] Abnormal data refers to the anomalies that exist in the target business data, such as single-point anomalies such as a surge in target business traffic, combined anomalies such as protocol and port mismatches, and statistical distribution anomalies such as a sudden drop in IP address entropy. The target classifier can distinguish between normal data and abnormal data based on sample features such as explicit associations between features and statistical laws (such as distribution deviations and conditional probability anomalies). If the target classifier determines that abnormal data exists in the target business data by detecting the target business data, it can generate an anomaly assessment result for the abnormal data (such as evaluating the anomaly detection situation through an anomaly score) and an alarm message. Among them, the alarm information is used to remind business-related personnel to inquire about the abnormal status and take timely processing measures based on the abnormal situation to avoid losses to the target business.

[0116] The data anomaly detection method provided in this embodiment uses normal structured sample data and abnormal structured sample data for iterative classification training to produce a corresponding iterative classifier. Then, according to the sample detection results generated by the iterative classifier and the generation process of the abnormal structured sample data, interactive training of the abnormal structured sample data generation process and the classifier training process is performed to ensure that the target classifier can accurately identify abnormal details and improve the anomaly detection effect in the target business data.

[0117] As a specific application embodiment of the present disclosure, normal tabular data is collected as normal structured sample data, feature mapping is performed on the normal tabular data to obtain corresponding tabular mapping data, and perturbation data is added to the tabular mapping data to jointly constrain the generation process of abnormal tabular data according to the original loss function of the diffusion model and the designed new loss function, thereby obtaining abnormal tabular data with distribution characteristics similar to those of normal tabular data, and the abnormal tabular data is different from the normal tabular data. Subsequently, classification training is performed on the normal tabular data and the abnormal tabular data to obtain a corresponding classifier. At the same time, the classification detection results of the classifier are used to reversely train the generation process of the abnormal tabular data, so as to continue training the classifier using the newly generated abnormal tabular data. The generation process of the abnormal tabular data and the training process of the classifier are interactively trained in sequence to obtain a target classifier. The final target classifier is deployed to an online server for actual application. When abnormal traffic is detected in the target business, the target classifier will output a relatively high score and generate corresponding alarm information. Relevant personnel of the target business can check the abnormal status in time through the alarm information to avoid losses to the target business.

[0118] This embodiment also provides a data anomaly detection device for implementing the above-mentioned embodiments and preferred embodiments. Details already described will not be repeated. As used below, the term "module" may refer to a combination of software and / or hardware that implements a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation using hardware, or a combination of software and hardware, is also possible and contemplated.

[0119] This embodiment provides a data anomaly detection device, such as Figure 5 Shown, including:

[0120] The acquisition module 501 is configured to acquire a normal structured sample data set, where the normal structured sample data set includes a plurality of normal structured sample data.

[0121] The feature mapping module 502 is configured to perform feature mapping on each normal structured sample data to obtain sample mapping data corresponding to each normal structured sample data.

[0122] The abnormality generation module 503 is used to guide the generation process of abnormal structured samples according to the sample distribution characteristics of each sample mapping data, and obtain abnormal structured sample data corresponding to the normal structured sample data.

[0123] The classification training module 504 is used to perform classifier training using the normal structured sample data and the abnormal structured sample data to obtain a target classifier.

[0124] The anomaly detection module 505 is used to perform anomaly detection on target business data using a target classifier to generate anomaly detection results.

[0125] In some optional implementations, the feature mapping module 502 includes:

[0126] The attribute acquisition unit is used to obtain the characteristic attributes corresponding to the normal structured sample data.

[0127] The linear transformation unit is used to perform feature transformation on normal structured sample data based on feature attributes using a preset feature mapping method, map the normal structured sample data to different linear spaces, and obtain sample mapping data generated in each linear space.

[0128] In some optional implementations, the exception generation module 503 includes:

[0129] The disturbance adding unit is used to add disturbance data to the normal structured sample data according to the sample distribution characteristics.

[0130] The generation guidance unit is configured to guide the generation process of the abnormal structured samples using the perturbation data and constrain the generation process of the abnormal structured samples according to a preset constraint method to obtain abnormal structured sample data. The sample distribution of the abnormal structured sample data is similar to the sample distribution of the normal structured sample data, and the abnormal structured sample data is different from the normal structured sample data.

[0131] In some optional implementations, the preset constraint mode includes a first constraint and a second constraint, and the generating guidance unit includes:

[0132] The type constraint subunit is used to perform type constraint on the normal structured sample data and the disturbance data using the first constraint to generate first sample data, where the first sample data has the same type as the normal structured sample data.

[0133] The feature distribution constraint subunit is used to use the second constraint to constrain the feature distribution of the first sample data and the disturbance data to generate second sample data. The sample distribution features of the second sample data are similar to those of the normal structured sample data and the second sample data are different from the normal structured sample data.

[0134] The abnormal structured sample determining subunit is configured to determine the second sample data as abnormal structured sample data.

[0135] In some optional embodiments, the sample distribution characteristics include sample change characteristics and sample concentration characteristics, the sample change characteristics of the abnormal structured sample data are similar to those of the normal structured sample data, and the characteristics in the sample concentration are dissimilar, or the sample change characteristics of the abnormal structured sample data are dissimilar to those of the normal structured sample data, and the characteristics in the sample concentration are similar, or the sample change characteristics of the abnormal structured sample data are partially similar to those of the normal structured sample data, and the characteristics in the sample concentration are partially similar.

[0136] In some optional implementations, the classification training module 504 includes:

[0137] The iterative training unit is used to perform iterative classification training using normal structured sample data and abnormal structured sample data to obtain an iterative classifier.

[0138] The interactive training unit is used to interactively train the abnormal structured sample data generation process and the classifier training process according to the sample detection results generated by the iterative classifier and the abnormal structured sample data generation process to obtain a target classifier.

[0139] In some optional embodiments, the interactive training unit includes:

[0140] The first parameter updating subunit is used to freeze the feature mapping parameters and the perturbation parameters for abnormal structured sample generation in the current iteration, and use the cross-training constraint parameters to update the classification parameters of the iterative classifier to obtain the iterative classifier after the classification parameters are updated; wherein, the feature mapping parameters are used to map the normal structured sample data to the target linear space; the abnormal structured sample generation parameters are used to control the process of generating corresponding abnormal structured sample data according to the normal structured sample data.

[0141] The second parameter updating subunit is used to freeze the classification parameters of the iterative classifier, and use the abnormal structured sample generation constraint parameters and the cross-training constraint parameters to update the feature mapping parameters and the perturbation parameters generated by the abnormal structured sample.

[0142] In some optional embodiments, the above device further includes:

[0143] The alarm module is used to generate an abnormality assessment result and alarm information for the abnormal data if the target classifier identifies the existence of abnormal data in the target business data.

[0144] The further functional description of each of the above modules and units is the same as that of the above corresponding embodiments and will not be repeated here.

[0145] The data anomaly detection device provided by the embodiments of the present disclosure can execute the data anomaly detection method provided by any embodiment of the present disclosure, and has the functional modules and beneficial effects corresponding to the execution method. By performing feature mapping on the normal structured sample data, the corresponding sample mapping data is generated, and the abnormal structured sample data is generated according to the sample distribution characteristics of the sample mapping data, thereby expanding the number of samples of the abnormal structured sample data and realizing the effective acquisition of the abnormal structured sample data. Subsequently, the classifier is trained using the normal structured sample data and the abnormal structured sample data to generate a target classifier for anomaly detection. Since the training of the target classifier has sufficient normal structured sample data and abnormal structured sample data, when the target business data is subjected to anomaly detection by the target classifier, the abnormal structured data existing in the target business data can be accurately detected, effectively improving the anomaly detection effect of the structured data.

[0146] Figure 6 A schematic diagram of the structure of a computer device provided in an embodiment of the present disclosure.

[0147] The following specific reference Figure 6, which shows a schematic diagram of the structure of a computer device suitable for implementing the embodiments of the present disclosure. The computer device may include a processor (e.g., a central processing unit, a graphics processing unit, etc.) 601, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 602 or a program loaded from a memory 608 into a random access memory (RAM) 603. Various programs and data required for the operation of the computer device are also stored in the RAM 603. The processor 601, ROM 602, and RAM 603 are connected to each other via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.

[0148] Typically, the following devices may be connected to the I / O interface 605: an input device 606 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 607 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 608 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 609. The communication device 609 may allow the computer device to communicate with other devices wirelessly or by wire to exchange data. Although Figure 6 A computer device having various devices is shown, but it should be understood that it is not required to implement or possess all of the devices shown, and more or fewer devices may be implemented or possessed instead.

[0149] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network via the communication device 609, or installed from the memory 608, or installed from the ROM 602. When the computer program is executed by the processor 601, the above-mentioned functions defined in the data anomaly detection method of the embodiment of the present disclosure are performed.

[0150] Figure 6 The computer device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present disclosure.

[0151] The embodiments of the present disclosure also provide a computer-readable storage medium. The above-mentioned method according to the embodiments of the present disclosure can be implemented in hardware, firmware, or implemented as a computer code that can be recorded in a storage medium, or implemented as a computer code that is originally stored in a remote storage medium or a non-temporary machine-readable storage medium and downloaded through a network and will be stored in a local storage medium, so that the method described herein can be stored in such software processing on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only storage memory, a random access memory, a flash memory, a hard disk or a solid-state drive, etc.; further, the storage medium can also include a combination of the above-mentioned types of memory. It can be understood that a computer, a processor, a microprocessor controller or programmable hardware includes a storage component that can store or receive software or computer code. When the software or computer code is accessed and executed by a computer, a processor or hardware, the data anomaly detection method shown in the above embodiment is implemented.

[0152] A portion of the present disclosure may be applied as a computer program product, such as a computer program instruction, which, when executed by a computer, can call or provide the method and / or technical solution according to the present disclosure through the operation of the computer. Those skilled in the art should understand that the form in which the computer program instruction exists in a computer-readable medium includes but is not limited to a source file, an executable file, an installation package file, etc. Accordingly, the way in which the computer program instruction is executed by the computer includes but is not limited to: the computer directly executes the instruction, or the computer compiles the instruction and then executes the corresponding compiled program, or the computer reads and executes the instruction, or the computer reads and installs the instruction and then executes the corresponding installed program. Here, the computer-readable medium can be any available computer-readable storage medium or communication medium that can be accessed by the computer.

[0153] Although the embodiments of the present disclosure have been described with reference to the accompanying drawings, those skilled in the art may make various modifications and variations without departing from the spirit and scope of the present disclosure, and such modifications and variations are all within the scope defined by the appended claims.

Claims

1. A data anomaly detection method, characterized in that: The method comprises: Acquire a normal structured sample data set, where the normal structured sample data set includes a plurality of normal structured sample data; Performing feature mapping on each of the normal structured sample data to obtain sample mapping data corresponding to each of the normal structured sample data; According to the sample distribution characteristics of each of the sample mapping data, the generation process of the abnormal structured sample is guided to obtain the abnormal structured sample data corresponding to the normal structured sample data; Performing classifier training using the normal structured sample data and the abnormal structured sample data to obtain a target classifier; The target classifier is used to perform anomaly detection on the target business data to generate anomaly detection results.

2. The method according to claim 1, characterized in that The performing feature mapping on each of the normal structured sample data to obtain sample mapping data corresponding to each of the normal structured sample data includes: Obtaining characteristic attributes corresponding to the normal structured sample data; By utilizing a preset feature mapping method, feature transformation is performed on the normal structured sample data based on the feature attributes, and the normal structured sample data is mapped to different linear spaces to obtain the sample mapping data generated in each of the linear spaces.

3. The method according to claim 1, characterized in that The step of guiding the generation process of abnormal structured samples according to the sample distribution characteristics of each sample mapping data to obtain abnormal structured sample data corresponding to the normal structured sample data includes: Adding disturbance data to the normal structured sample data according to the sample distribution characteristics; Using the disturbance data to guide the generation process of the abnormal structured sample, and constraining the generation process of the abnormal structured sample according to a preset constraint method to obtain the abnormal structured sample data; The sample distribution of the abnormal structured sample data is similar to the sample distribution of the normal structured sample data, and the abnormal structured sample data is different from the normal structured sample data.

4. The method according to claim 3, characterized in that The preset constraint method includes a first constraint and a second constraint, and constraining the generation process of the abnormal structured sample according to the preset constraint method to obtain the abnormal structured sample data includes: Performing type constraints on the normal structured sample data and the disturbance data using the first constraint to generate first sample data, where the first sample data is of the same type as the normal structured sample data; Performing feature distribution constraints on the first sample data and the perturbation data using the second constraint to generate second sample data, where the second sample data has a sample distribution feature similar to that of the normal structured sample data and is different from the normal structured sample data; The second sample data is determined as the abnormal structured sample data.

5. The method according to claim 3 or 4, characterized in that The sample distribution characteristics include sample variation characteristics and sample concentration characteristics; The abnormal structured sample data and the normal structured sample data have similar sample change characteristics, but dissimilar sample set characteristics; or, The abnormal structured sample data and the normal structured sample data have different sample change characteristics, but similar sample set characteristics; or, The abnormal structured sample data and the normal structured sample data are similar in sample change feature part, and are similar in sample concentration feature part.

6. The method according to claim 1, characterized in that The method of performing classifier training using the normal structured sample data and the abnormal structured sample data to obtain a target classifier includes: Performing iterative classification training using the normal structured sample data and the abnormal structured sample data to obtain an iterative classifier; According to the sample detection results generated by the iterative classifier and the generation process of the abnormal structured sample data, the generation process of the abnormal structured sample data and the classifier training process are interactively trained to obtain the target classifier.

7. The method according to claim 6, characterized in that The interactive training of the abnormal structured sample data generation process and the classifier training process according to the sample detection results generated by the iterative classifier and the abnormal structured sample data generation process includes: Freeze the feature mapping parameters and the perturbation parameters generated by abnormal structured samples in the current iteration; wherein the feature mapping parameters are used to map the normal structured sample data to the target linear space; and the perturbation parameters generated by the abnormal structured samples are used to control the process of generating corresponding abnormal structured sample data according to the normal structured sample data; Updating the classification parameters of the iterative classifier using the cross-training constraint parameters to obtain an iterative classifier with updated classification parameters; Freeze the classification parameters of the iterative classifier, and use the abnormal structured sample generation constraint parameters and the cross-training constraint parameters to update the feature mapping parameters and the disturbance parameters generated by the abnormal structured sample.

8. The method according to claim 1, characterized in that Also includes: If the target classifier identifies that abnormal data exists in the target business data, an abnormality assessment result and alarm information for the abnormal data are generated.

9. A data anomaly detection device, characterized in that: The device comprises: An acquisition module, configured to acquire a normal structured sample data set, wherein the normal structured sample data set includes a plurality of normal structured sample data; A feature mapping module, configured to perform feature mapping on each of the normal structured sample data to obtain sample mapping data corresponding to each of the normal structured sample data; an abnormality generation module, configured to guide the generation process of abnormal structured samples according to the sample distribution characteristics of each sample mapping data, and obtain abnormal structured sample data corresponding to the normal structured sample data; A classification training module, configured to perform classifier training using the normal structured sample data and the abnormal structured sample data to obtain a target classifier; The anomaly detection module is used to use the target classifier to perform anomaly detection on the target business data and generate anomaly detection results.

10. A computer device, characterized in that: include: A memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the data anomaly detection method according to any one of claims 1 to 8 by executing the computer instructions.

11. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a computer to execute the data anomaly detection method according to any one of claims 1 to 8.

12. A computer program product, characterized in that The method comprises computer instructions for causing a computer to execute the data anomaly detection method according to any one of claims 1 to 8.