Abnormal data detection method and device

The unsupervised model and large model collaboratively generate synthetic exception samples, and use the nuclearized fuzzy rough set theory to select exception support samples, train the discriminator, which solves the problem of lack of abnormal sample information in the existing technology and achieves more accurate abnormal data detection.

CN120145273APending Publication Date: 2025-06-13BEIJING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510306606.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-14
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The existing anomaly data detection model lacks key information about the abnormal samples during training, resulting in inaccurate detection results.

Method used

Generate pseudo-exception samples based on pre-trained unsupervised models and generate synthetic exception samples using large models. Then, anomaly support samples are selected based on the nuclearized fuzzy rough set theory, which is used to supervisedly train the discriminator to build more accurate decision boundaries.

Benefits of technology

It improves abnormal detection performance, can detect abnormal data more accurately, and significantly improves the accuracy of detection results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120145273A_ABST
    Figure CN120145273A_ABST
Patent Text Reader

Abstract

The invention discloses an abnormal data detection method and device. The method comprises the following steps: determining a pseudo abnormal sample in a test set based on a pre-trained unsupervised model; the sample formats in the test set and the training set for training the unsupervised model are the same; inputting a predetermined prompt template into a preset large model to generate a synthesis abnormal sample; the prompt template comprises a task and target component, a data description component, an analysis component and an output component; the data description component stores a training set and data features and labels of all samples in the pseudo-abnormal samples; selecting an exception support sample from the synthesis exception samples based on a nucleation fuzzy rough set theory; training a pre-constructed discriminator based on the abnormal support sample and the training set to obtain a trained discriminator; and inputting to-be-detected target data into the trained discriminator, and outputting abnormal data in the target data. According to the invention, abnormal data can be accurately detected.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data detection, and particularly relates to a method and device for detecting abnormal data. Background Art

[0002] The detection of abnormal data aims to identify data points that significantly deviate from most instance points in a dataset, and plays an important role in fields such as medical diagnosis, network intrusion detection, financial fraud detection, and industrial fault detection. However, due to the general lack of labeled abnormal data in real scenarios, existing abnormal detection models can only learn the characteristics of normal samples during the training process and detect abnormal data by evaluating the deviation of the data from the normal distribution. This type of detection model lacks the key information of abnormal samples, resulting in inaccurate speculation of abnormalities and affecting the data detection results.

[0003] Based on this, there is an urgent need for a method and device for detecting abnormal data to solve the above problems. Summary of the Invention

[0004] The present invention provides a method and device for detecting abnormal data, which can accurately detect abnormal data. The technical solutions are as follows:

[0005] In a first aspect, an embodiment of the present invention provides a method for detecting abnormal data, the method comprising:

[0006] Determining pseudo-abnormal samples in a test set based on a pre-trained unsupervised model; the test set and the samples in the training set used to train the unsupervised model have the same sample format;

[0007] Inputting a pre-determined prompt template into a preset large model to generate synthetic abnormal samples; the prompt template includes a task and target component, a data description component, an analysis component, and an output component; the data description component stores the data features and labels of each sample in the training set and the pseudo-abnormal samples;

[0008] Selecting abnormal support samples from the synthetic abnormal samples based on the kernelized fuzzy rough set theory;

[0009] Training a pre-constructed discriminator based on the abnormal support samples and the training set to obtain a trained discriminator;

[0010] Inputting the target data to be detected into the trained discriminator, and outputting the abnormal data in the target data.

[0011] In a second aspect, an embodiment of the present invention further provides a device for detecting abnormal data, the device comprising:

[0012] A determination unit, configured to determine pseudo-abnormal samples in a test set based on a pre-trained unsupervised model; the test set and the samples in the training set used to train the unsupervised model have the same sample format;

[0013] A generation unit, configured to input a pre-determined prompt template into a preset large model to generate synthetic abnormal samples; the prompt template includes a task and target component, a data description component, an analysis component, and an output component; the data description component stores the data features and labels of each sample in the training set and the pseudo-abnormal samples;

[0014] A selection unit, configured to select abnormal support samples from the synthetic abnormal samples based on the kernelized fuzzy rough set theory;

[0015] A training unit, configured to train a pre-constructed discriminator based on the abnormal support samples and the training set to obtain a trained discriminator;

[0016] An output unit, configured to input target data to be detected into the trained discriminator and output the abnormal data in the target data.

[0017] In a third aspect, an embodiment of the present invention further provides an electronic device, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the method described in any embodiment of this specification is implemented.

[0018] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed on a computer, the computer is made to execute the method described in any embodiment of this specification.

[0019] In a fifth aspect, an embodiment of the present invention further provides a computer program product, including a computer program, and when the computer program is executed by a processor, the steps of the method described above are implemented.

[0020] An embodiment of the present invention provides a method and device for detecting abnormal data. First, an unsupervised detection method is used to generate pseudo-abnormal samples, thereby providing high-confidence abnormal samples for the large model as abnormal references. Then, a prompt template is designed to help the large model understand the task context and data characteristics, so that a large number of synthetic abnormal samples can be generated. Then, the kernelized fuzzy rough set theory is used to guide the selection of abnormal support samples to ensure that the selected samples are close to the real abnormal while maintaining diversity. Finally, the discriminator is supervised by using the selected abnormal support samples and normal samples in the training set. By learning the key differences between normal and abnormal samples, a more accurate and detailed decision boundary is constructed to improve the abnormal detection ability of the discriminator. After the discriminator is trained, the target data can be detected to accurately screen out the abnormal data in the target data. It can be seen that this application can improve the performance of abnormal detection and accurately detect abnormal data. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0022] Figure 1 is a flowchart of a method for detecting abnormal data provided by an embodiment of the present invention;

[0023] Figure 2 is a structural diagram of a device for detecting abnormal data provided by an embodiment of the present invention;

[0024] Figure 3 is a hardware architecture diagram of a computer device provided by an embodiment of the present invention;

[0025] Figure 4 is a framework diagram corresponding to the detection method of this application provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0026] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0027] The following describes the specific implementation manners of the above concepts.

[0028] Please refer to Figure 1 , a method for detecting abnormal data provided by an embodiment of the present invention, the method comprising:

[0029] Step 100, based on a pre-trained unsupervised model, determine pseudo-abnormal samples in the test set; the sample formats in the test set and the training set used to train the unsupervised model are the same;

[0030] Step 102, input a pre-determined prompt template into a preset large model to generate synthetic abnormal samples; the prompt template includes a task and target component, a data description component, an analysis component, and an output component; the data description component stores the data features and labels of each sample in the training set and the pseudo-abnormal samples;

[0031] Step 104, select abnormal support samples from the synthetic abnormal samples based on the kernelized fuzzy rough set theory;

[0032] Step 106, train a pre-constructed discriminator based on the abnormal support samples and the training set to obtain a trained discriminator;

[0033] Step 108, input the target data to be detected into the trained discriminator, and output the abnormal data in the target data.

[0034] In this embodiment, first, an unsupervised detection method is used to generate pseudo-abnormal samples, so as to provide high-confidence abnormal samples for the large model as abnormal references. Then, a prompt template is designed to help the large model understand the task context and data features, so that a large number of synthetic abnormal samples can be generated. Then, the kernelized fuzzy rough set theory is used to guide the selection of abnormal support samples, ensuring that the selected samples are close to real anomalies while maintaining diversity. Finally, the selected abnormal support samples and the normal samples in the training set are used to train the discriminator in a supervised manner. By learning the key differences between normal and abnormal samples, a more accurate and detailed decision boundary is constructed to improve the abnormal detection ability of the discriminator. After the discriminator is trained, the target data can be detected to accurately screen out the abnormal data in the target data. It can be seen that this application can improve the abnormal detection performance and accurately detect abnormal data.

[0035] The following describes Figure 1 the execution manners of the following steps.

[0036] First, for step 100, it includes:

[0037] Input the test set into the trained unsupervised model to obtain the abnormal score of each sample in the test set;

[0038] Sort the anomaly scores of each sample in descending order, and determine the first quantity of samples with higher rankings as pseudo-anomaly samples.

[0039] In this step, the samples in the test set and the training set are normal samples. The specific form of the unsupervised model is not limited. After training with the training set, the optimal parameters of the unsupervised model can be obtained. The unsupervised anomaly detection model can effectively identify samples that deviate significantly from the normal distribution. Therefore, test samples with high anomaly scores can be regarded as reliable approximations of true anomalies.

[0040] The set of anomaly scores of each sample in the test set can be expressed by the following formula:

[0041]

[0042] In the formula, represents the test set; θ * is the optimal parameter obtained by training the unsupervised model Q on the training set .

[0043] In addition, the first quantity is determined according to the user's requirements. For example, the first quantity accounts for 5% of the total number of samples in the test set. Let the first quantity be K, that is, the first K samples after sorting are used as pseudo-anomaly samples. The set of pseudo-anomaly samples can be expressed by the following formula:

[0044]

[0045] In the formula, x i represents the i-th test sample, and a i represents the anomaly score of the i-th test sample. The set of pseudo-anomaly samples can be used as a reference anomaly.

[0046] For step 102, input the pre-determined prompt template into the pre-set large model to generate synthetic anomaly samples, including:

[0047] Step A1, use the task and target component to help the large model understand the task background and target;

[0048] Step A2, use the data description component to help the large model understand the data characteristics and data labels of each sample in the training set and the pseudo-anomaly samples;

[0049] Step A3, based on the understanding results, use the analysis component to guide the large model to analyze each sample, and generate initial synthetic anomaly samples based on the analysis results;

[0050] Step A4, use the output component to format the initial synthetic anomaly samples to obtain synthetic anomaly samples that meet the requirements.

[0051] In this step, the test set and the training set are tabular data, and this prompt template is formulated for tabular data. The large model takes this prompt template as input and outputs synthetic abnormal samples.

[0052] In some embodiments, for step A3, it includes:

[0053] Based on the understanding result, guide the large model to analyze the feature distribution of each dimension of each sample, and assign an abnormal score to each dimension feature based on the feature distribution;

[0054] Summarize the abnormal scores of each dimension in each sample to obtain the comprehensive abnormal score of each sample;

[0055] Determine the score range composed of the comprehensive abnormal scores of each pseudo-abnormal sample as the abnormal score range;

[0056] Guide the large model to generate synthetic abnormal samples whose abnormal scores fall within the abnormal score range.

[0057] In this step, for the features of each dimension of each sample, their abnormal scores are determined based on their positions in the overall distribution. For example, samples located at the tail of the feature distribution are assigned higher abnormal scores. In addition, the comprehensive abnormal score of each sample can be the direct sum or weighted sum of the abnormal scores of each dimension, which is not specifically limited in this application. Additionally, since both the training set and the pseudo-abnormal samples are labeled, the highest score and the lowest score corresponding to the pseudo-abnormal samples can be known, and the score range between the lowest score and the highest score is used as the abnormal score range (such as 75% - 95%). Finally, based on the learning and understanding of the features of normal samples and abnormal samples, the large model is guided to generate synthetic abnormal samples whose abnormal scores fall within the abnormal score range.

[0058] It should be noted that in step 102, the synthetic data sets directly generated by the large model often show distribution biases, which limits their applicability to downstream tasks. To solve this problem, the support sample selector composed of step 104 can be used to control the quality of the generated abnormal samples and ensure that the synthesized samples are similar to real abnormalities.

[0059] In some embodiments, for step 104, selecting abnormal support samples from synthetic abnormal samples based on the kernelized fuzzy rough set theory includes:

[0060] Step B1, merge the pseudo-abnormal samples, the training set, and the synthetic abnormal samples to obtain a comprehensive sample set;

[0061] Step B2, calculate the kernelized fuzzy relationships between the samples in the comprehensive sample set;

[0062] Step B3: Calculate the fuzzy information granules of each sample based on the kernelized fuzzy rough set theory and the similarity between samples.

[0063] Step B4: Calculate the initial approximation accuracy of each fuzzy information granule.

[0064] Step B5: Improve the initial approximation accuracy based on the proportion of the fuzzy information granules and the initial approximation accuracy of each sample in the comprehensive sample set to obtain the improved final approximation accuracy.

[0065] Step B6: Use the synthetic abnormal samples with the final approximation accuracy greater than the set threshold as abnormal support samples.

[0066] For step B2, the following formula is used to calculate the kernelized fuzzy relationship between samples in the comprehensive sample set:

[0067]

[0068]

[0069] Where

[0070]

[0071] In the formula, R represents the kernelized fuzzy relationship between samples in the comprehensive sample set; r ij represents the similarity between sample x i and x j , r ij ∈(0,1], i = 1~n, j = 1~n, n represents the total number of samples in the comprehensive sample set ; represents the eigenvalue of the m-th dimension of sample x i before normalization, m = 1~M, M represents the total dimension of each sample; represents the eigenvalue of the m-th dimension of sample x i after normalization; represents the eigenvalue of the m-th dimension of sample x j before normalization; represents the eigenvalue of the m-th dimension of sample x j after normalization; and represent the maximum and minimum values of the m-th dimension features of all samples in the comprehensive sample set respectively; represents the feature set of sample x i after normalization of all dimensions; represents the feature set of sample x j after normalization of all dimensions; represents sample x i and x jEuclidean distance on all dimensional features; δ represents the Gaussian kernel parameter.

[0072] In the process of calculating the similarity between samples, first, the feature values of each dimension of each sample are normalized to scale them to a common range, and then the data is mapped to a high-dimensional space through the Gaussian kernel function, and the relationship between samples is learned in the high-dimensional space. In addition, the greater the similarity between samples, the more similar the two samples are.

[0073] The inventor found in work that relying solely on the similarity between sample features cannot accurately capture their relationship in complex distributions. Therefore, this step uses fuzzy rough sets to handle the uncertain relationship between samples, maps them from the feature space to the score space, and finally selects the supporting samples in the score space. The implementation process is B3 to B6.

[0074] For step B3, the following formula is used to calculate the fuzzy information granule of each sample:

[0075]

[0076] In the formula, [x i R represents the fuzzy information granule of sample x i ; represents the similarity score between sample x i and sample x j .

[0077] Through the above formula, the set of fuzzy information granules corresponding to all samples can be obtained The expression is as follows:

[0078] For step B4, the following formula is used to calculate the initial approximation accuracy of each fuzzy information granule:

[0079]

[0080] In the formula, α([x i R ) represents the initial approximation accuracy of the fuzzy information granule of sample x i ; | R S [x i R | represents the lower approximation value, which represents the degree to which each sample "definitely" belongs to the same class as sample x i ; represents the upper approximation value, which represents the degree to which each sample "possibly" belongs to the same class as sample x i .

[0081] In this step, based on the initial approximation accuracy α([x​​​i R ) can successfully map the samples from the feature space to the score space. However, α([x i R ) still cannot be directly used to select samples because it is vulnerable to the number of samples in the equivalence class, making it difficult to distinguish normal samples and abnormal samples with similar initial approximation accuracy values. For example, for an abnormal sample x 1 , it is found through the lower approximation value calculation that there is 1 sample that must belong to the same category as it, and through the upper approximation value calculation, it is found that there are 2 samples that may belong to the same category as it. At this time, the initial approximation accuracy value of x 1 is calculated to be 0.5. For a normal sample x 2 , it is found through the lower approximation value calculation that there are 50 samples that must belong to the same category as it, and through the upper approximation value calculation, it is found that there are 100 samples that may belong to the same category as it. At this time, the initial approximation accuracy value of this sample is also calculated to be 0.5. In this case, α([x i R ) cannot show the difference between these two samples. Therefore, it is necessary to improve the initial accuracy, such as in step B5.

[0082] For step B5, the following is used to calculate the final approximation accuracy of each sample after improvement:

[0083]

[0084] Among them, |[x i R | ∈ [1, n].

[0085] In the formula, α′([x i R ) represents the final approximation accuracy of the fuzzy information granule of sample x i ; |[x i R | is an intermediate parameter; n is the total number of samples in the comprehensive sample set.

[0086] For step B6, the following formula is used to determine the abnormal support samples:

[0087]

[0088] In the formula, represents the set of abnormal support samples; p represents the set threshold; represents the synthesized abnormal sample.

[0089] In some embodiments, the set threshold is determined in the following manner:

[0090] ​​​​​​Based on the labels of each sample, determine the first precision range based on the final approximate precision of each normal sample, and determine the second precision range based on the final approximate precision of each abnormal sample;

[0091] Determine a set threshold according to the first precision range and the second precision range. The set threshold is between the first precision range and the second precision range, and the difference between the set threshold and the maximum approximate precision in the first precision range is not greater than a set value. This set value is set according to user requirements, such as 3%. For example, assume the first precision range is 0 - 50%, the second precision range is 60% - 95%, and the set value is 3%. Then the set threshold can take any value between 50% - 53%. In this way, the selected abnormal support samples can include both synthetic abnormal samples that are significantly deviated from the normal distribution, making the samples have a high abnormal probability; and synthetic abnormal samples that are closer to the normal distribution, improving the diversity of sample selection. Using the abnormal support samples selected by this method for discriminator training can improve the training accuracy of the model.

[0092] Finally, for steps 106 and 108:

[0093] Since the abnormal support samples selected in step 104 have diversity, the discriminator trained by them has a high precision. Using the trained discriminator for data detection can improve the accuracy of the detection results.

[0094] It should be noted that, as Figure 4 shown, the unsupervised model and the large model can form an abnormal sample generator for generating synthetic abnormal samples, and the solution involved in step 104 can form a support sample selector. The abnormal sample generator, the support sample selector, and the discriminator constitute the model architecture of the method of this application.

[0095] To prove the effectiveness of the method of this application, the inventor compared the detection effects of the method of this application with those of five baseline models (OCSVM, DeepSVDD, NeuTral AD, ICL, and MCM). The F1 scores of each detection method on different datasets (Wbc, Pima, Cardio, Ionosphere, Thyroid, Satellite, Yeast, Stamps, Cardiotocography, Fault) are shown in Table 1:

[0096] Table 1 F1 scores of each detection method on different datasets

[0097]

[0098] As can be seen from Table 1, compared with the five baseline models, the method of the present application achieves state-of-the-art performance in anomaly detection. Specifically, in terms of the F1 score, the method of the present application obtains the highest F1 score on 8 datasets and the second highest F1 score on 1 dataset, significantly outperforming the overall performance of other baseline models. The previous state-of-the-art detection model MCM only obtains 1 optimal result and 2 sub-optimal results in terms of the F1 score.

[0099] In addition, the method proposed by the present invention is a good unsupervised anomaly detection framework, which can not only effectively improve the performance of existing detection methods, but also show good robustness to different detection methods, as shown in Table 2.

[0100] Table 2 Performance comparison of different existing methods before and after being incorporated into the framework proposed by the present invention

[0101]

[0102] As can be seen from the above embodiments, the method of the present application has the following effects:

[0103] 1. The present application proposes a new unsupervised anomaly detection framework, which optimizes the decision boundary of the model by generating anomaly support samples that are very similar to real anomalies, thereby improving the anomaly detection performance of the model.

[0104] 2. In the absence of real anomalies, the present invention generates synthetic anomaly samples through the cooperation of an unsupervised model and a large model, and uses the kernelized fuzzy rough set theory to select anomaly support samples to ensure the effectiveness and diversity of the anomaly support samples.

[0105] 3. A large number of experimental results show that the present invention achieves state-of-the-art performance in unsupervised anomaly detection of tabular data, effectively improving the anomaly detection performance of the model on multiple datasets.

[0106] The method proposed by the present invention achieves state-of-the-art performance in anomaly detection, as shown in Table 1. In terms of the F1 score, the method proposed by the present invention obtains the highest F1 score on 8 datasets and the second highest F1 score on 1 dataset, significantly outperforming the overall performance of other baseline models. The previous state-of-the-art detection model MCM only obtains 1 optimal result and 2 sub-optimal results in terms of the F1 score.

[0107] As Figure 2 、 Figure 3 shown, the embodiments of the present invention provide a detection device for abnormal data. The device embodiments can be implemented by software, or by hardware or a combination of software and hardware. From the hardware level, as Figure 2As shown in the figure, it is a hardware architecture diagram of a computing device where a detection device for abnormal data provided by an embodiment of the present invention is located. In addition to Figure 2 the shown processor, memory, network interface, and non-volatile memory, the computing device where the device is located in the embodiment usually may further include other hardware, such as a forwarding chip responsible for processing packets, etc. Taking software implementation as an example, as Figure 3 shown, as a device in a logical sense, it is formed by the CPU of its computing device reading the corresponding computer program in the non-volatile memory into the memory for running.

[0108] Please refer to Figure 3 , an embodiment of the present invention provides a detection device for abnormal data, and the device includes:

[0109] A determination unit 300, configured to determine pseudo-abnormal samples in a test set based on a pre-trained unsupervised model; the sample formats in the test set and the training set used to train the unsupervised model are the same;

[0110] A generation unit 302, configured to input a pre-determined prompt template into a preset large model to generate synthetic abnormal samples; the prompt template includes a task and target component, a data description component, an analysis component, and an output component; the data description component stores the data features and labels of each sample in the training set and the pseudo-abnormal samples;

[0111] A selection unit 304, configured to select abnormal support samples from the synthetic abnormal samples based on the kernelized fuzzy rough set theory;

[0112] A training unit 306, configured to train a pre-constructed discriminator based on the abnormal support samples and the training set to obtain a trained discriminator;

[0113] An output unit 308, configured to input target data to be detected into the trained discriminator and output the abnormal data in the target data.

[0114] In some embodiments, the determination unit 300 is configured to perform the following operations:

[0115] Input the test set into the trained unsupervised model to obtain the abnormal score of each sample in the test set;

[0116] Sort the abnormal scores of each sample in descending order, and determine the first number of samples with the highest ranking as pseudo-abnormal samples.

[0117] In some embodiments, the generation unit 302 is configured to perform the following operations:

[0118] Use the task and target component to help the large model understand the task background and target;

[0119] Use the data description component to help the large model understand the data features and data labels of each sample in the training set and pseudo-anomaly samples;

[0120] Based on the understanding result, use the analysis component to guide the large model to analyze each sample, and generate initial synthetic anomaly samples based on the analysis result;

[0121] Use the output component to format the initial synthetic anomaly samples to obtain synthetic anomaly samples that meet the requirements.

[0122] In some embodiments, when the generation unit 302 executes the operation of guiding the large model to analyze each sample based on the understanding result, using the analysis component, and generating initial synthetic anomaly samples based on the analysis result, it is used to perform the following operations:

[0123] Based on the understanding result, guide the large model to analyze the feature distribution of each dimension of each sample, and assign anomaly scores to each dimension of features based on the feature distribution;

[0124] Summarize the anomaly scores of each dimension in each sample to obtain the comprehensive anomaly score of each sample;

[0125] Determine the score range composed of the comprehensive anomaly scores of each pseudo-anomaly sample as the anomaly score range;

[0126] Guide the large model to generate synthetic anomaly samples whose anomaly scores fall within the anomaly score range.

[0127] In some embodiments, the selection unit 304 is used to perform the following operations:

[0128] Merge the pseudo-anomaly samples, the training set, and the synthetic anomaly samples to obtain a comprehensive sample set;

[0129] Calculate the kernelized fuzzy relationship between each sample in the comprehensive sample set;

[0130] Based on the kernelized fuzzy rough set theory and the similarity between each sample, calculate the fuzzy information granule of each sample;

[0131] Calculate the initial approximation accuracy of each fuzzy information granule;

[0132] Improve the initial approximation accuracy based on the proportion of the fuzzy information granule and the initial approximation accuracy of each sample in the comprehensive sample set to obtain the improved final approximation accuracy;

[0133] Use the synthetic anomaly samples with the final approximation accuracy greater than the set threshold as anomaly support samples.

[0134] In some embodiments, the following formula is used to calculate the kernelized fuzzy relationship between each sample in the comprehensive sample set:

[0135]

[0136] Among them,

[0137]

[0138] In the formula, r ij represents the similarity between samples x i and x j , where i = 1 to n, j = 1 to n, and n represents the total number of samples in the comprehensive sample set ; represents the eigenvalue of the m-th dimension of sample x i before normalization, where m = 1 to M, and M represents the total dimension of each sample; represents the eigenvalue of the m-th dimension of sample x i after normalization; represents the eigenvalue of the m-th dimension of sample x j before normalization; represents the eigenvalue of the m-th dimension of sample x j after normalization; and respectively represent the maximum and minimum values of the m-th dimension features of all samples in the comprehensive sample set; represents the feature set of sample x i after normalization in all dimensions; represents the feature set of sample x j after normalization in all dimensions; represents the Euclidean distance between sample x i and x j in all-dimensional features; δ represents the Gaussian kernel parameter.

[0139] In some embodiments, the fuzzy information granules of each sample are calculated using the following formula:

[0140]

[0141] In the formula, [x i R represents the fuzzy information granule of sample x i ; represents the similarity score between sample x i and sample x j .

[0142] ​It should be noted that: The detection device for abnormal data provided in the above embodiments is only illustrated by dividing into the above functional modules. In practical applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the detection device for abnormal data provided in the above embodiments and the embodiments of the detection method for abnormal data belong to the same concept. For the specific implementation process, please refer to the method embodiments and will not be elaborated here.

[0143] An embodiment of the present application further provides a computer device. Please refer to Figure 3 , which includes a processor and a memory. At least one instruction, at least one program, a code set or an instruction set is stored in the memory. The at least one instruction, at least one program, the code set or the instruction set is loaded and executed by the processor to implement the detection method for abnormal data provided in each of the above method embodiments.

[0144] An embodiment of the present application further provides a computer-readable storage medium. At least one instruction, at least one program, a code set or an instruction set is stored on the computer-readable storage medium. The at least one instruction, at least one program, the code set or the instruction set is loaded and executed by the processor to implement the detection method for abnormal data provided in each of the above method embodiments.

[0145] An embodiment of the present application further provides a computer program product. The computer program product includes a computer program. The processor of the computer device reads the computer program from the computer-readable storage medium, and the processor executes the computer program, so that the computer device executes the detection method for abnormal data in any one of the above embodiments.

[0146] For the convenience of description, when describing the above system or device, various modules or units are described separately according to functions. Of course, when implementing the present application, the functions of each unit can be implemented in the same or multiple software and / or hardware.

[0147] From the description of the above embodiments, those skilled in the art can clearly understand that the present application can be implemented by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods in each embodiment or some parts of the embodiments of the present application.

[0148] Finally, it should also be noted that in this text, relational terms such as first, second, third, and fourth are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprising", "including", or any other variant thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article, or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article, or device comprising the element.

[0149] The above are only the preferred embodiments of the present application. It should be pointed out that for those of ordinary skill in the art, without departing from the principle of the present application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present application.

Claims

1. A method for detecting abnormal data, characterized in that: The method comprises: Based on a pre-trained unsupervised model, determining pseudo-abnormal samples in a test set; the test set has the same sample format as the sample format in the training set used to train the unsupervised model; Inputting a predetermined prompt template into a preset large model to generate a synthetic abnormal sample; the prompt template includes a task and target component, a data description component, an analysis component and an output component; the data description component stores data features and labels of each sample in the training set and the pseudo abnormal sample; Selecting abnormal support samples from the synthetic abnormal samples based on kernelized fuzzy rough set theory; Training a pre-built discriminator based on the abnormal support sample and the training set to obtain a trained discriminator; The target data to be detected is input into the trained discriminator, and the abnormal data in the target data is output.

2. The method according to claim 1, characterized in that: The method of determining pseudo abnormal samples in the test set based on the pre-trained unsupervised model includes: Inputting the test set into the trained unsupervised model to obtain an anomaly score for each sample in the test set; The anomaly score of each sample is sorted in descending order, and the first number of samples ranked at the top are determined as pseudo anomaly samples.

3. The method according to claim 1, characterized in that The step of inputting a predetermined prompt template into a preset large model to generate a synthetic abnormal sample includes: Using the task and goal components to help the large model understand the task context and goals; Using the data description component to help the large model understand the data features and data labels of each sample in the training set and the pseudo-abnormal samples; Based on the understanding results, the analysis component is used to guide the large model to analyze each sample, and an initial synthetic abnormal sample is generated based on the analysis results; The output component is used to format the initial synthetic abnormal sample to obtain a synthetic abnormal sample that meets the requirements.

4. The method according to claim 3, characterized in that: Based on the understanding result, the analysis component is used to guide the large model to analyze each sample, and an initial synthetic abnormal sample is generated based on the analysis result, including: Based on the understanding results, the large model is guided to analyze the feature distribution of each dimension of each sample, and an abnormal score is assigned to each dimension feature based on the feature distribution; Summarize the anomaly scores of each dimension in each sample to obtain the comprehensive anomaly score of each sample; The score range composed of the comprehensive abnormal scores of each pseudo-abnormal sample is determined as the abnormal score range; The large model is guided to generate synthetic anomaly samples whose anomaly scores fall within the anomaly score range.

5. The method according to claim 1, characterized in that Selecting abnormal support samples from the synthetic abnormal samples based on kernelized fuzzy rough set theory includes: Merging the pseudo abnormal sample, the training set and the synthetic abnormal sample to obtain a comprehensive sample set; Calculating the kernelized fuzzy relationship between samples in the comprehensive sample set; Based on kernelized fuzzy rough set theory and the similarity between samples, the fuzzy information granules of each sample are calculated; Calculate the initial approximation accuracy of each fuzzy information granule; The initial approximate accuracy of each sample is improved based on the proportion of the fuzzy information granules and the initial approximate accuracy in the comprehensive sample set to obtain an improved final approximate accuracy; The synthetic anomaly samples whose final approximation accuracy is greater than the set threshold are taken as anomaly support samples.

6. The method according to claim 5, characterized in that The kernelized fuzzy relationship between samples in the comprehensive sample set is calculated using the following formula: in, In the formula, r ij Represents sample x i and x j Similarity between them, i = 1 ~ n, j = 1 ~ n, n represents the comprehensive sample set The total number of samples in ; Represents the sample x before normalization i The eigenvalue of the mth dimension, m = 1 to M, where M represents the total dimension of each sample; Represents the normalized sample x i The eigenvalue of the mth dimension; Represents the sample x before normalization j The eigenvalue of the mth dimension; Represents the normalized sample x j The eigenvalue of the mth dimension; They represent the maximum and minimum values ​​of the m-th dimension features of all samples in the comprehensive sample set respectively; Represents sample x i The feature set after all dimensions are normalized; Represents sample x j The feature set after all dimensions are normalized; Represents sample x i and x j Euclidean distance on all dimensional features; δ represents the Gaussian kernel parameter.

7. The method according to claim 6, characterized in that The fuzzy information granule of each sample is calculated using the following formula: In the formula, [x i ] R Represents sample x i The fuzzy information particles; Represents sample x i With sample x j The similarity score between .

8. A device for detecting abnormal data, characterized in that: The device comprises: A determination unit, used to determine pseudo-abnormal samples in a test set based on a pre-trained unsupervised model; the test set has the same sample format as the sample format in the training set used to train the unsupervised model; A generating unit, used for inputting a predetermined prompt template into a preset large model to generate a synthetic abnormal sample; the prompt template includes a task and target component, a data description component, an analysis component and an output component; the data description component stores data features and labels of each sample in the training set and the pseudo abnormal sample; A selection unit, configured to select an abnormal support sample from the synthetic abnormal samples based on kernelized fuzzy rough set theory; A training unit, used for training a pre-built discriminator based on the abnormal support sample and the training set to obtain a trained discriminator; The output unit is used to input the target data to be detected into the trained discriminator and output the abnormal data in the target data.

9. A computing device, comprising a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the method according to any one of claims 1 to 7 is implemented.

10. A computer-readable storage medium having a computer program stored thereon, which, when executed in a computer, causes the computer to execute the method according to any one of claims 1 to 7.