Fuzzy beta-coverage relation-based entropy and anomaly detection method thereof
By constructing the entropy of fuzzy β-covering relations and its anomaly detection method based on the theory of fuzzy β-covering rough sets, the problem of information loss in rough entropy when processing covered data is solved, and efficient anomaly detection for complex data is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHAANXI UNIV OF SCI & TECH
- Filing Date
- 2026-01-19
- Publication Date
- 2026-05-05
AI Technical Summary
Existing coarse entropy methods rely on binary equivalence relations to partition the sample set, which is only applicable to nominal data and cannot effectively characterize the large amount of data that exists in the form of overlay in reality, resulting in serious information loss when processing continuous data.
We propose an entropy and anomaly detection method based on fuzzy β-coverage relationship. By transferring the traditional fuzzy β-neighborhood operator to the variable-scale fuzzy β-coverage approximation space, we construct fuzzy β-co-neighborhood and fuzzy β-coverage relationship, calculate fuzzy β-coverage entropy and fuzzy β-coverage relative entropy, and construct anomaly score to detect anomalous samples.
It can effectively process nominal, numerical and mixed data, avoid information loss caused by discretization, significantly improve the ability to express the uncertainty of complex data, and improve the accuracy of anomaly detection.
Smart Images

Figure CN121980454A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of data anomaly detection technology, specifically relating to entropy based on fuzzy β-coverage relationship and its anomaly detection method. Background Technology
[0002] The core objective of anomaly detection is to identify anomalous data that differs significantly from the majority of samples in a dataset. Therefore, measuring the degree of anomalousness of samples in the data is crucial. Rough entropy, an uncertainty measure based on rough set theory, is used to characterize the discriminative power of a subset of attributes on samples and is thus frequently used in anomaly detection tasks.
[0003] However, classical rough sets rely on binary equivalence relations to partition the sample set, making it difficult to effectively characterize the large amount of data in reality that exists in a covering form. To address this, Zakowski proposed covering rough sets in the 1980s, a significant extension of Pawlak rough set theory. Unlike Pawlak rough sets, which strictly rely on equivalence relations, covering rough sets relax equivalence relations to a general covering structure, allowing overlap between sets without guaranteeing that each element belongs to only one set. This structural design significantly enhances the model's classification flexibility, particularly suitable for incomplete data where one or more attribute values are missing. However, both of these models can only handle nominal data, i.e., discrete data, necessitating discretization when processing continuous data, resulting in significant information loss. To overcome this limitation, Dubois et al. and Deng et al., based on Zadeh's fuzzy set theory, extended Pawlak rough sets and covering rough sets into fuzzy rough sets and fuzzy covering rough sets, respectively. By combining the advantages of rough set theory and fuzzy set theory, these models can minimize information loss in data processing, becoming an important mathematical tool for handling uncertain and incomplete information, and have been widely applied in practical application scenarios such as feature selection, language knowledge acquisition, and decision-making.
[0004] Given that the definition of fuzzy covering is too strict, Ma creatively replaced the "1" in the definition with a parameter β (0 < β < 1), thus generalizing fuzzy covering and obtaining the definition of fuzzy β-cover. In fact, when β = 1, fuzzy β-cover is a fuzzy cover. Therefore, this is not only a theoretical extension of fuzzy covering, but also improves the ability of fuzzy covering rough sets to solve practical problems.
[0005] In summary, fuzzy covering rough sets are theoretically a fuzzification and generalization of covering rough sets. Fuzzy coverings are a family of fuzzy subsets that satisfy specific conditions by replacing classical sets with fuzzy sets, cleverly overcoming the limitations of rough set theory in both theory and application. Therefore, based on the theory of fuzzy β-covering rough sets, this paper proposes a new method for constructing rough entropy based on fuzzy β-covering relations and applies it to the task of anomaly sample detection. Summary of the Invention
[0006] The purpose of this invention is to address the problem that existing rough entropy methods rely on binary equivalence relations to partition sample sets, and are only applicable to nominal data, making it difficult to effectively characterize the large amount of data existing in the form of covering in reality. A new solution is proposed. First, a method is proposed to extend the traditional fuzzy β-neighborhood operator to a variable-scale fuzzy β-covering approximation space. Based on this, a fuzzy β-covering relation based on fuzzy β-co-neighborhood is constructed, which can characterize the differences between samples from a reverse perspective, better meeting the needs of anomaly detection tasks. Furthermore, fuzzy β-covering entropy and fuzzy β-covering relative entropy are proposed, and an anomaly detection method is designed accordingly. Experiments conducted on 10 datasets from different fields show that the proposed method has significant advantages compared to seven existing anomaly detection methods, effectively detecting anomalous samples in the data. This invention not only expands the application of fuzzy β-covering rough set theory in the field of anomaly detection but also effectively compensates for the shortcomings of traditional methods in handling data existing in the form of covering, providing reliable technical support for practical applications.
[0007] The specific technical solution adopted by this invention is as follows: Entropy and anomaly detection methods based on fuzzy β-coverage relationships include the following steps: Step 1: Propose a method to transfer the traditional fuzzy β-neighborhood operator to the variable-scale fuzzy β-covering approximation space; Step 2: Propose two parameterized fuzzy β-co-neighborhoods; Step 3: Propose fuzzy β-coverage relations and fuzzy β-coverage information granules based on fuzzy β-co-neighborhood; Step 4: Propose fuzzy β-coverage entropy and fuzzy β-coverage relative entropy; Step 5: Propose a sequence-based fuzzy β-coverage relative entropy and weighting function; Step 6: Based on the content of Steps 1 to 5, construct anomaly scores and propose an anomaly detection method based on fuzzy β-coverage entropy. According to the input information table, obtain the output and output the detected anomaly samples.
[0008] The technical effects achieved by this invention are as follows: First, this invention proposes a method for constructing entropy based on fuzzy β-covering relations. Existing rough entropy methods rely on binary equivalence relations to partition sample sets and are only applicable to nominal data, failing to effectively characterize the large amount of data existing in a covering form in reality. Based on fuzzy β-covering rough set theory, this invention proposes a novel method for constructing rough entropy based on fuzzy β-covering relations. This not only overcomes the limitation of rough entropy being only applicable to nominal attribute data but also characterizes data existing in a covering form. It inherits the advantages of rough entropy while cleverly overcoming its limitations in practical applications, theoretically enhancing its ability to express the uncertainty of complex data.
[0009] Secondly, this invention proposes an anomaly detection method based on fuzzy β-coverage entropy. The fuzzy β-coverage entropy obtained by the above method, when applied to anomaly detection tasks, can effectively characterize the large amount of data existing in the form of coverage in reality. Compared to traditional coarse entropy methods, this invention can not only process nominal data but also directly process numerical data and even mixed data, effectively avoiding data loss caused by discretization. Attached Figure Description
[0010] Figure 1 This is a flowchart of the entropy and anomaly detection method based on fuzzy β-coverage relationship of the present invention. Detailed Implementation
[0011] To make the objectives and advantages of this invention clearer, the invention will be specifically described below with reference to embodiments. It should be understood that the following text is merely used to describe one or more specific embodiments of the invention and does not strictly limit the scope of protection specifically claimed by the invention.
[0012] like Figure 1 As shown, the entropy and anomaly detection method based on fuzzy β-coverage relationship includes the following steps: Step 1: Propose a method to transfer the traditional fuzzy β-neighborhood operator to the variable-scale fuzzy β-covering approximation space; Step 2: Propose two parameterized fuzzy β-co-neighborhoods; Step 3: Propose fuzzy β-coverage relations and fuzzy β-coverage information granules based on fuzzy β-co-neighborhood; Step 4: Propose fuzzy β-coverage entropy and fuzzy β-coverage relative entropy; Step 5: Propose a sequence-based fuzzy β-coverage relative entropy and weighting function; Step 6: Based on the content of Steps 1 to 5, construct anomaly scores and propose an anomaly detection method based on fuzzy β-coverage entropy. According to the input information table, obtain the output and output the detected anomaly samples.
[0013] Preferably, in step 1: Definition 1: Let... It is a non-empty domain. Describing the domain The set consisting of all fuzzy sets on; if the family of fuzzy sets... ( , Satisfies: for any and All have Then it is called For the domain Fuzzy β-coverage on For fuzzy β-covered approximate space; In addition, Describes a non-empty index set, if It is a fuzzy beta cover. in It is a blur -cover( ), then it is called For the variable-scale fuzzy β-covered approximation space; The variable-scale fuzzy β-cover approximation space is a reasonable generalization of the fuzzy β-cover approximation space; specifically, for any given variable-scale fuzzy β-cover approximation space... If for all All meet Then the variable-scale fuzzy β-covering approximate space With a fuzzy β-covered approximate space Equivalent; at this point, Blur in cover satisfy Therefore, the fuzzy β-neighborhood operator defined in the variable-scale fuzzy β-cover approximation space can be directly applied to the fuzzy β-cover approximation space; conversely, if the fuzzy β-neighborhood operator defined in the fuzzy β-cover approximation space is to be applied to the variable-scale fuzzy β-cover approximation space, the fuzzy β-neighborhood operator must be... cover View as a blur cover Based on this, the fuzzy β-neighborhood operator defined on the fuzzy β-cover approximation space and the fuzzy β-neighborhood operator defined on the variable-scale fuzzy β-cover approximation space can be mutually transformed.
[0014] Preferably, in step 2: Definition 2: Let... It is a variable-scale fuzzy β-covering approximation space. It is an overlapping function. It is a grouping function; for any ,definition Parameterized fuzzy β-co-neighborhood and for: ; ; in ;and ; .
[0015] Preferably, in step 3: Definition 3: Let... It is a variable-scale fuzzy β-covering approximation space. For any Define the l-th type of fuzzy β-covering relation for: ; Definition 4: Let... For a variable-scale fuzzy β-covered approximation space, and For any Defined by fuzzy β-coverage relation Induced generation domain The fuzzy β-covering grain structure on it is: , in It is composed of fuzzy beta-coverage relations The generated fuzzy β-coverage information granules; a fuzzy β-coverage relation can derive a granular structure, and Defined as the domain of discourse A fuzzy set on , satisfying The formula for calculating fuzzy information granularity is as follows: Obviously there is .
[0016] Preferably, in step 4: Definition 5: Let... For a variable-scale fuzzy β-covered approximation space, and definition The fuzzy β-coverage entropy is: ; For any ,definition The relative entropy of the fuzzy β-coverage is: ; in , indicating from Remove samples back The fuzzy β-coverage entropy.
[0017] Preferably, in step 5: Definition 6: Let... For a variable-scale fuzzy β-coverage approximation space, the attribute sequence is defined as follows: ,in , Define the sequence of attribute subsets as ,in , , ,and , ; Let the definition be: For a variable-scale fuzzy β-covered approximation space, and For any , and ,definition The relative entropy and weighting function for the sequence-based fuzzy β-coverage are as follows: ; in Indicates the use of the first Class-fuzzy beta-coverage relation The fuzzy β-coverage entropy, It is a sample About attributes The relative entropy of fuzzy β-coverage It is a sample Regarding attribute subsets The relative entropy of fuzzy β-coverage.
[0018] Preferably, in step 6: Let the definition be: For a variable-scale fuzzy β-covering approximation space, define The abnormal score is: ; Input: Input information form As shown in Table 1, in , , , ,
[0019] Output: Detected abnormal samples.
[0020] Preferably, step 6 specifically includes the following steps: Step 61: Inducing a variable-scale fuzzy β-coverage approximation space: In the preprocessing stage, all data are represented as distinct integers, and all attribute values are normalized to the [0,1] interval using a min-max normalization method; the resulting information table... For a variable-scale fuzzy β-covering approximation space ,in ; Step 62: Calculate all In a single attribute The relative entropy of the fuzzy β-coverage is calculated through the following steps:
[0021] Step 63: Calculate the attribute sequence AS and the attribute subset sequence ASS; Step 64: Calculate all In attribute subset The relative entropy of the fuzzy β-coverage is calculated through the following steps:
[0022] Step 65: Calculate the anomaly score to obtain the anomaly sample set. ,in The threshold for abnormal samples set by experts.
[0023] In summary: Figure 1 As shown, when using entropy based on fuzzy β-coverage relationship and its anomaly detection method, in order to analyze the application of the method in rotating machinery vibration signals, this invention uses 10 datasets from https: / / github.com / BElloney / Outlier-detection after processing for anomaly detection tasks. The following is a description of the datasets: (1) The Annealing dataset is a dataset related to the steel annealing process. It contains 38 features such as chemical composition content, physical properties, and size parameters, with a total of 798 samples. Among them, the "1" and "U" classes are merged into anomaly classes. (2) The Autos dataset contains relevant data from Ward's Automotive Yearbook in 1985, including 25 features such as various vehicle specifications and insurance risk ratings, with a total of 205 samples. Among them, the "-2" and "-1" classes are merged into anomaly classes. (3) The Bands dataset focuses on the roller stripe problem in rotary gravure printing, covering 39 features such as viscosity, ink temperature, and printing speed, with a total of 328 samples. Among them, the "band" class was downsampled to 16 samples as an anomaly class. (4) The Breast cancer dataset comes from the University Medical Center of the Cancer Institute in Ljubljana. It contains 9 features including patient age, tumor size, and malignancy, with a total of 286 samples. Among them, the “recurrence-events” class is an abnormal class. (5) The Iris dataset contains four features: sepal length, sepal width, petal length, and petal width, with a total of 111 samples. Among them, the “Iris-virginica” class was downsampled to 11 samples as an anomaly class. (6) The Wisconsin breast cancer dataset, also known as the WBC dataset, is based on clinical cases and includes nine features such as tumor thickness, cell size uniformity, and edge adhesion. There are a total of 483 samples, of which 39 are "malignant" and are abnormal. (7) The Wisconsin diagnostic breast cancer dataset, also known as the WDBC dataset, is a diagnostic dataset based on digital images extracted from fine-needle aspiration (FNA) of breast masses. It contains 31 features extracted from cells, totaling 396 samples, of which 39 samples of the "M" class are downsampled to the abnormal class. (8) The Wisconsin prognostic breast cancer dataset, also known as the WPBC dataset, focuses on follow-up data of breast cancer patients after diagnosis. It contains 30 cellular features, 2 clinical features and 1 time feature, with a total of 198 samples. The “R” class is considered as the abnormal class. (9) The Yeast dataset is mainly used to predict the cellular localization of proteins. It contains 8 features, including the McGeoch signal sequence recognition method score, with a total of 1141 samples. Among them, “ERL” is selected as the anomaly class. (10) The Zoo dataset is a classic dataset focusing on animal classification, containing 16 physiological and behavioral characteristics of animals such as mammals, flying animals, and aquatic animals, with a total of 101 samples. Among them, the "reptile", "amphibian" and "insect" classes are merged into anomaly classes. Table 1 comprehensively shows the specific information of the above-mentioned processed dataset.
[0024]
[0025] Finally, the anomaly detection method based on fuzzy β-coverage entropy proposed in this invention was used to detect anomalous samples in the above dataset. The AUC of this method was compared with that of seven other anomaly detection methods: Local Anomaly Detection (LOF) based on density, OCSVM based on support vector machines, IFOres (Isolation Forest), HBOS based on histograms, ECOD based on empirical cumulative distribution functions, KFRAD based on kernel fuzzy rough sets, and FREAD based on fuzzy rough entropy. The results are shown in Table 2. A closer AUC is to 1 indicates better overall performance in the anomaly detection task.
[0026]
[0027] Clearly, the method proposed in this invention achieves the highest average AUC value. Furthermore, it demonstrates significant advantages on the two challenging datasets, Autos and WPBC: On the Autos dataset, the proposed method achieves a 13.40% performance improvement in AUC compared to the second-best method. On the WPBC dataset, the proposed method achieves an AUC of 0.6112, becoming the only method to break the 0.6 AUC threshold, representing a 12.35% performance improvement compared to the second-best method, HBOS. In contrast, other methods generally achieve AUC values close to or even lower than random guessing (AUC=0.5) on these two datasets. This significant difference indicates that in complex data scenarios where traditional methods struggle or perform poorly, the proposed method can more effectively capture anomalies, providing a unique and valuable solution for anomaly detection tasks.
[0028] The above description is merely a preferred embodiment of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention. Structures, devices, and operating methods not specifically described or explained in this invention are implemented according to conventional methods in the art unless otherwise specified or limited.
Claims
1. An entropy and anomaly detection method based on fuzzy β-coverage relationship, characterized in that: Includes the following steps: Step 1: Propose a method to transfer the traditional fuzzy β-neighborhood operator to the variable-scale fuzzy β-covering approximation space; Step 2: Propose two parameterized fuzzy β-co-neighborhoods; Step 3: Propose fuzzy β-coverage relations and fuzzy β-coverage information granules based on fuzzy β-co-neighborhood; Step 4: Propose fuzzy β-coverage entropy and fuzzy β-coverage relative entropy; Step 5: Propose a sequence-based fuzzy β-coverage relative entropy and weighting function; Step 6: Based on the content of Steps 1 to 5, construct anomaly scores and propose an anomaly detection method based on fuzzy β-coverage entropy. According to the input information table, obtain the output and output the detected anomaly samples.
2. The entropy and anomaly detection method based on fuzzy β-coverage relationship according to claim 1, characterized in that: In step 1: Definition 1: Let... It is a non-empty domain. Describing the domain The set consisting of all fuzzy sets on; if the family of fuzzy sets... ( , Satisfies: for any and All have Then it is called For the domain Fuzzy β-coverage on For fuzzy β-covered approximate space; In addition, Describes a non-empty index set, if It is a fuzzy beta cover. in It is a blur -cover( ), then it is called For the variable-scale fuzzy β-covered approximation space; The variable-scale fuzzy β-cover approximation space is a reasonable generalization of the fuzzy β-cover approximation space; specifically, for any given variable-scale fuzzy β-cover approximation space... If for all All meet Then the variable-scale fuzzy β-covering approximate space With a fuzzy β-covered approximate space Equivalent; at this point, Blur in cover satisfy Therefore, the fuzzy β-neighborhood operator defined in the variable-scale fuzzy β-cover approximation space can be directly applied to the fuzzy β-cover approximation space; conversely, if the fuzzy β-neighborhood operator defined in the fuzzy β-cover approximation space is to be applied to the variable-scale fuzzy β-cover approximation space, the fuzzy β-neighborhood operator must be... cover View as a blur cover Based on this, the fuzzy β-neighborhood operator defined on the fuzzy β-cover approximation space and the fuzzy β-neighborhood operator defined on the variable-scale fuzzy β-cover approximation space can be mutually transformed.
3. The entropy and anomaly detection method based on fuzzy β-coverage relationship according to claim 2, characterized in that: In step 2: Definition 2: Let... It is a variable-scale fuzzy β-covering approximation space. It is an overlapping function. It is a grouping function; for any ,definition Parameterized fuzzy β-co-neighborhood and for: ; ; in ;and ; 。 4. The entropy and anomaly detection method based on fuzzy β-coverage relationship according to claim 3, characterized in that: In step 3: Definition 3: Let... It is a variable-scale fuzzy β-covering approximation space. For any Define the l-th type of fuzzy β-covering relation for: ; Definition 4: Let... For a variable-scale fuzzy β-covered approximation space, and For any Defined by fuzzy β-coverage relation Induced generation domain The fuzzy β-covering grain structure on it is: , in It is composed of fuzzy beta-coverage relations The generated fuzzy β-coverage information granules; a fuzzy β-coverage relation can derive a granular structure, and Defined as the domain of discourse A fuzzy set on , satisfying The formula for calculating fuzzy information granularity is as follows: Obviously there is .
5. The entropy and anomaly detection method based on fuzzy β-coverage relationship according to claim 4, characterized in that: In step 4: Definition 5: Let... For a variable-scale fuzzy β-covered approximation space, and definition The fuzzy β-coverage entropy is: ; For any ,definition The relative entropy of the fuzzy β-coverage is: ; in , indicating from Remove samples back The fuzzy β-coverage entropy.
6. The entropy and anomaly detection method based on fuzzy β-coverage relationship according to claim 5, characterized in that: In step 5: Definition 6: Let... For a variable-scale fuzzy β-coverage approximation space, the attribute sequence is defined as follows: ,in , Define the sequence of attribute subsets as ,in , , ,and , ; Let the definition be: For a variable-scale fuzzy β-covered approximation space, and For any , and ,definition The relative entropy and weighting function for the sequence-based fuzzy β-coverage are as follows: ; in Indicates the use of the first Class-fuzzy beta-coverage relation The fuzzy β-coverage entropy, It is a sample About attributes The relative entropy of fuzzy β-coverage It is a sample Regarding attribute subsets The relative entropy of fuzzy β-coverage.
7. The entropy and anomaly detection method based on fuzzy β-coverage relationship according to claim 6, characterized in that: In step 6: Let the definition be: For a variable-scale fuzzy β-covering approximation space, define The abnormal score is: ; Input: Input information form Output: Detected abnormal samples.
8. The entropy and anomaly detection method based on fuzzy β-coverage relationship according to claim 7, characterized in that: The implementation process of step 6 specifically includes: Step 61: Inducing a variable-scale fuzzy β-coverage approximation space: In the preprocessing stage, all data are represented as distinct integers, and all attribute values are normalized to the [0,1] interval using a min-max normalization method; the resulting information table... For a variable-scale fuzzy β-covering approximation space ,in ; Step 62: Calculate all In a single attribute The relative entropy of fuzzy β-coverage Step 63: Calculate the attribute sequence AS and the attribute subset sequence ASS; Step 64: Calculate all In attribute subset The relative entropy of fuzzy β-coverage Step 65: Calculate the anomaly score to obtain the anomaly sample set. ,in The threshold for abnormal samples set by experts.