Data hierarchical fusion method based on fuzzy rough set

CN118260714BActive Publication Date: 2026-09-29DALIAN UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410448720.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-04-15
Publication Date
2026-09-29
Estimated Expiration
2044-04-15

AI Technical Summary

Technical Problem

最后,引入基于核函数的模糊相似关系,解决了粗糙集理论在连续型属性或带扰动的属性上不兼容的问题

Benefits of technology

[0041]本发明的有益效果:数据融合能够整合多个数据集的数据以便进行数据分析和挖掘,但在带来方便的同时,也带来了一系列的问题。直接对多个数据集的数据进行融合会导致得到的融合数据可用性较低,特别是会造成当前研究较少的属性冗余问题,因此,本发明提出一种基于模糊粗糙集的数据分级融合方法。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118260714B_ABST
    Figure CN118260714B_ABST
Patent Text Reader

Abstract

A data hierarchical fusion method based on fuzzy rough set is provided.The attribute importance is graded by wavelet clustering algorithm, the attribute importance is defined by information entropy, skewness coefficient and kurtosis coefficient, and the feature space is quantified.The quantified feature space is transformed by wavelet, the density threshold is determined according to the data distribution after wavelet transformation, and the cluster label is assigned, then the lookup table is made and the original data is mapped to the corresponding cluster.The attribute redundancy is removed by fuzzy rough set, the attribute with the highest importance is selected as the attribute to be reduced, the remaining attributes are traversed, if the attribute is continuous, the fuzzy similarity of the attribute is calculated by kernel function, then the discernibility matrix and discernibility function of the attribute are calculated by rough set, otherwise the discernibility matrix and discernibility function are directly calculated;according to the importance grading result, the attribute set with the lowest importance is determined, and the new attribute to be reduced is selected, until there is no important attribute left or the attribute set is empty.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of information processing and relates to a data hierarchical fusion method based on fuzzy rough sets. Specifically, it relates to an attribute importance hierarchical method based on wavelet clustering and an attribute redundancy removal method based on fuzzy rough sets. Background Technology

[0002] Data fusion is a promising field that integrates information from multiple datasets for in-depth analysis and mining, resulting in more accurate results. In many areas such as payments, healthcare, and agriculture, data fusion has become an indispensable tool. However, while data fusion brings numerous conveniences, it also faces challenges. Due to the uncertainty of data sources and quality, directly fusing these data often leads to poor quality of the fused data, and even attribute redundancy. For example, one dataset might record users' home addresses, while another dataset records their postal codes. Since the postal codes can be obtained from the home addresses, this is redundant. Attribute redundancy is a pressing issue in data fusion. Excessive redundant data not only wastes computational resources but also causes distortion of analysis results, decreased model performance, and increased complexity in subsequent analysis and mining work.

[0003] Currently, one of the important methods for solving attribute redundancy is attribute reduction, which is mainly divided into methods based on information entropy, evidence theory, and rough sets. The first two methods are aimed at the relationship between two attributes, using methods to find the degree of correlation between two attribute variables, thereby hiding attributes with similarity higher than a certain threshold. However, in real life, during data fusion, there are often situations where multiple attributes jointly contain information about a specific attribute. For example, a personal ID number contains information such as date of birth, gender, and province. In such cases, the first two methods cannot solve the attribute redundancy problem.

[0004] Existing rough set-based attribute reduction methods can handle redundancy issues with multiple attributes, but currently, these methods are mainly applied to decision-making systems. The aim of this method is to reduce a specific decision attribute to obtain a minimal set of conditional attributes that can classify that attribute for subsequent decision-making. However, real-world data mining tasks require high-quality fused datasets to support various analyses and decisions, not just those specific to a single decision. Furthermore, current rough set theory is not applicable to continuous data, while most real-world datasets contain continuous data. Additionally, in data fusion, different attributes contain varying amounts of information, and data distribution also influences attribute importance. Therefore, after obtaining the attribute reduction results, how to select attributes—discarding those with less information and retaining those with greater value—is also a problem that needs to be considered. Summary of the Invention

[0005] To effectively address the attribute redundancy problem in data fusion and retain the most valuable attributes, this invention proposes a hierarchical data fusion method based on fuzzy rough sets. First, the scheme proposes an importance-based hierarchical strategy using wavelet clustering, considering the influence of data skewness and kurtosis on attribute sensitivity, and clusters the fused data according to attribute importance. Then, it proposes using a rough set-based attribute reduction method to solve the attribute redundancy problem where a single attribute is implied in multiple attributes during data fusion. Finally, it introduces a kernel-based fuzzy similarity relation to address the incompatibility of rough set theory with continuous or perturbed attributes.

[0006] The technical solution of this invention:

[0007] A data hierarchical fusion method based on fuzzy rough sets, comprising the following steps:

[0008] Define variables:

[0009] Table 1 Commonly Used Variables and Explanations

[0010]

[0011] Continued from Table 1: Commonly Used Variables and Explanations

[0012]

[0013] The specific steps are as follows:

[0014] (1) Use wavelet clustering algorithm to cluster the importance of attributes, use attribute sensitivity, skewness coefficient and kurtosis coefficient as features of attribute importance, and select the attribute with the highest importance as the attribute to be reduced;

[0015] Clustering attributes based on their importance follows the specific process:

[0016] (1.1) First, attribute importance features are obtained using attribute sensitivity, skewness coefficient, and kurtosis coefficient, and then the feature space is quantified using the coefficient of variation. Grid division;

[0017] Attribute sensitivity (AS): Sensitivity to attribute A in dataset D i ( Given a dataset D containing a set of all attributes, the sensitivity of an attribute is defined as the ratio of the difference between the attribute's maximum discrete entropy and its information entropy to the attribute's maximum discrete entropy. The formula is as follows:

[0018]

[0019] Among them, H(A) i ) is attribute x i Information entropy, H max (A i ) is attribute A i The maximum discrete entropy.

[0020] AS i ∈(0,1), attribute sensitivity AS i The smaller the value, the more sensitive the attribute; conversely, the larger the value, the less sensitive the attribute.

[0021] Skewness coefficient (SK): Used to measure a certain attribute A in dataset D. i The degree of skewness, when calculating the skewness coefficient for ungrouped original attributes, is defined by the following formula:

[0022]

[0023] Where n represents the number of data entries in dataset D, x j Let attribute A represent the j-th record in dataset D. i The value of s represents attribute A i The standard deviation of all values, Represents attribute A i The average of all possible values.

[0024] |SK|=0 indicates that the data is symmetrically distributed, |SK|>0 indicates that the data is right-skewed, and |SK|<0 indicates that the data is left-skewed.

[0025] Kuroism coefficient (K): Used to measure a certain attribute A in dataset D. i The peak level is defined by the following formula:

[0026]

[0027] Where n represents the number of data entries in dataset D, x j Let attribute A represent the j-th record in dataset D. i The value of s represents attribute A i The standard deviation of all values, Represents attribute A i The average of all possible values.

[0028] K=0 indicates that the data follows a normal distribution; K>0 indicates a leptokurtic distribution, where the data is more concentrated; K<0 indicates a flat distribution, where the data is more dispersed.

[0029] Coefficient of variation (c) v ): Used to measure a specific attribute A in dataset D. i The degree of dispersion of the probability distribution is defined by the following formula:

[0030]

[0031] Where s represents attribute A i The standard deviation of all values, Represents attribute A i The average of all possible values.

[0032] (1.2) After obtaining the above calculation results, attribute A i {AS} i Using ,|SK|,K} as features, for the attribute set Clustering is performed. First, the feature space is... Perform wavelet transform to obtain the transformed feature space. Then, based on the feature space after wavelet transform... The distribution of data determines the threshold, and grids with a density greater than the threshold are marked as dense. Then, dense and connected grids are grouped into a cluster and numbered. Finally, the data in the grid is labeled with the cluster number it belongs to.

[0033] (1.3) Establish a mapping table to map cluster labels to the original feature space. Original feature space The data in the dataset is mapped to their respective clusters according to the cluster labels, and the importance level of each cluster is determined based on the attribute sensitivity of the centroid of each cluster. Then, the attribute in the cluster with the highest importance is selected as the attribute to be reduced.

[0034] (2) After selecting the attributes to be reduced, calculate the fuzzy similarity for the remaining continuous attributes. Discrete attributes do not need to be processed. Then, use the attribute reduction algorithm based on rough set to calculate the minimum decision set of the attributes and select the attribute decision set with the lowest sensitivity as the attribute reduction set.

[0035] The attribute redundancy removal is performed using an attribute reduction algorithm based on fuzzy rough sets. The specific process is as follows:

[0036] (2.1) First, select the most sensitive attribute from the cluster with the highest importance as the attribute to be reduced. Then, iterate through the remaining attributes. If the attribute is a continuous attribute, calculate the fuzzy similarity between any two data objects under the attribute. If the attribute is a discrete attribute, no processing is required. Repeat this process until the fuzzy similarity has been calculated for all continuous attributes.

[0037] Fuzzy similarity relation for continuous data objects: used to determine whether two arbitrary continuous data objects are similar. The formula is defined as follows:

[0038]

[0039] Among them, A c It is an arbitrary continuous attribute, where x and y are two arbitrary data objects. This indicates that data objects x and y are in property A c The conditions are similar. It is a Gaussian kernel function, and ε is a threshold, ε∈[0,1].

[0040] (2.2) Calculate the discrimination matrix and discrimination function of the attribute to be reduced, and select the attribute decision set with the lowest sensitivity based on the clustering results of attribute importance.

[0041] The beneficial effects of this invention are as follows: Data fusion can integrate data from multiple datasets for data analysis and mining, but while bringing convenience, it also brings a series of problems. Directly fusing data from multiple datasets leads to low usability of the fused data, especially causing attribute redundancy, a problem that is currently less studied. Therefore, this invention proposes a hierarchical data fusion method based on fuzzy rough sets.

[0042] When classifying the importance of attributes, considering that the amount of information contained in each attribute is different, resulting in different levels of importance for each attribute, and that the distribution of attributes also affects the importance of attributes, using attribute sensitivity, skewness coefficient, and kurtosis coefficient as features of attribute importance can more accurately measure the importance of attributes.

[0043] By using kernel functions to improve the fuzzy similarity relationship of continuous attributes, attributes within a certain range are considered similar. This can solve the problem that rough sets are not suitable for continuous attributes, and also improve the ability to resist noise interference. Attached Figure Description

[0044] Figure 1 This is a structural diagram of the data hierarchical fusion method based on fuzzy rough sets described in this invention.

[0045] Figure 2 This is a flowchart of the attribute importance classification method based on wavelet clustering described in this invention.

[0046] Figure 3 This is a flowchart of the attribute redundancy removal method based on fuzzy rough sets described in this invention. Detailed Implementation

[0047] The specific embodiments of the present invention will be further described below with reference to the accompanying drawings and technical solutions.

[0048] A data hierarchical fusion method based on fuzzy rough sets is proposed. This method includes an attribute importance hierarchical method based on wavelet clustering and an attribute redundancy removal method based on fuzzy rough sets.

[0049] Reference Figure 2 The specific operation process of the attribute importance ranking method based on wavelet clustering is as follows:

[0050] Step 1. Calculate the feature vector of attribute importance: {AS i ,|SK i |,K i The formula is as follows:

[0051]

[0052]

[0053]

[0054] Among them, H(A) i ) represents attribute A i Information entropy, H max (A i ) represents attribute A i The maximum discrete entropy, where n represents the number of data entries in dataset D. 's' represents the average value of the attribute, and 's' represents the standard deviation of the data.

[0055] Step 2. Quantize the feature space using the coefficient of variation, transform the original space into a grid, and then assign objects to corresponding cells to form a new feature space. The formula for calculating the coefficient of variation is as follows:

[0056]

[0057] Where s represents the standard deviation of the data. This represents the average value of the attribute.

[0058] Step 3. Perform wavelet transform on the feature space to compress the original data.

[0059] Step 4. Determine the threshold based on the distribution of data in the grid after wavelet transform. If the data is approximately symmetrically distributed, take the median as the threshold; if the data is left-skewed or right-skewed, it means the mode has a greater influence. In this case, if it is left-skewed, take the median plus the weighted value of the mode as the threshold; if it is right-skewed, take the median minus the weighted value of the mode as the threshold.

[0060] Step 5. Find the grids with a density greater than the threshold in the wavelet-transformed space and mark them as dense.

[0061] Step 6. For density-connected grids, treat them as a cluster and label them with the cluster number.

[0062] Step 7. Create a mapping table of cells before and after the transformation, and map the cluster labels to the original grid space.

[0063] Step 8. Map the raw data to their respective clusters according to their cluster labels.

[0064] Step 9. Calculate the center point of each cluster and determine the importance level of the cluster based on the characteristics of the center point.

[0065] Reference Figure 3 The specific operation process of the attribute redundancy removal method based on fuzzy rough sets is as follows:

[0066] Step 10. Select the attribute A with the highest sensitivity from the cluster with the highest importance. s As a property to be reduced.

[0067] Step 11. For the remaining attribute set Perform the traversal; if attribute A i If attribute A is a continuous attribute, then for any data objects x and y, calculate their fuzzy similarity; if attribute A i For discrete attributes, proceed directly to step 12. The formula for calculating fuzzy similarity is as follows:

[0068]

[0069] Where x and y are two arbitrary data objects. This indicates that data objects x and y are in property A i The conditions are similar. It is a Gaussian kernel function, and ε is a threshold, ε∈[0,1].

[0070] Step 12. Calculate the attribute to be reduced, A. s Discrimination matrix Among them l pq The p-th record and the q-th record are distinct attributes, and n represents the number of records in dataset D.

[0071] Step 13. Calculate the attribute to be reduced, A. s Discriminant function

[0072] Step 14. Set the pending meeting attribute A s Discriminant function Transform from conjunction normal form to disjunctive normal form.

[0073] Step 15. Based on the results of Step 9, select the attribute decision set with the lowest sensitivity.

[0074] Step 16. Transfer the attribute set Record as a new attribute set Simultaneously update the attribute set The modulus m.

[0075] Step 17. Repeat steps 10-16 until there are no remaining sensitive attributes or attribute sets. until.

Claims

1. A data hierarchical fusion method based on fuzzy rough sets, characterized in that, The steps are as follows: (1) Use wavelet clustering algorithm to cluster the importance of attributes, and use attribute sensitivity, skewness coefficient and kurtosis coefficient as features of attribute importance, and select the attribute with the highest importance as the attribute to be reduced; Clustering attributes based on their importance follows the specific process: (1.1) First, the importance of attributes is obtained by using attribute sensitivity, skewness coefficient, and kurtosis coefficient, and then the feature space is quantified by using the coefficient of variation. Divide the grid; Attribute sensitivity AS : For dataset D Attributes in The information content of an attribute differs, so the amount of attribute information is defined as the ratio of the difference between the maximum discrete entropy of the attribute and the attribute information entropy to the maximum discrete entropy of the attribute. , For dataset D The set of all attributes in the set is defined by the following formula: in, For attributes Information entropy For attributes The maximum discrete entropy; , The smaller the value, the more sensitive the attribute; conversely, the larger the value, the less sensitive the attribute. skewness coefficient SK : Used for measuring datasets D A certain attribute The degree of skewness, when calculating the skewness coefficient for ungrouped original attributes, is defined by the following formula: in, n Represents the dataset D The number of data entries in the data. Represents the dataset D The Middle j The attributes corresponding to each record The value, s Representing attributes The standard deviation of all values, Representing attributes The average of all possible values; This indicates that the data is symmetrically distributed. This indicates that the data is right-skewed. This indicates that the data is left-skewed. kurtosis coefficient K : Used for measuring datasets D A certain attribute The peak level is defined by the following formula: This indicates that the data follows a normal distribution; This indicates a peaked distribution, meaning the data is more concentrated. This indicates a flattened distribution, meaning the data is more dispersed. coefficient of variation Used to measure datasets D A certain attribute The degree of dispersion of the probability distribution is defined by the following formula: (1.2) After obtaining the above calculation results, the attributes of As a feature, the attribute set Clustering is performed; firstly, the feature space is... Perform wavelet transform to obtain the feature space after wavelet transform. Then, based on the feature space after wavelet transform The distribution of data determines the threshold, and grids with a density greater than the threshold are marked as dense. Then, dense and connected grids are grouped into a cluster and numbered. Finally, the data in the grid is labeled with the cluster number it belongs to. (1.3) Establish a mapping table to map cluster labels to the original feature space. , the original feature space The data in the data is mapped to their respective clusters according to the cluster labels, and the importance level of each cluster is determined according to the attribute sensitivity of the centroid of each cluster. Then, the attribute in the cluster with the highest importance is selected as the attribute to be reduced. (2) After selecting the attributes to be reduced, calculate the fuzzy similarity for the remaining continuous attributes. Discrete attributes do not need to be processed. Then, use the attribute reduction algorithm based on fuzzy rough set to calculate the minimum decision set of the attributes and select the attribute decision set with the lowest sensitivity as the attribute reduction set. The attribute redundancy removal is performed using an attribute reduction algorithm based on fuzzy rough sets. The specific process is as follows: (2.1) First, select the most sensitive attribute from the cluster with the highest importance as the attribute to be reduced. Then, iterate through the remaining attributes. If the attribute is a continuous attribute, calculate the fuzzy similarity between any two data objects under the attribute. If the attribute is a discrete attribute, no processing is required. Repeat this process until the fuzzy similarity has been calculated for all continuous attributes. Fuzzy similarity relation for continuous data objects: used to determine whether two arbitrary continuous data objects are similar. The formula is defined as follows: in, It is an arbitrary continuous attribute. They are two arbitrary data objects. Represents data objects In attributes The conditions are similar. It is a Gaussian kernel function. It is a threshold. ; (2.2) Calculate the discrimination matrix and discrimination function of the attribute to be reduced, and select the attribute decision set with the lowest sensitivity based on the clustering results of attribute importance.

Citation Information

Patent Citations

  • Neighborhood rough set reduction-based spectrum clustering method and system

    CN107169500A

  • Short-term load prediction method based on C-means clustering fuzzy rough set

    CN110245783A