Data cleaning method

Through the improved multidimensional spatial anomaly detection method and nonlinear weighted correlation coefficient filling method, the shortcomings of traditional methods in abnormal detection and missing value filling are solved, and the accuracy and quality of data cleaning are significantly improved.

CN120179998APending Publication Date: 2025-06-20HAISHI (YANTAI) INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510258162.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-05
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

Traditional anomaly detection methods cannot flexibly adapt to the local density changes of data, resulting in mis-detection or missed detection; existing similarity calculation methods fail to effectively capture the nonlinear relationship between data points, especially in high-dimensional data, making it difficult to accurately judge the similarity of data points; traditional missing value filling methods ignore the complex correlation between features, resulting in low accuracy of filling results.

Method used

An improved multidimensional spatial anomaly detection method is used, combining the neighbor number scheduling mechanism based on local variance and a nonlinear dynamic weighting kernel function to identify and eliminate abnormal data points; the missing values ​​are filled by using the nonlinear weighting correlation coefficient filling method, and the influence between features is adjusted by introducing nonlinear weighting factors.

Benefits of technology

It significantly improves the accuracy of abnormal data point detection, reduces false detection and missed detection; improves the accuracy of abnormal detection in high-dimensional data; enhances the accuracy of missing value filling, captures the complex correlation between features, and improves data quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120179998A_ABST
    Figure CN120179998A_ABST
Patent Text Reader

Abstract

The invention relates to the field of data processing, in particular to a data cleaning method. Comprising the following steps: performing preliminary preprocessing on an original data set to obtain a data set, calculating abnormal data points in a local abnormal factor identification data set by adopting an improved multi-dimensional space anomaly detection method, and removing the abnormal data points; after the abnormal data points are removed, a non-linear weighted correlation coefficient filling method is adopted to fill the missing value, and the filled missing value is obtained; and after the missing value is filled, formatting and uniformly processing the data set. The problems that a traditional anomaly detection method usually adopts a fixed neighbor number to judge the anomaly of a data point and cannot flexibly adapt to local density change of data, false detection or missing detection is easily caused, and the accuracy of a detection result is affected are solved; the technical problem that an existing similarity calculation method mainly depends on a simple Euclidean distance or a linear kernel function and cannot effectively capture a nonlinear relation between data points is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing, and particularly to a data cleaning method. Background Art

[0002] With the continuous increase in the amount of data and the increasing complexity of data analysis requirements, data cleaning technology has become a core link in data processing. Currently, for outlier identification and missing value imputation in data cleaning, there are already a variety of mature technology applications. For example, outlier detection based on statistical methods has been widely used in the field of data preprocessing. In terms of missing value imputation, traditional methods such as mean imputation and regression imputation, although relatively simple, often cannot effectively capture the non-linear relationships between features when the data complexity is high. Therefore, methods based on neighboring values, k-nearest neighbor imputation, and model-based imputation methods have been gradually introduced. The imputation methods provide a basic framework for data cleaning, can improve the data quality to a certain extent, and promote the progress of data processing technology.

[0003] However, the existing data cleaning methods have the following technical problems: Traditional outlier detection methods usually use a fixed number of neighbors to judge the abnormality of data points, cannot flexibly adapt to the local density changes of data, and are prone to false detection or missed detection, thus affecting the accuracy of the detection results; The existing similarity calculation methods mainly rely on simple Euclidean distance or linear kernel functions, and fail to effectively capture the non-linear relationships between data points. Especially in high-dimensional data, it is difficult to accurately judge the similarity of data points, resulting in insufficient accuracy of outlier detection; Traditional missing value imputation methods such as mean imputation or nearest neighbor imputation ignore the complex correlations between features. When dealing with data with highly correlated features, the accuracy of the imputation results is low, which may have an adverse impact on subsequent operations. Summary of the Invention

[0004] The present invention provides a data cleaning method to solve the technical problems that traditional outlier detection methods usually use a fixed number of neighbors to judge the abnormality of data points, cannot flexibly adapt to the local density changes of data, are prone to false detection or missed detection, thus affecting the accuracy of the detection results; The existing similarity calculation methods mainly rely on simple Euclidean distance or linear kernel functions, and fail to effectively capture the non-linear relationships between data points. Especially in high-dimensional data, it is difficult to accurately judge the similarity of data points, resulting in insufficient accuracy of outlier detection; Traditional missing value imputation methods such as mean imputation or nearest neighbor imputation ignore the complex correlations between features. When dealing with data with highly correlated features, the accuracy of the imputation results is low, which may have an adverse impact on subsequent operations.

[0005] A data cleaning method of the present invention specifically includes the following technical solutions:

[0006] A data cleaning method, comprising the following steps:

[0007] S1: Perform preliminary preprocessing on the original data set to obtain a data set. Use an improved multi-dimensional space anomaly detection method to calculate the local anomaly factor to identify the abnormal data points in the data set, and eliminate the abnormal data points;

[0008] S2: After eliminating the abnormal data points, use the non-linear weighted correlation coefficient filling method to fill in the missing values to obtain the filled missing values; after filling in the missing values, format and unify the data set.

[0009] Preferably, the S1 specifically includes:

[0010] The improved multi-dimensional space anomaly detection method identifies the abnormal data points in the data set by combining the neighbor number scheduling mechanism based on local variance and the non-linear dynamic weighted kernel function.

[0011] Preferably, the S1 specifically includes:

[0012] The neighbor number scheduling mechanism based on local variance adjusts the neighbor number according to the local variance of each data point. By dynamically adjusting the neighbor number, the local density change of the data points in different data regions is captured; the non-linear dynamic weighted kernel function measures the similarity between data points by combining the Euclidean distance and the exponential function.

[0013] Preferably, the S1 specifically includes:

[0014] In the specific implementation process of the improved multi-dimensional space anomaly detection method, a neighbor number scheduling mechanism based on local variance is adopted. By calculating the local variance of the distance distribution between each data point and its neighbors, the neighbor number is dynamically adjusted. The specific formula is as follows:

[0015]

[0016] Where k i is the adjusted neighbor number of the i-th data point; d i is the i-th data point; i is the data point index variable; round(·) is the rounding operation; k base is the initial neighbor number; exp is the exponential function; σ i is the local variance of the data point d i ; σ max is the maximum local variance.

[0017] Preferably, the S1 specifically includes:

[0018] In the specific implementation process of the improved multi-dimensional space anomaly detection method, the non-linear dynamic weighted kernel function is used to map data points into a high-dimensional space, capture the non-linear relationship between data points, and obtain the similarity between different data points. The specific formula is as follows:

[0019]

[0020] Where K(d i , d j ) is the kernel function value between the i-th data point and the j-th data point, representing the similarity between data points d i and data point d j ; d i and d j are the i-th data point and the j-th data point respectively; i and j are data point index variables; D ij is the Euclidean distance between data points d i and data point d j ; σ is the scale parameter in the non-linear dynamic weighted kernel function; α is the balance coefficient.

[0021] Preferably, the S1 specifically includes:

[0022] Calculate the local outlier factor based on the kernel function value between data points and the adjusted number of neighbors to measure the outlier degree of data points. The specific formula is as follows:

[0023]

[0024] Where LOF(d i ) is the local outlier factor of data point d i ; represents the i -th neighbor of data point d ; and i are data point index variables; is the neighbor set representing data point d i ; and lrd(d i ) are the local densities of data points and data point d i respectively; is the Euclidean distance between data point d i and data point ; is the similarity between data point d i and data point ; k i is the adjusted number of neighbors of the i-th data point.

[0025] Preferably, the S1 specifically includes:

[0026] Set an anomaly threshold, compare the local anomaly factor with the anomaly threshold. When the local anomaly factor is greater than the anomaly threshold, determine that the data point is an abnormal data point and eliminate the abnormal data point.

[0027] Preferably, the S2 specifically includes:

[0028] In the implementation process of the non - linear weighted correlation coefficient filling method, calculate the non - linear weighted correlation coefficient between features and adjust the influence between data points through the non - linear weighted factor.

[0029] Preferably, the S2 specifically includes:

[0030] In the implementation process of the non - linear weighted correlation coefficient filling method, fill in the missing values based on the non - linear weighted correlation coefficient, introduce a filling formula. The filling formula captures the non - linear relationship between features by using the sum of squares of feature values and the non - linear weighted correlation coefficient, and obtains the filled missing values. The specific implementation formula is:

[0031]

[0032] Where, is the value after filling the missing feature p of the data point d i , that is, the filled missing value; m is the total number of features, which is the number of features of the data point; C pq is the non - linear weighted correlation coefficient between feature p and feature q; is the average value of feature q in the data set; x iq is the feature value of feature q in the data point d i in.

[0033] The beneficial effects of the technical solution of the present invention are:

[0034] 1. The present invention introduces a neighbor number scheduling mechanism based on local variance. By dynamically adjusting the neighbor number of each data point, it effectively adapts to the local density changes in different data regions. Compared with the traditional method with a fixed neighbor number, the neighbor number scheduling mechanism based on local variance adaptively adjusts the neighbor number according to the local variance of the data point, significantly improving the accuracy of abnormal data point detection, reducing false detection and missed detection, and thus improving the quality of data cleaning.

[0035] 2. The present invention introduces a non - linear dynamic weighted kernel function. By combining the Euclidean distance and the exponential function, it accurately calculates the similarity between data points. As an important part of the improved multi - dimensional space anomaly detection method, the non - linear dynamic weighted kernel function can more flexibly capture the non - linear relationship between data points, enhance the determination accuracy of the similarity between data points, and significantly improve the accuracy of anomaly detection in high - dimensional data.

[0036] 3. The present invention adopts a nonlinear weighted correlation coefficient filling method to fill missing values. By introducing a nonlinear weighting factor based on feature differences, the influence between different features is dynamically adjusted, thereby overcoming the limitations of the nonlinear filling method when processing complex data. Compared with traditional mean filling or nearest neighbor filling methods, the nonlinear weighted correlation coefficient filling method can more accurately capture the complex correlation between features, thereby improving the accuracy of missing value filling and data quality. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] Figure 1 The present invention is a flow chart of a data cleaning method. DETAILED DESCRIPTION

[0038] In order to further explain the technical means and effects adopted by the present invention to achieve the predetermined purpose of the invention, the technical scheme in the embodiment of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiment of the present invention. Obviously, the described embodiment is only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.

[0039] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs.

[0040] A specific scheme of a data cleaning method provided by the present invention is described in detail below with reference to the accompanying drawings.

[0041] Refer to the attached Figure 1 , which shows a flow chart of a data cleaning method provided by an embodiment of the present invention, the method comprising the following steps:

[0042] S1: Perform preliminary preprocessing on the original data set to obtain a data set, use an improved multi-dimensional space anomaly detection method to calculate the local anomaly factor to identify abnormal data points in the data set, and remove the abnormal data points;

[0043] First, preprocess the input original dataset. The preprocessing includes: deduplication to ensure that each data record is unique and avoid biases caused by duplicate data; removing irrelevant features, that is, according to the business requirements and data analysis objectives of the specific implementation scenario, screening out features that are irrelevant or redundant to the analysis, such as IDs, timestamps, etc., to reduce the computational burden and improve data processing efficiency; filtering the noise in the data, deleting records with obvious errors or excessive missing values, and processing invalid extreme values, such as abnormal points where the numerical values exceed the reasonable range; marking missing values and filling them with the mean value. The above preprocessing is a well-known means to those skilled in the art and will not be elaborated here. After preprocessing, a dataset is obtained. The dataset contains multiple data points, and each data point contains multiple features. The number of data points and the features in the data points are determined according to the specific implementation scenario.

[0044] An improved multi-dimensional space anomaly detection method is used to calculate the local anomaly factor to identify the abnormal data points in the dataset. The improved multi-dimensional space anomaly detection method can more accurately identify abnormal data points by combining the neighbor number scheduling mechanism based on local variance and the non-linear dynamic weighted kernel function, improving the accuracy and robustness of anomaly detection.

[0045] The neighbor number scheduling mechanism based on local variance flexibly adjusts the neighbor number according to the local variance of each data point, and can better adapt to the local density changes in different data regions, thus effectively capturing local anomalies. The neighbor number refers to the number used to determine the neighbor set of each data point during the anomaly detection process. By dynamically adjusting the neighbor number, the local density changes of data points in different data regions can be more accurately captured, improving the accuracy of identifying abnormal data points.

[0046] The non-linear dynamic weighted kernel function can measure the similarity between data points in the high-dimensional space by combining the Euclidean distance and the exponential function, improving the flexibility and accuracy of abnormal data point detection.

[0047] The specific calculation process of the improved multi-dimensional space anomaly detection method is as follows:

[0048] First, adopt the neighbor number scheduling mechanism based on local variance, and dynamically adjust the neighbor number by calculating the local variance of the distance distribution between each data point and its neighbors. The specific formula is as follows:

[0049]

[0050] where k i is the adjusted neighbor number of the i-th data point, which is dynamically adjusted according to the local variance σ i of the data point d i ; d iis the i-th data point; i is the data point index variable; round(·) is a rounding operation used to ensure that the calculated number of neighbors is an integer value; k base is the initial number of neighbors, which serves as a benchmark for adjusting the number of neighbors and is set according to specific implementation scenarios; exp is the exponential function; σ i is the local variance of the data point d i , which is calculated by computing the Euclidean distance between the data point d i and the initial neighbors of the data point and then calculating the variance of the Euclidean distance; σ max is the maximum local variance, which is obtained by calculating the local variances of all data points in the dataset and selecting the maximum value.

[0051] The data points are mapped to a high-dimensional space using a non-linear dynamic weighted kernel function to capture the non-linear relationships between the data points and improve the accuracy of detecting abnormal data points. The specific formula is as follows:

[0052]

[0053] where K(d i , d j ) is the kernel function value between the i-th data point and the j-th data point, representing the similarity between the data point d i and the data point d j ; d i and d j are the i-th data point and the j-th data point respectively; i and j are the data point index variables; D ij is the Euclidean distance between the data point d i and the data point d j ; σ is the scale parameter in the non-linear dynamic weighted kernel function, which is set by the expert experience method; α is the balance coefficient used to adjust the non-linear influence of the Euclidean distance D ij , and it is a constant that is set according to the requirements of the implementation scenario.

[0054] Finally, based on the kernel function values between the data points and the adjusted number of neighbors, the local outlier factor is calculated to measure the degree of abnormality of the data points. When the local outlier factor is large, it indicates that there is a difference in the local density between the data point and its neighbors, suggesting that the data point may be an abnormal data point. The formula for the local outlier factor is as follows:

[0055]

[0056] where LOF(d i ) is the local outlier factor of the data point d i , which is used to measure the data point d iWhen the local anomaly factor is large, it means that the data point has a significantly different local density from its neighbors, and the probability of being an abnormal data point is high. Represents data point d i No. Neighbors; and i are data point index variables; is the data point d i The neighbor set of , including the data point d i The nearest k i neighbors, and the distance is measured by Euclidean distance; and lrd(d i ) are data points and data point d i The local density measures the density around the data point. The calculation of local density is a well-known method for those skilled in the art and will not be described in detail here. is the data point d i and data points The Euclidean distance between is the data point d i and data points The similarity between i is the adjusted number of neighbors of the ith data point.

[0057] The anomaly threshold is set according to the specific implementation scenario, and the local anomaly factor is compared with the anomaly threshold. If the local anomaly factor is greater than the anomaly threshold, the data point is determined to be an anomaly data point and the anomaly data point is removed.

[0058] s2: After the abnormal data points are removed, the nonlinear weighted correlation coefficient filling method is used to fill the missing values ​​to obtain the filled missing values; after filling the missing values, the data set is formatted and uniformly processed.

[0059] After outliers are removed, the nonlinear weighted correlation coefficient filling method is used to fill in the missing values ​​of the markers. The missing values ​​refer to the missing feature values ​​of certain abnormal data points in the data set.

[0060] The nonlinear weighted correlation coefficient filling method is a method for filling missing values ​​by using the nonlinear relationship between features. First, the nonlinear weighted correlation coefficient between features is calculated, and the difference of feature values ​​is considered, and the influence between data points is adjusted by nonlinear weighting factors. Then, the nonlinear weighted correlation coefficient and other features in the data points are used to fill missing values. The nonlinear weighted correlation coefficient filling method can more accurately reflect the complex relationship in the data, thereby improving the accuracy of filling, and is suitable for situations with complex data structures.

[0061] The specific process of the non - linear weighted correlation coefficient filling method is as follows:

[0062] First, calculate the non - linear weighted correlation coefficient between features. By introducing a non - linear weighted factor based on feature differences, the non - linear weighted correlation coefficient measures the non - linear relationship between features, thus capturing the complex correlation between features more accurately.

[0063] The calculation formula of the non - linear weighted correlation coefficient is as follows:

[0064]

[0065] Among them, C pq is the non - linear weighted correlation coefficient between feature p and feature q, used to measure the non - linear relationship between feature p and feature q; p and q are index variables of features in the data points; n is the number of data points; i is the data point index variable; w(d i , p, q) is the non - linear weighted factor, which is obtained by using the non - linear weighted factor calculation formula based on the difference between feature p and feature q in data point d i , and is used to adjust the influence of the feature difference between data points on the calculation of the non - linear weighted correlation coefficient; x ip and x iq are the feature values of feature p and feature q in data point d i respectively; and are the average values of feature p and feature q in the data set respectively.

[0066] The calculation formula of the non - linear weighted factor is as follows:

[0067]

[0068] Among them, w(d i , p, q) is the non - linear weighted factor; x ip and x iq are the feature values of data point d i on feature p and feature q respectively.

[0069] The calculation formula of the non - linear weighted factor makes the influence of data points with large feature differences on weighting more stable by non - linearly processing the differences in feature values of different features in the data points, avoiding the excessive influence of extreme differences, thus capturing the non - linear relationship between features more accurately and improving the effect of data filling and analysis.

[0070] Based on the non - linear weighted correlation coefficient to fill in the missing values, the filling formula can better capture the non - linear relationship between features by using the sum of the squares of feature values and the non - linear weighted correlation coefficient, thus improving the accuracy of the filling result. The filling formula is as follows:

[0071]

[0072] Among them, is the value after filling the missing feature p of the data point d i , that is, the missing value after filling; m is the total number of features, which is the number of features of the data point and is set according to the specific implementation scenario; C pq is the non - linear weighted correlation coefficient between feature p and feature q; is the average value of feature q in the data set; x iq is the data point d i and the feature value of feature q in it.

[0073] After filling the missing values, according to the structural requirements of the data set set in the specific implementation scenario, format and unify the data set. Specifically, check and unify the data types. For example, convert categorical variables into appropriate formats by methods such as one - hot encoding or label encoding. Unify the ranges of numerical features, and scale the features to a unified scale by using standardization or normalization methods to avoid the influence of the feature values of some features on subsequent calculations due to large or small numerical values. Formatting and unifying processing are well - known means to those skilled in the art and will not be elaborated here.

[0074] In summary, a data cleaning method is completed.

[0075] The order of the invention embodiments is only for description and does not represent the superiority or inferiority of the embodiments. The processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0076] Each embodiment in this specification is described in a progressive manner. The same or similar parts between each embodiment can be referred to each other, and the key point of each embodiment is to illustrate the differences from other embodiments.

[0077] The above - mentioned embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of each embodiment of the present invention, and should all be included in the protection scope of the present invention.

Claims

1. A data cleaning method, characterized in that: The following steps are involved: S1: Perform preliminary preprocessing on the original data set to obtain a data set, use an improved multi-dimensional space anomaly detection method to calculate the local anomaly factor to identify abnormal data points in the data set, and remove the abnormal data points; s2: After the abnormal data points are removed, the nonlinear weighted correlation coefficient filling method is used to fill the missing values ​​to obtain the filled missing values; after filling the missing values, the data set is formatted and uniformly processed.

2. A data cleaning method according to claim 1, characterized in that: The S1 specifically includes: The improved multi-dimensional space anomaly detection method identifies abnormal data points in a data set by combining a neighbor number scheduling mechanism based on local variance and a nonlinear dynamic weighted kernel function.

3. A data cleaning method according to claim 2, characterized in that: The s1 specifically includes: The local variance-based neighbor number scheduling mechanism adjusts the number of neighbors according to the local variance of each data point, and captures the local density changes of data points in different data areas by dynamically adjusting the number of neighbors; the nonlinear dynamic weighted kernel function measures the similarity between data points by combining the Euclidean distance and the exponential function.

4. A data cleaning method according to claim 3, characterized in that: The S1 specifically includes: In the specific implementation process of the improved multi-dimensional space anomaly detection method, a neighbor number scheduling mechanism based on local variance is adopted. By calculating the local variance of the distance distribution between each data point and its neighbors, the number of neighbors is dynamically adjusted. The specific formula is as follows: Among them, k i is the adjusted number of neighbors of the ith data point; d i is the i-th data point; i is the data point index variable; round(·) is the rounding operation; k base is the initial number of neighbors; exp is the exponential function; σ i is the data point d i The local variance of max is the maximum local variance.

5. A data cleaning method according to claim 4, characterized in that: The s1 specifically includes: In the specific implementation process of the improved multi-dimensional space anomaly detection method, a nonlinear dynamic weighted kernel function is used to map data points to a high-dimensional space, capture the nonlinear relationship of data points, and obtain the similarity between different data points. The specific formula is as follows: Among them, K(d i , d j ) is the kernel function value between the i-th data point and the j-th data point, indicating that the data point d i and data point d j The similarity between i and d j are the i-th data point and the j-th data point respectively; i and j are the data point index variables respectively; D ij is the data point d i and data point d j The Euclidean distance between them; σ is the scale parameter in the nonlinear dynamic weighted kernel function; α is the balance coefficient.

6. A data cleaning method according to claim 5, characterized in that: The s1 specifically includes: Based on the kernel function value between data points and the adjusted number of neighbors, the local anomaly factor is calculated to measure the degree of anomaly of the data point. The specific formula is as follows: Among them, LOF(d i ) is the local anomaly factor of data point di; Represents data point d i No. Neighbors; and i are data point index variables; is the data point d i The set of neighbors of ; and lrd(d i ) are data points and data point d i The local density of is the data point d i and data points The Euclidean distance between is the data point d i and data points The similarity between i is the adjusted number of neighbors of the ith data point.

7. A data cleaning method according to claim 6, characterized in that: The s1 specifically includes: An abnormal threshold is set, and the local abnormal factor is compared with the abnormal threshold. When the local abnormal factor is greater than the abnormal threshold, the data point is determined to be an abnormal data point and the abnormal data point is removed.

8. A data cleaning method according to claim 1, characterized in that: The s2 specifically includes: In the implementation process of the nonlinear weighted correlation coefficient imputation method, the nonlinear weighted correlation coefficients between the features are calculated, and the influence between the data points is adjusted by the nonlinear weighting factor.

9. A data cleaning method according to claim 8, characterized in that: The s2 specifically includes: In the process of implementing the nonlinear weighted correlation coefficient filling method, the missing values ​​are filled based on the nonlinear weighted correlation coefficient, and a filling formula is introduced. The filling formula captures the nonlinear relationship between the features by using the square of the eigenvalue and the nonlinear weighted correlation coefficient to obtain the missing values ​​after filling. The specific implementation formula is: in, is the data point d i The missing feature p is the value after filling, that is, the missing value after filling; m is the total number of features, which is the number of features of the data point; C pq is the nonlinear weighted correlation coefficient between feature p and feature q; is the average value of feature q in the data set; x iq is the data point d i The eigenvalue of feature q in .