Data cleaning method based on directed neighborhood distance
A data cleaning and distance technology, applied in database models, relational databases, electrical digital data processing, etc., can solve problems such as inability to apply variable density data, uneven density of measured data, and lack of active and effective noise reduction.
- Summary
- Abstract
- Description
- Claims
- Application Information
AI Technical Summary
Problems solved by technology
Method used
Image
Examples
Embodiment
[0053] Such as figure 1 As shown, a data cleaning method based on the distance of the directed neighborhood includes the following steps:
[0054] Step 1, input the original data matrix X=[x 1 ,x 2 ,...,x m ] T m×n , where x 1 ~x m be m samples, m is the number of samples, and n is the data dimension (in this embodiment, the original data is taken from the UCI machine learning measured data Seeds, in the present embodiment, 5% random noise is added, m=201, n=7, The original data of Seeds is 7-dimensional, such as figure 2 shown). Set the number of neighbors k, data outlier rate τ, neighbor coefficient δ, density scaling factor ρ and density adjustment factor ε (generally 10 -4 ≤ε≤10 -2 ).
[0055] Step 2, calculate the distance between two samples in the original data matrix X based on the traditional Euclidean distance, and obtain the Euclidean distance matrix D.
[0056] Step 3, based on the Euclidean distance matrix D, select the k nearest neighbor samples of e...
PUM
Login to View More Abstract
Description
Claims
Application Information
Login to View More 


