Data cleaning method based on directed neighborhood distance

A data cleaning and distance technology, applied in database models, relational databases, electrical digital data processing, etc., can solve problems such as inability to apply variable density data, uneven density of measured data, and lack of active and effective noise reduction.

Pending Publication Date: 2021-02-02
ARMY ENG UNIV OF PLA
View PDF0 Cites 0 Cited by
  • Summary
  • Abstract
  • Description
  • Claims
  • Application Information

AI Technical Summary

Problems solved by technology

[0003] However, in practical applications, the density of measured data may not be uniform, and most existing outlier data removal algorithms are not well suited for variable density data.
The noise of the real environment will cause data offset, resulting in a decrease in the effectiveness of subsequent data analysis algorithms
In summary, the data cleaning algorithm has the following dilemmas: (1) It is difficult to apply to variable density data, especially to identify outlier data from variable density data; (2) Most of the existing methods pursue the robustness of the algorithm to noise , adaptability, lack of active and effective methods to reduce the impact of noise

Method used

the structure of the environmentally friendly knitted fabric provided by the present invention; figure 2 Flow chart of the yarn wrapping machine for environmentally friendly knitted fabrics and storage devices; image 3 Is the parameter map of the yarn covering machine
View more

Image

Smart Image Click on the blue labels to locate them in the text.
Viewing Examples
Smart Image
  • Data cleaning method based on directed neighborhood distance
  • Data cleaning method based on directed neighborhood distance
  • Data cleaning method based on directed neighborhood distance

Examples

Experimental program
Comparison scheme
Effect test

Embodiment

[0053] Such as figure 1 As shown, a data cleaning method based on the distance of the directed neighborhood includes the following steps:

[0054] Step 1, input the original data matrix X=[x 1 ,x 2 ,...,x m ] T m×n , where x 1 ~x m be m samples, m is the number of samples, and n is the data dimension (in this embodiment, the original data is taken from the UCI machine learning measured data Seeds, in the present embodiment, 5% random noise is added, m=201, n=7, The original data of Seeds is 7-dimensional, such as figure 2 shown). Set the number of neighbors k, data outlier rate τ, neighbor coefficient δ, density scaling factor ρ and density adjustment factor ε (generally 10 -4 ≤ε≤10 -2 ).

[0055] Step 2, calculate the distance between two samples in the original data matrix X based on the traditional Euclidean distance, and obtain the Euclidean distance matrix D.

[0056] Step 3, based on the Euclidean distance matrix D, select the k nearest neighbor samples of e...

the structure of the environmentally friendly knitted fabric provided by the present invention; figure 2 Flow chart of the yarn wrapping machine for environmentally friendly knitted fabrics and storage devices; image 3 Is the parameter map of the yarn covering machine
Login to View More

PUM

No PUM Login to View More

Abstract

The invention discloses a data cleaning method based on a directed neighborhood distance. The method comprises the steps of obtaining an original data matrix added with noise, calculating an Euclideandistance matrix, calculating a shared neighbor distance matrix, removing outlier samples,constructing a density scaling matrix based on the directed neighborhood distance, and scaling the Euclidean distance matrix by using the density scaling matrix to obtain a directed neighborhood distance matrix. The method is suitable for outlier data elimination of variable density data, the influence of noise on the data can be reversely reduced, and the accuracy of subsequent clustering, pattern recognition and manifold dimension reduction algorithms can be effectively improved.

Description

technical field [0001] The method belongs to the field of data mining, and specifically relates to a data cleaning method based on a directed neighborhood distance. Background technique [0002] With the rapid development of information technology, the channels for data generation and acquisition have increased, and the data styles have become diversified, such as audio and video data, image data, text data, etc. Mining useful information from data is the premise and necessary means of data analysis. Data cleaning is the primary task of data mining. The main purpose is to eliminate outlier data and adjust the mean, variance, and amplitude of the data. Commonly used methods for removing outlier data include local outlier factor (LOF) detection, boxplot method, Grubbs test method, clustering method, and data-based variance (covariance) distribution method. Commonly used data adjustment methods include Normalization, standardization, logistic regression, linear scaling, etc. ...

Claims

the structure of the environmentally friendly knitted fabric provided by the present invention; figure 2 Flow chart of the yarn wrapping machine for environmentally friendly knitted fabrics and storage devices; image 3 Is the parameter map of the yarn covering machine
Login to View More

Application Information

Patent Timeline
no application Login to View More
IPC IPC(8): G06F16/215G06F16/28
CPCG06F16/215G06F16/285
Inventor梁少军
OwnerARMY ENG UNIV OF PLA