A method and device for detecting the boundary of high-dimensional clustering data
By calculating k-nearest neighbor objects of high-dimensional data and generating index vectors, the problem of clustering boundary detection in high-dimensional space is solved, more efficient boundary recognition is achieved, and detection performance and accuracy are improved.
Patent Information
- Application Number
- CN202111178171.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-10-09
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2041-10-09
AI Technical Summary
The prior art is difficult to effectively identify cluster boundaries in high-dimensional space, resulting in a degradation of detection performance and affecting subsequent applications.
By calculating the k-nearest neighbors of the data point, obtaining the equilibrium coefficient and distance correction coefficient, generating an index vector, and using the index vector to determine the position of the boundary point in the high-dimensional data matrix.
The cluster boundary detection of high-dimensional data is realized, with better detection performance and higher accuracy, and can take into account the boundary recognition of two-dimensional and high-dimensional data.
Smart Images

Figure CN114037000B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of data mining, and particularly relates to a method and device for detecting the boundary of high-dimensional clustering data. Background Art
[0002] Data mining refers to the process of searching for information hidden in a large amount of data through algorithms. As a very useful technology in data mining, clustering analysis is mainly used to find implicit data distribution patterns and association rules from a large amount of data, so as to effectively conduct data mining. In the clustering analysis technology, the clustering data boundary, as a special pattern, focuses on those data that are distributed at the clustering edge, have a clear class membership, but are different from the data within the class to a certain extent. In the real world, it has wide practical significance, such as the carriers of a certain recessive genetic disease or recessive virus in a large medical dataset; abnormal gene fragments in gene expression profile data; abnormal handwritten signatures; target intruders in surveillance videos, etc. Existing domestic and foreign research teams have achieved certain success in the clustering boundary of low-dimensional space using geometric theory.
[0003] In 1996, M.Ester et al. first proposed the concept of clustering boundary, opening the door to clustering boundary detection. In 2006, Xia C Y et al. proposed the BORDER algorithm, which uses the reverse k-nearest neighbor technology to extract the clustering boundary. Since the number of reverse k-nearest neighbors of both the boundary points and noise points of the clustering is less than that of the central points, there are often more clustering noise points mixed in the detection results of this algorithm.
[0004] To make up for the deficiencies of the BORDER algorithm, Qiu Baozhi et al. proposed the BRIM algorithm in 2007. This algorithm identifies the boundary based on the characteristics that the neighborhood distribution of boundary points is uneven while the neighborhood distribution of clustering core points is approximately uniform. However, this algorithm is easily affected by the noise near the clustering boundary, especially unable to accurately extract the boundaries of variable density and multi-density clustering.
[0005] Xue Lixiang et al. proposed the BAND algorithm in 2009. This algorithm extracts boundary points based on the coefficient of variation of data objects, so it can overcome the shortcomings of the BRIM algorithm. However, since the coefficient of variation of noise points around the clustering may be the same as that of some boundaries, this algorithm will misjudge clustering noise points as boundaries.
[0006] The BRINK algorithm uses weighted Euclidean distance to measure the similarity between data points and also achieves good boundary detection results. However, as the data dimension increases, the sparsity of high-dimensional space causes the measurement of this similarity to gradually fail.
[0007] Cao Xiaofeng et al. proposed the Lever algorithm in 2016. This algorithm equates the distribution of high-dimensional data in the k-nearest neighbor space to a balance problem of a lever. The distribution of clustering core points is more balanced and stable than that of clustering boundary points. However, when dealing with high-dimensional data, this algorithm may encounter the problem that its divergence coefficient overflows and exceeds the calculation range.
[0008] In summary, the respective disadvantages of the existing clustering boundary inspection technologies will reduce the detection performance of clustering boundary points to varying degrees, and cannot effectively identify the clustering boundaries in high-dimensional data, which has an adverse impact on subsequent applications. Summary of the Invention
[0009] To solve the above problems existing in the prior art, the present invention provides a method and device for detecting the boundary of high-dimensional clustering data. The technical problems to be solved by the present invention are realized through the following technical solutions:
[0010] In a first aspect, the present invention provides a method for detecting the boundary of high-dimensional clustering data, including:
[0011] S1: Obtain a data matrix to be detected;
[0012] S2: Calculate the k-nearest neighbor objects of all data points in the data matrix to be detected;
[0013] S3: Calculate the balance coefficient and distance correction coefficient of the data points to be detected according to the k-nearest neighbor objects to obtain a balance coefficient vector and a distance correction coefficient vector;
[0014] S4: Calculate the product of the balance coefficient vector and the distance correction coefficient vector, and sort the obtained product vector to obtain an index vector;
[0015] S5: Determine the index positions of the boundary points in the data matrix to be detected according to the index vector to complete the detection of the clustering data boundary.
[0016] In an embodiment of the present invention, step S2 includes:
[0017] For each data point to be detected, calculate the Euclidean distance between it and the remaining data points to be detected, and select the k data points corresponding to the smallest k Euclidean distances as the k-nearest neighbor objects of the current data point to be detected, where k represents the number of nearest neighbors and k≥3.
[0018] In an embodiment of the present invention, in step S3, the formula for calculating the balance coefficient of the data points to be detected according to the k-nearest neighbor objects is:
[0019]
[0020] where a iDenote the data point to be detected as x i 's balance coefficient, x ij Denote the data point to be detected as x i The element value of the j-th dimension of, Denote the data point x i The element value of the j-th dimension of the p-th k-nearest neighbor object of, M represents the number of dimensions of the data.
[0021] In one embodiment of the present invention, in step S3, the formula for calculating the distance correction coefficient of the data point to be detected according to the k-nearest neighbor object is:
[0022]
[0023] Where, b i Denote the distance correction coefficient of the data point to be detected as x i , ‖·‖2 represents calculating the Euclidean distance, Denote the data x i 's p-th k-nearest neighbor object.
[0024] In one embodiment of the present invention, step S4 includes:
[0025] Multiply the corresponding elements of the balance coefficient vector a and the distance correction coefficient vector b to obtain the product vector c = ab, where c i = a i b i , i = 1, …, N;
[0026] Sort the elements in the product vector c from largest to smallest to obtain the sorted vector d;
[0027] Record the indices of the elements in the sorted vector d in the product vector c to obtain the index vector loc; where,
[0028] The i-th element d of the sorted vector d i Is denoted as loc i Denote the i-th element of the index vector loc, Denote the loc i -th element in the product vector c.
[0029] In one embodiment of the present invention, step S5 includes:
[0030] Set two threshold values λ1 and λ2 with different numerical values;
[0031] Take the elements of the position indices of the index vector loc between [floor(Nλ1), floor(Nλ2)] as row indices, and find the corresponding data in the data matrix to be detected, which are the clustering boundary points;
[0032] Among them, floor represents rounding down, N represents the total number of data to be detected, and 0 < λ1 < λ2 < 1.
[0033] In a second aspect, the present invention provides a high-dimensional clustering data boundary detection device, including:
[0034] A data acquisition module for acquiring a data matrix to be detected;
[0035] A first calculation module for calculating the k-nearest neighbor objects of all data points in the data matrix to be detected;
[0036] A second calculation module for calculating the balance coefficient and distance correction coefficient of the data points to be detected according to the k-nearest neighbor objects, and obtaining a balance coefficient vector and a distance correction coefficient vector;
[0037] A third calculation module for calculating the product of the balance coefficient vector and the distance correction coefficient vector, and sorting the obtained product vector to obtain an index vector;
[0038] A boundary detection module for determining the index position of the boundary points in the data matrix to be detected according to the index vector, so as to complete the clustering data boundary detection.
[0039] Advantages of the present invention:
[0040] 1. Compared with the prior art, the clustering data boundary detection method provided by the present invention can not only perform boundary detection on two-dimensional plane data, but also effectively identify the clustering boundaries of high-dimensional data;
[0041] 2. When calculating the balance coefficient, the high-dimensional data boundary detection method provided by the present invention adds a way of weighting by mass on the basis of the lever balance idea, so that the performance of this method for detecting clustering boundaries is better and the accuracy is higher.
[0042] The following will further describe the present invention in detail with reference to the drawings and embodiments. Description of the Drawings
[0043] Figure 1 is a schematic diagram of a high-dimensional clustering data boundary detection method provided by an embodiment of the present invention;
[0044] Figure 2 is a schematic diagram of the structure of a high-dimensional clustering data boundary detection device provided by an embodiment of the present invention;
[0045] Figure 3 is a schematic diagram of clustering data of different categories in a simulation experiment;
[0046] Figure 4It is a detection result graph of the boundaries of clustering data of different categories using the method of the present invention. Detailed implementation manners
[0047] The present invention will be further described in detail below with reference to specific embodiments, but the implementation manners of the present invention are not limited thereto.
[0048] Embodiment 1
[0049] Please refer to Figure 1 , Figure 1 which is a schematic diagram of a high-dimensional clustering data boundary detection method provided by an embodiment of the present invention, and includes:
[0050] S1: Obtain the data matrix to be detected.
[0051] Specifically, the data to be detected is organized into a data matrix with dimensions N×M, where N is the total number of data to be detected, M is the number of dimensions of the data, and generally, N≥50 and M≥2.
[0052] S2: Calculate the k-nearest neighbor objects of all data points in the data matrix to be detected.
[0053] First, set the number of nearest neighbors k. Preferably, k≥3.
[0054] Then, for each data point to be detected, calculate the Euclidean distance between it and the remaining data points to be detected, and select the k data points corresponding to the smallest k Euclidean distances as the k-nearest neighbor objects of the current data point to be detected.
[0055] Specifically, for each data point x to be detected i , calculate its Euclidean distance from the remaining N - 1 data, and select the k data points corresponding to the smallest k Euclidean distances as the k-nearest neighbor objects of x i . The result is recorded as an N×k matrix, denoted as A. The k elements in the i-th row of this matrix are respectively the row index values of the k-nearest neighbor objects of the i-th data point x to be detected in the N×M data matrix. Denote i as the p-th k-nearest neighbor object of data x , then i where A represents the element value of the i-th row and p-th column of matrix A. i,p
[0056] S3: Calculate the balance coefficient and distance correction coefficient of the data points to be detected according to the k-nearest neighbor objects, and obtain the balance coefficient vector and distance correction coefficient vector.
[0057] First, for each data x in the data matrix N×M to be detected i , calculate its balance coefficient a according to its k-nearest neighbor objectsi , The calculation formula is as follows:
[0058]
[0059] Where x ij represents the element value of the j-th dimension of the data point x to be detected i , represents the element value of the j-th dimension of the p-th k-nearest neighbor object of the data point x i .
[0060] Next, for each data x in the data matrix N×M to be detected i , according to its k-nearest neighbor objects, calculate its distance correction coefficient b i , and the calculation formula is as follows:
[0061]
[0062] Where ‖·‖2 represents calculating the Euclidean distance.
[0063] Record the N balance coefficients as the balance coefficient vector a, and record the N distance correction coefficients as the distance correction coefficient vector b.
[0064] In this embodiment, when calculating the balance coefficient, a method of weighted by mass is added on the basis of the lever balance idea, making the performance of this method for detecting the clustering boundary better and the accuracy higher.
[0065] S4: Calculate the product of the balance coefficient vector and the distance correction coefficient vector, and sort the obtained product vector to obtain an index vector.
[0066] First, multiply the corresponding elements of the balance coefficient vector a and the distance correction coefficient vector b to obtain the product vector c = ab, where c i = a i b i , i = 1, …, N;
[0067] Then, sort the elements in the product vector c from largest to smallest to obtain the sorted vector d.
[0068] Finally, record the indexes of the elements in the sorted vector d in the product vector c to obtain the index vector loc. Where the i-th element d of the sorted vector d i can be expressed as loc i represents the i-th element of the index vector loc, represents the loc i -th element in the product vector c.
[0069] S5: Determine the index positions of the boundary points in the data matrix to be detected according to the index vector, so as to complete the clustering data boundary detection.
[0070] First, set two thresholds λ1 and λ2 with different values.
[0071] Then, take the elements of the position index of the index vector loc between [floor(Nλ1), floor(Nλ2)] as row indexes, and find the corresponding data in the data matrix N×M to be detected, which are the clustering boundary points, that is, the data corresponding to the elements between [floor(Nλ1), floor(Nλ2)] as row indexes in the data matrix N×M to be detected are the clustering boundary points of the original data matrix.
[0072] Wherein, floor represents rounding down, N represents the total number of data to be detected, and 0 < λ1 < λ2 < 1.
[0073] Compared with the prior art, the clustering data boundary detection method provided in this embodiment can not only perform boundary detection on two-dimensional plane data, but also effectively identify the clustering boundaries of high-dimensional data.
[0074] Embodiment 2
[0075] Based on the above Embodiment 1, this embodiment also provides a high-dimensional clustering data boundary detection device. Please refer to Figure 2 , Figure 2 is a schematic structural diagram of a high-dimensional clustering data boundary detection device provided by an embodiment of the present invention, which includes:
[0076] A data acquisition module 1, configured to acquire a data matrix to be detected;
[0077] A first calculation module 2, configured to calculate the k-nearest neighbor objects of all data points in the data matrix to be detected;
[0078] A second calculation module 3, configured to calculate the balance coefficient and the distance correction coefficient of the data points to be detected according to the k-nearest neighbor objects, and obtain a balance coefficient vector and a distance correction coefficient vector;
[0079] A third calculation module 4, configured to calculate the product of the balance coefficient vector and the distance correction coefficient vector, and sort the obtained product vector to obtain an index vector;
[0080] A boundary detection module 5, configured to determine the index positions of the boundary points in the data matrix to be detected according to the index vector, so as to complete the clustering data boundary detection.
[0081] The high-dimensional clustering data boundary detection device provided in this embodiment can implement the high-dimensional clustering data boundary detection method provided in the above embodiment, and the detailed process will not be repeated here.
[0082] Therefore, the high-dimensional clustering data boundary detection device provided in this embodiment also has the advantages of not only being able to take into account both two-dimensional data and high-dimensional data for boundary detection, but also having better detection performance and higher accuracy.
[0083] Embodiment III
[0084] The beneficial effects of the present invention will be verified and described below through simulation experiments.
[0085] 1. Simulation conditions
[0086] The hardware platform for the simulation experiment in this embodiment is: Processor: 3.2GHz Intel Core i7, Memory 16GB.
[0087] The software platform for this simulation experiment is: Windows10 and Python 3.8.
[0088] The data set used in simulation experiment 1) is two-dimensional clustering data randomly generated by the "make_blobs" function in the sklearn.datasets functional package; the data set used in simulation experiment 2) is the medical data set Biomed, the data dimension of this data set is 4, and there are a total of 209 observation objects, including 134 normal objects and 75 virus carrier objects. Among the 134 normal objects, there are 30 potential virus infected persons, and these 30 potential virus infected persons are marked as the real data boundaries.
[0089] 2. Simulation content and result analysis
[0090] 1) Two-dimensional clustering data boundary detection simulation experiment
[0091] In this experiment, the "make_blobs" function in the sklearn.datasets functional package was used to randomly generate 3 two-dimensional clustering data, as Figure 3 shown, the clustering data of a total of 3 categories represented by triangles, squares, and circles.
[0092] Using the method of the present invention, the outlier points and boundary points of these 3 clusters can be detected at one time, and the results are as Figure 4 shown, the fork-shaped points are the clustering boundary points detected by the present invention, and the diamond-shaped points are the clustering outlier points detected by the present invention.
[0093] From Figure 3 and Figure 4 it can be seen that the method provided by the present invention can well detect the outlier points and boundary points in two-dimensional clustering data, and this experiment proves that the present invention can take into account the boundary detection function of two-dimensional clustering data.
[0094] 2) Simulation Experiment on the Performance of Detecting the Boundary of High-Dimensional Clustering Data
[0095] In this experiment, the medical dataset Biomed is used to detect the boundary of high-dimensional clustering data. The experimental parameters include accuracy, recall, and F-measure, and their respective definitions are as follows:
[0096]
[0097]
[0098] The experimental results are shown in the following table.
[0099]
[0100] Experiment 2) adopted the medical dataset Biomed, and the dimension number of each data is 4. This experiment compared the running results of the existing algorithms BAND, BORDER, BRINK, BERGE, Spinver, and Lever with the method of the present invention on this dataset. As can be seen from the above table, under the condition of taking into account both accuracy and recall, the present invention obtained the best experimental results, and the experimental results proved the effectiveness of the present invention.
[0101] The above content is a further detailed description of the present invention in combination with specific preferred implementation manners. It cannot be determined that the specific implementation of the present invention is only limited to these descriptions. For those of ordinary skill in the technical field to which the present invention belongs, without departing from the concept of the present invention, several simple deductions or substitutions can still be made, and all should be regarded as belonging to the protection scope of the present invention.
Claims
1. A method for detecting the boundary of high-dimensional clustering data, characterized in that Applied to the identification of potential virus-infected objects in a medical dataset, including: S1: Obtain the medical data matrix to be detected; wherein, the medical data matrix includes a four-dimensional medical dataset, and the four-dimensional medical dataset includes normal objects, virus carriers, and potential virus-infected objects; S2: Calculate the k-nearest neighbor objects of all data points in the medical data matrix; S3: Calculate the balance coefficient and distance correction coefficient of the data point to be detected based on the k-nearest neighbor objects to obtain a balance coefficient vector and a distance correction coefficient vector; Among them, the formula for calculating the balance coefficient of the data point to be detected based on the k-nearest neighbor objects is: Where a i represents the balance coefficient of the data point x i , x ij represents the element value of the j-th dimension of the data point x i , represents the element value of the j-th dimension of the p-th k-nearest neighbor object of the data point x i , and M represents the number of dimensions of the data; The formula for calculating the distance correction coefficient of the data point to be detected based on the k-nearest neighbor objects is: where b i represents the distance correction coefficient of the data point x i to be detected, ‖·‖2 represents the calculation of the Euclidean distance, represents the data x i of the p-th k-nearest neighbor object; S4: Calculate the product of the balance coefficient vector and the distance correction coefficient vector, and sort the obtained product vector to obtain an index vector; S5: Determine the index position of the boundary point in the data matrix to be detected according to the index vector to complete the clustering data boundary detection, thereby identifying potential virus-infected objects.
2. The high-dimensional clustering data boundary detection method according to claim 1, wherein Step S2 includes: For each data point to be detected, calculate the Euclidean distance between it and the remaining data points to be detected, and select the k data points corresponding to the smallest k Euclidean distances as the k-nearest neighbor objects of the current data point to be detected, where k represents the number of nearest neighbors, and k≥3.
3. The high-dimensional clustering data boundary detection method according to claim 1, characterized in that Step S4 includes: Multiply the corresponding elements of the balance coefficient vector a and the distance correction coefficient vector b to obtain the product vector c = ab, where c i = a i b i , i = 1, …, N; Sort the elements in the product vector c from largest to smallest to obtain a sorted vector d; Record the indexes of the elements in the sorted vector d in the product vector c to obtain an index vector loc; wherein, The i-th element d of the sorting vector d i is denoted as loc i denotes the i-th element of the index vector loc, denotes the loc-th i element in the product vector c.
4. The high-dimensional clustering data boundary detection method according to claim 1, wherein Step S5 includes: Set two threshold values λ1 and λ2 with different values; Use the elements with the position indexes of the index vector loc between [floor(Nλ1), floor(Nλ2)] as row indexes, and find the corresponding data in the data matrix to be detected, which are the clustering boundary points; Among them, floor represents rounding down, N represents the total number of data to be detected, and 0<λ1<λ2<1.
5. A high-dimensional clustering data boundary detection device for implementing the method according to any one of claims 1-4, characterized in that, The device is applied to the identification of potential virus-infected objects in a medical dataset, including: A data acquisition module (1) for obtaining the medical data matrix to be detected; wherein, the medical data matrix includes a four-dimensional medical dataset, and the four-dimensional medical dataset includes normal objects, virus carriers, and potential virus-infected objects; A first calculation module (2) for calculating the k-nearest neighbor objects of all data points in the medical data matrix; A second calculation module (3) for calculating the balance coefficient and distance correction coefficient of the data point to be detected based on the k-nearest neighbor objects to obtain a balance coefficient vector and a distance correction coefficient vector; A third calculation module (4) for calculating the product of the balance coefficient vector and the distance correction coefficient vector, and sorting the obtained product vector to obtain an index vector; A boundary detection module (5) for determining the index position of the boundary point in the data matrix to be detected according to the index vector to complete the clustering data boundary detection, thereby identifying potential virus-infected objects.