A data processing method, device, equipment and storage medium

By using similarity density and similarity relationship data for clustering in high-dimensional data sets, the problem that distance-based clustering in the prior art is difficult to achieve efficient clustering, and the quality of clustering results is improved.

CN114064811BActive Publication Date: 2025-05-02CHINA MOBILE (SUZHOU) SOFTWARE TECH CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202010746217.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-07-29
Publication Date
2025-05-02
Estimated Expiration
2040-07-29

AI Technical Summary

Technical Problem

In the prior art, clustering is performed based on the distance between the data point and the clustering center, making it difficult to realize effective clustering in high-dimensional data sets, and the clustering results are of poor quality.

Method used

By obtaining similarity data for the data set, the similarity density and similarity relationship data for each data point are determined, and clustered based on these data, rather than relying on distance relationships.

Benefits of technology

The quality of clustering results of high-dimensional data sets is improved, making clustering results more accurate and effective.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114064811B_ABST
    Figure CN114064811B_ABST
Patent Text Reader

Abstract

The present application proposes a data processing method, firstly obtaining similarity data of a data set, wherein the similarity data represents the similarity between data points in the data set, and then determining the similarity density and similarity relationship data of each data point in the data set based on the similarity data; finally clustering the data set based on the similarity density and similarity relationship data of each data point in the data set. Since the data processing method clusters the similarity density and similarity relationship data of each data point in the data set determined by the similarity data of the data set, and does not cluster based on the distance between the data, the method is used to cluster high-dimensional data sets with relatively sparse data distribution, and the quality of the clustering results obtained is better. The present application also proposes a data processing device, an electronic device, and a computer storage medium.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to data processing technology, and in particular to a data processing method, device, electronic equipment and computer storage medium. Background Art

[0002] In the related art, the clustering method for high-dimensional data is mainly based on the distance between data to perform clustering. Specifically, it is necessary to assign each data point to the cluster center closest to it according to the distance between each data point and each sub-cluster center to achieve data clustering. Since the data distribution of high-dimensional data sets is relatively sparse and the distances between data points are almost equal, it is difficult to achieve clustering based on the distance between each data point and each sub-cluster center, and the quality of the obtained clustering results is also poor. Summary of the invention

[0003] The present application hopes to provide a technical solution for data processing to solve the problem in the related art that clustering based on the distance between each data point and each sub-cluster center is not only difficult to implement but also the quality of the obtained clustering results is also poor.

[0004] The present application provides a data processing method, the method comprising:

[0005] Acquire similarity data of a data set, wherein the similarity data represents similarities between data points in the data set;

[0006] Determine the similarity density and similarity relationship data of each data point in the data set according to the similarity data; the similarity density represents the number of data points in the corresponding similarity neighborhood of each data point, and the similarity relationship data represents the magnitude relationship between the similarities of each data point and at least two other data points;

[0007] The data set is clustered according to the similarity density and similarity relationship data of each data point in the data set.

[0008] In one implementation, determining the similarity density of each data point in the data set according to the similarity data includes:

[0009] The similarity neighborhood of each data point is determined according to a preset similarity threshold, and the similarity density of each data point in the data set is determined according to the similarity neighborhood of each data point and the similarity between each data point in the data set.

[0010] In one embodiment, clustering the data set according to the similarity density and similarity relationship data of each data point in the data set includes:

[0011] Determine the clustering parameters of each data point according to the similarity data of the data set and the similarity relationship data of each data point in the data set; each data point includes a first data point and a second data point, the first data point represents a data point in the data set with a non-maximum similarity density, and the second data point is a data point in the data set with a maximum similarity density;

[0012] The data set is clustered according to the similarity density and clustering parameters of each data point in the data set.

[0013] In one implementation, determining the clustering parameters of each data point based on the similarity data of the data set and the similarity relationship data of each data point in the data set includes:

[0014] Selecting a data point with the smallest first calculated value from among the data points with a similarity density greater than the first data point, and using the first calculated value of the selected data point as a clustering parameter of the first data point, wherein the first calculated value is negatively correlated with the second calculated value, and the second calculated value represents the similarity between the data point with a similarity density greater than the first data point and the first data point;

[0015] Among the data points other than the second data point in the data set, the data point with the largest third calculated value is selected, and the third calculated value of the selected data point is used as the clustering parameter of the second data point. The third calculated value is negatively correlated with the fourth calculated value, and the fourth calculated value represents the similarity between the second data point and the data points other than the second data point in the data set.

[0016] In one embodiment, clustering the data set according to the similarity density and clustering parameters of each data point in the data set includes:

[0017] According to the clustering parameters and similarity density of each data point in the data set, determining the data point in the data set whose similarity density is greater than a first threshold and whose clustering parameter is greater than a second threshold as the clustering center of the data set;

[0018] Clustering the other data points outside the cluster center according to the similarity density of each other data point outside the cluster center in the data set. In one embodiment, clustering the other data points outside the cluster center according to the similarity density of each other data point outside the cluster center in the data set includes:

[0019] According to the similarity density of each other data point outside the cluster center, a data point set corresponding to each other data point outside the cluster center is determined among the other data points outside the cluster center in the data set, wherein the similarity density of the data points in the data point set is greater than the similarity density of each corresponding other data point;

[0020] In each of the data point sets, determining a third data point having the largest similarity density;

[0021] It is determined that the category of each other data point outside the cluster center is the same as the category of the corresponding third data point.

[0022] In one implementation, the obtaining of similarity data of the data set includes:

[0023] Acquire relevant data of the data set, wherein the relevant data is a matrix;

[0024] The similarity data of the data set is determined according to the column vector of the relevant data of the data set.

[0025] The present application also provides a data processing device, which includes: an acquisition module, a determination module and a clustering module, wherein:

[0026] The acquisition module is used to acquire similarity data of the data set, wherein the similarity data represents the similarity between data points in the data set;

[0027] The determination module is used to determine the similarity density and similarity relationship data of each data point in the data set according to the similarity data; the similarity density represents the number of data points in the corresponding similarity neighborhood of each data point, and the similarity relationship data represents the magnitude relationship between the similarities of each data point and at least two other data points;

[0028] The clustering module is used to cluster the data set according to the similarity density and similarity relationship data of each data point in the data set.

[0029] In one implementation, the determination module is used to determine the similarity density of each data point in the data set according to the similarity data, including:

[0030] The similarity neighborhood of each data point is determined according to a preset similarity threshold, and the similarity density of each data point in the data set is determined according to the similarity neighborhood of each data point and the similarity between each data point in the data set.

[0031] In one embodiment, the clustering module is used to cluster the data set according to the similarity density and similarity relationship data of each data point in the data set, including:

[0032] Determine the clustering parameters of each data point according to the similarity data of the data set and the similarity relationship data of each data point in the data set; each data point includes a first data point and a second data point, the first data point represents a data point in the data set with a non-maximum similarity density, and the second data point is a data point in the data set with a maximum similarity density;

[0033] The data set is clustered according to the similarity density of each data point in the data set, the data points and the clustering parameters.

[0034] In one embodiment, the clustering module is used to determine the clustering parameters of each data point according to the similarity data of the data set and the similarity relationship data of each data point in the data set, including:

[0035] Selecting a data point with the smallest first calculated value from among the data points with a similarity density greater than the first data point, and using the first calculated value of the selected data point as a clustering parameter of the first data point, wherein the first calculated value is negatively correlated with the second calculated value, and the second calculated value represents the similarity between the data point with a similarity density greater than the first data point and the first data point;

[0036] Among the data points other than the second data point in the data set, the data point with the largest third calculated value is selected, and the third calculated value of the selected data point is used as the clustering parameter of the second data point. The third calculated value is negatively correlated with the fourth calculated value, and the fourth calculated value represents the similarity between the second data point and the data points other than the second data point in the data set.

[0037] In one embodiment, the clustering module is used to cluster the data set according to the similarity density and clustering parameters of each data point in the data set, including:

[0038] According to the clustering parameters and similarity density of each data point in the data set, determining the data point in the data set whose similarity density is greater than a first threshold and whose clustering parameter is greater than a second threshold as the clustering center of the data set;

[0039] According to the similarity density of each other data point outside the cluster center in the data set, the other data points outside the cluster center are clustered.

[0040] In one embodiment, the clustering module is used to cluster other data points outside the cluster center according to the similarity density of each other data point outside the cluster center in the data set, including:

[0041] According to the similarity density of each other data point outside the cluster center, a data point set corresponding to each other data point outside the cluster center is determined among the other data points outside the cluster center in the data set, wherein the similarity density of the data points in the data point set is greater than the similarity density of each corresponding other data point;

[0042] In each of the data point sets, determining a third data point having the largest similarity density;

[0043] It is determined that the category of each other data point outside the cluster center is the same as the category of the corresponding third data point.

[0044] In one implementation, the acquisition module is used to acquire similarity data of a data set, including:

[0045] Acquire relevant data of the data set, wherein the relevant data is a matrix;

[0046] The similarity data of the data set is determined according to the column vector of the relevant data of the data set.

[0047] The present application also provides an electronic device, comprising a processor and a memory for storing a computer program that can be run on the processor; wherein,

[0048] When the processor is used to run the computer program, it executes any one of the data processing methods described above.

[0049] The present application also provides a computer storage medium on which a computer program is stored. When the computer program is executed by a processor, any of the above-mentioned data processing methods is implemented.

[0050] It can be seen that the data processing method in the present application first obtains the similarity data of the data set, the similarity data represents the similarity between the data points in the data set, and then determines the similarity density and similarity relationship data of each data point in the data set based on the similarity data, the similarity density represents the number of data points of each data point in the similarity neighborhood, and the similarity relationship data represents the size relationship between the similarities of each data point and at least two other data points; finally, according to the similarity density and similarity relationship data of each data point in the data set, the data set is clustered. Since the data processing method clusters the similarity density and similarity relationship data of each data point in the data set determined by the similarity data of the data set, and does not cluster according to the distance between the data, the method is used to cluster high-dimensional data sets with relatively sparse data distribution, and the quality of the clustering results obtained is better; further, the method of constructing a similarity matrix through the correlation (representation coefficient) between data makes the similarity between data points belonging to the same subspace as large as possible, and the similarity between data points in different subspaces as small as possible.

[0051] It is to be understood that both the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] The accompanying drawings herein are incorporated into and constitute a part of the specification. These drawings illustrate embodiments consistent with the present invention and, together with the specification, are used to explain the technical solutions of the present invention.

[0053] Figure 1 A flowchart of a data processing method of the present application;

[0054] Figure 2 A schematic diagram of a similarity matrix constructed for the first 10 categories of data from the Olivetti Research Laboratory (ORL) face dataset in this application;

[0055] Figure 3 is the data point x in category 1 of the ORL face dataset of this application i Schematic diagram of the similarity between data points in category 1 and category 2;

[0056] Figure 4 This is a schematic diagram of the change in the accuracy of the clustering results corresponding to the change in the similarity difference control parameter γ value from 2 to 9 in the present application;

[0057] Figure 5 This is a schematic diagram of the change in the accuracy of the clustering results corresponding to the application when γ=7 and the similarity threshold ε changes from 0.01 to 0.06;

[0058] Figure 6 This is a schematic diagram of the structure of the data processing device of the present application;

[0059] Figure 7 This is a schematic diagram of the structure of the electronic device of the present application. DETAILED DESCRIPTION

[0060] The present invention will be further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the embodiments provided herein are only used to explain the present invention and are not intended to limit the present invention. In addition, the embodiments provided below are partial embodiments for implementing the present invention, rather than providing all embodiments for implementing the present invention. In the absence of conflict, the technical solutions recorded in the embodiments of the present invention can be implemented in any combination.

[0061] It should be noted that, in the embodiments of the present invention, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a method or apparatus including a series of elements includes not only the elements explicitly recorded, but also includes other elements not explicitly listed, or also includes elements inherent to the implementation of the method or apparatus. In the absence of further restrictions, an element defined by the sentence "includes a ..." does not exclude the existence of other related elements (such as steps in a method or units in a device, for example, a unit may be a part of a circuit, a part of a processor, a part of a program or software, etc.) in the method or apparatus including the element.

[0062] The term "and / or" herein is only a description of the association relationship of the associated objects, indicating that there may be three relationships. For example, U and / or V may indicate the existence of U alone, the existence of U and V at the same time, and the existence of V alone. In addition, the term "at least one" herein indicates any combination of at least two of any one or more of a plurality of. For example, at least one of U, V, and W may indicate any one or more elements selected from the set consisting of U, V, and W.

[0063] For example, the data processing method provided by the embodiment of the present invention includes a series of steps, but the data processing method provided by the embodiment of the present invention is not limited to the recorded steps. Similarly, the data processing device provided by the embodiment of the present invention includes a series of modules, but the device provided by the embodiment of the present invention is not limited to including the modules explicitly recorded, and may also include modules required to obtain relevant information or perform processing based on information.

[0064] The embodiments of the present invention can be applied to hardware such as terminals and servers or computer systems composed of hardware, and can operate with many other general or special computing system environments or configurations, or can be implemented by a processor running computer executable code. Here, the terminal can be a thin client, a thick client, a handheld or laptop device, a microprocessor-based system, a set-top box, a programmable consumer electronic product, a network personal computer, a small computer system, etc., and the server can be a server computer system, a small computer system, a large computer system, and a distributed cloud computing technology environment including any of the above systems, etc.

[0065] Electronic devices such as terminals and servers can be described in the general context of computer system executable instructions (such as program modules) executed by a computer system. In general, program modules can include routines, programs, object programs, components, logic, data structures, etc., which perform specific tasks or implement specific abstract data types. Computer systems / servers can be implemented in a distributed cloud computing environment, where tasks are performed by remote processing devices linked through a communication network. In a distributed cloud computing environment, program modules can be located on local or remote computing system storage media including storage devices.

[0066] With the rapid development of information technology, data is everywhere in our daily life. The huge scale and complex structure of data bring many challenges to data processing. How to effectively mine valuable information from data has become a major problem faced by people in the new era. With the introduction of classic clustering algorithms, clustering algorithms have been able to effectively solve the problem of clustering low-dimensional data. However, the application environment is changing with each passing day, and high-dimensional data can be seen everywhere in people's lives. Among them, the dimensions of various image data, video data, and text data are often as high as tens of thousands of dimensions. For example, a picture taken by a smartphone can reach tens of thousands of pixels; traditional clustering algorithms often cannot obtain ideal results when dealing with high-dimensional data clustering problems. The main problems faced in clustering high-dimensional data include: the data distribution in high-dimensional space is more sparse than that in low-dimensional space, the distance between data is almost equal, and there are some irrelevant attributes in the data. Therefore, it is usually impossible to achieve clustering based on the distance relationship between data in high-dimensional space, and most traditional clustering methods are based on distance clustering. How to design new clustering algorithms to solve the clustering problem of high-dimensional data has become the focus of research in fields such as data mining and machine learning.

[0067] The main purpose of clustering a data set is to divide the samples in the data set into several non-overlapping subsets, each subset represents a category, and the data points of the same category have a high similarity, while the data points of different categories have a low similarity. By clustering the data set, the data distribution characteristics hidden in the data set can be effectively discovered, thus laying a good foundation for further full and effective use of the data. According to different usage methods, clustering algorithms can be divided into the following categories: partition clustering algorithms, hierarchical clustering algorithms, density-based clustering algorithms, grid-based clustering algorithms, and model-based clustering algorithms. The choice of clustering algorithm depends largely on the characteristics of the data. In fact, an excellent clustering algorithm should be able to handle data sets with a variety of data distribution shapes, and use as few parameters as possible to assign data points to different subsets, thereby obtaining meaningful results.

[0068] The k-means algorithm is one of the earliest partitioning clustering algorithms and has been used in many application fields. The k-means algorithm clusters the data set by continuously iterating and updating the cluster center until convergence. The method itself is simple and effective. However, this simple partitioning strategy makes the k-means algorithm sensitive to noise points, which will be affected by the initial cluster center and can only be applied to the case of spherical cluster distribution. The spherical clustering means that the clustered points are spherical around the cluster center. The k-modes algorithm is an extension of the k-means algorithm. The algorithm uses difference instead of the distance in the k-means algorithm. The smaller the difference, the smaller the distance, so that it can process categorical attribute data.

[0069] The DBSCAN algorithm is another popular density-based clustering algorithm. The algorithm uses the number of data points contained in the neighborhood of each data point with a fixed radius as the density of the data point; then the density threshold is used to divide the data points into core data points and non-core data points, and the core data points are continuously condensed to form clusters by other core data points in the neighborhood. The significant advantages of the DBSCAN algorithm are fast clustering speed and the ability to effectively handle noise points and realize spatial clustering of arbitrary shapes. However, the DBSCAN algorithm is sensitive to two parameters set by the user. Slight differences in parameter settings may cause great changes in clustering results. In order to overcome the shortcomings of the DBSCAN algorithm, the OPTICS algorithm and the NBC algorithm have made further improvements to the DBSCAN algorithm.

[0070] Recently, a new clustering algorithm, Density Peaks Clustering (DPC), has been proposed. This algorithm combines the advantages of partition clustering algorithm and density clustering algorithm, and provides a new idea for clustering algorithm. The algorithm clusters the data by finding the cluster center. It can be considered that the cluster center of each category should have two characteristics:

[0071] 1) The cluster center itself has a high density and is surrounded by low-density data points;

[0072] 2) The distance between the cluster center and other data points with higher density is relatively larger.

[0073] Based on these two assumptions, two attributes are set for each data point: local density ρ i and the minimum distance δ to the data point with greater density i The larger the values ​​of these two attributes, the greater the possibility that the point is the cluster center. The algorithm manually selects data points with large local density and distance as cluster centers. The algorithm is simple and clear. Based on the distance information between data points, it obtains the similarity between data points and then clusters data sets of any shape. In the past two years, the DPC algorithm has been widely used in many fields, such as remote sensing technology, biophysics, time series mining, image processing, and computer vision.

[0074] Compared with other clustering algorithms, the DPC algorithm has the following three advantages:

[0075] Advantage 1: The DPC algorithm does not require prior knowledge of data distribution to cluster data, while other clustering algorithms do. Specifically, the k-means algorithm needs to know the number of data categories, and the DBSCAN algorithm requires the user to set the radius r neighborhood to contain at least the minimum number MinPts. The radius r neighborhood can be the range of a circle with a radius of r. In addition, the DPC algorithm uses fewer parameter initialization settings, and has a certain robustness to the clustering results of the parameters. Specifically, the DPC algorithm only performs a certain amount of optimization on the data point density ρ. i It is sensitive to the relative value of , but has a certain robustness to the setting of the intercept distance dc.

[0076] Advantage 2: The DPC algorithm can cluster data of any distribution shape, which highlights the superiority of the DPC algorithm over the k-means algorithm.

[0077] Advantage 3: The clustering results of the DPC algorithm for the same data set are always unique. The DPC algorithm does not contain random factors and can obtain consistent clustering results, while other clustering algorithms, such as the k-means algorithm and the Expectation-Maximization algorithm (EM), may obtain different clustering results due to different initial iteration states.

[0078] In related technologies, with the introduction of some classic clustering algorithms, clustering algorithms have been able to solve the problem of clustering data in low-dimensional space. However, with the continuous changes in the application environment, especially in the era of "big data", the huge scale of data and the complexity of the structure have posed increasingly severe challenges to clustering analysis. High-dimensional data are becoming more and more common, including various image data, biological gene expression data, and search engine data, which often have dimensions of tens of thousands. Traditional clustering algorithms are usually designed and developed for low-dimensional data. When analyzing and processing high-dimensional data, they usually encounter serious bottlenecks and cannot meet the sparsity of high-dimensional data and avoid the impact of the "dimensionality disaster", and the expected results cannot be obtained. How to design and develop high-dimensional data clustering algorithms to meet the growing needs is becoming an important research topic in fields such as data mining and pattern recognition.

[0079] Currently, these clustering methods all perform clustering based on the distance between data, and are unable to cluster high-dimensional data sets with sparse data distribution, resulting in poor clustering effects.

[0080] In order to solve the above technical problems, an embodiment of the present invention proposes a data processing method. Figure 1 A flowchart of a data processing method of the present application is shown in FIG. Figure 1 As shown, the process may include:

[0081] Step 101: Acquire similarity data of a data set, wherein the similarity data represents the similarity between data points in the data set.

[0082] In one embodiment, the data set may be a data set corresponding to high-dimensional data, specifically, the data set may be image data, biological gene expression data, search engine data, etc. Of course, the data set may also be a data set corresponding to low-dimensional data, which is not specifically limited here. The similarity data of the data set may be a similarity matrix corresponding to the data set, and each element in the similarity matrix represents the similarity between data points. For example, for the element C in the 2nd row and the 3rd column of the similarity matrix, 23 It represents the similarity between data points x2 and x3.

[0083] The implementation method for obtaining the similarity data of a data set may be to first obtain the relevant data of the data set, the relevant data matrix, and then construct the similarity data of the data set based on the relevant data of the data set. Of course, the similarity data of the data set may also be obtained by other methods, which are not specifically limited here.

[0084] In one example, the dataset can be represented as X = [x1, x2, ..., x n ]=[X1,X2,…,X k ]∈R m×n , the data set X is taken from k independent linear subspaces Among them, k represents the data type of the data set, m represents the dimension of the data set, n represents the number of data samples in the data set, and X i represents the i-th data sample in the data set, i can be 1, 2, 3...k, X i is m×n i The matrix, X i Each column of comes from the same subspace S i , and n1+n2+n3+……+n i =n, if the data sample is sufficient: n i >rank(X i ), it can be seen that the data set X is arranged according to the type of data samples. Of course, the data set X can also be arranged according to the type of data samples. The implementation method for obtaining the relevant data of the data set X can be to obtain the correlation matrix Z of the data set X * Specifically, the correlation matrix Z of the data set X can be determined by the objective function * , the objective function is shown in formula (1):

[0085]

[0086] Wherein, formula (1) represents the optimal solution Z of the correlation matrix Z under the constraint condition X = AZ + E. * , the correlation matrix Z * It can be called the low rank representation (LRR) of the dataset X with respect to the dictionary A. E is the identity matrix, the dictionary A can be the dataset X, λ is the balance parameter, |||| 2,1 represents the 21 norm, |||| * Represents kernel clustering.

[0087] The implementation method of determining the correlation matrix Z of the data set X by formula (1) can be to solve the objective function by using the augmented Lagrange method (ALM). Specifically, an auxiliary variable J can be introduced, and J=Z can be set, then the objective function can be converted into a first equivalent function, see formula (2):

[0088]

[0089] And use the augmented Lagrange multiplier method to reconstruct the first equivalent function to obtain the second equivalent function, see formula (3):

[0090]

[0091] Among them, |||| F represents the Frobenius norm, <Y A ,X-AZ-E> represents Y A The angle between X-AZ-E, <Y B ,ZJ> indicates Y B The angle between ZJ and Y A , Y B represents the Lagrange multiplier, μ represents the penalty parameter, which is used to control the convergence of the function; the optimal solution matrix Z of the second equivalent function can be obtained by the LRR optimization algorithm * , for example, by inputting a dataset X∈ m×n , and perform the following initialization: set the maximum number of iterations maxIter = 1000, the current number of iterations k = 0, initialize Z = J = 0, E = 0, Y A =0, Y B =0. μ=10 -6 max μ =10 10 , ρ=1.1,ε=10 -8 , and then execute the following While statement,

[0092] While(||ZJ|| ∞ >εor||X-AZ-E|| ∞ >ε)

[0093] 1: Fix other variables and update variable J

[0094]

[0095] 2: Fix other variables and update variable Z

[0096]

[0097] 3: Fix other variables and update variable E

[0098]

[0099] 4: Update the Laplace multiplier

[0100] Y A =Y A +μ(X-AZ-E) (7);

[0101] Y B =Y B +μ(ZJ) (8);

[0102] 5: Update parameter μ

[0103] μ=min(ρμ,max μ ) (9);

[0104] k=k+1 (10);

[0105] end while

[0106] To output the correlation matrix Z * .

[0107] For a data set X arranged according to the type of data samples, the correlation matrix Z * See formula (11):

[0108]

[0109] Yes i ×n i The matrix represents the subspace S i The correlation between the data in the subspace can be the representation coefficient between the data, where i = 1, 2, 3 ... k. The representation coefficient between the data in the same subspace is large, and the representation coefficient between different subspaces is zero.

[0110] Correlation matrix Z * The block diagonal structure reveals the subspace properties of the data. The number of blocks corresponds to the number of subspaces, and the size of each block represents the dimension of the corresponding subspace. The data in the same block belongs to the same subspace. Of course, for data sets X that are not arranged according to the type of data samples, the correlation matrix Z obtained by the above solution is * It is not a block diagonal structure, and the correlation matrix Z * It is still possible to represent the correlation between data points in the data set.

[0111] It can be seen that the low-rank representation of the data set is to process high-dimensional data from another perspective. The purpose of the low-rank representation is to mine data clusters in different subspaces in the same data set. Since the data in the same subspace can be regarded as generated by the same set of bases, it means that the generation structure of the data in the same subspace is the same, so the data points in the same subspace can be considered to be the same category. The low-rank representation divides the original data space into different subspaces and searches for the possibility of the existence of different categories in the subspace. It can more accurately describe the relationship between the original data, reveal the essential structure of the data set, and has strong robustness to noisy data sets.

[0112] The implementation method of constructing the similarity data of the data set based on the relevant data of the data set may be to determine the similarity data of the data set based on a column vector of the relevant data of the data set.

[0113] In one example, determining the similarity data of the data set according to the column vector of the relevant data of the data set may be determining the similarity data of the data set according to the generalized angle cosine value of the column vector of the relevant data of the data set. The similarity data may be a similarity matrix C. For example, the elements of the similarity matrix C may be obtained by the following formula (12):

[0114]

[0115] In formula (12), C ij represents the element in the i-th row and j-th column of the similarity matrix C, where γ>0 is used to control the similarity difference. Denotes the correlation matrix Z * The i-th column of Denotes the correlation matrix Z * The jth column of the similarity matrix C, element C ij Represents data point x i With x j The similarity, C ij The larger the value of the data point x i With x j The more similar.

[0116] It can be seen that the method of constructing the similarity matrix through the correlation (representation coefficient) between data makes the similarity between data points belonging to the same subspace as large as possible, and the similarity between data points in different subspaces as small as possible.

[0117] Figure 2 This is a schematic diagram of the similarity matrix constructed for the first 10 categories of data in the Olivetti Research Laboratory (ORL) face dataset in this application. Figure 2 As shown, the similarity matrix has a good block diagonal structure, and the similarity between data points of the same category is greater than the similarity between data points of different categories. The greater the brightness of the area in the diagonal block, the greater the similarity value of the area.

[0118] Figure 3 is the data point x in category 1 of the ORL face dataset of this application i A schematic diagram of the similarity between data points of category 1 and category 2, such as Figure 3 As shown, the horizontal axis represents 20 data points, the first 10 data points are 10 data points in category 1, the last 10 data points are 10 data points in category 2, and the vertical axis is x. i The similarity between the 20 data points is shown in Table 1. i Data points belonging to the same category as x i The similarity between the data points that do not belong to the same category and x i The similarity between .

[0119] Step 102: Determine the similarity density and similarity relationship data of each data point in the data set based on the similarity data; the similarity density represents the number of data points in the corresponding similarity neighborhood of each data point, and the similarity relationship data represents the size relationship between the similarities of each data point and at least two other data points.

[0120] An implementation method for determining the similarity density of each data point in the data set based on the similarity data may be to determine the similarity neighborhood of each data point based on a preset similarity threshold, and determine the similarity density of each data point in the data set based on the similarity neighborhood of each data point and the similarity between each data point in the data set. Specifically, a similarity threshold may be preset, for example, the preset similarity threshold is 0.7, and the similarity neighborhood of each data point may be determined according to the preset similarity threshold. For any data point A in the data set, a range consisting of all data points whose similarity between data point A and other data points in the data set except A is greater than the similarity threshold 0.7 is determined as the similarity neighborhood of data point A. Based on the similarity neighborhood of each data point and the similarity between each data point in the data set, the similarity density of each data point in the data set is determined. Based on the similarity neighborhood of data point A and the similarity between data point A and other data points except data point A, the number of data points of data point A in the similarity field is determined. For example, when it is determined that the number of all data points whose similarity between data point A and other data points in the data set except A is greater than the similarity threshold 0.7 is 20, the similarity density of data point A is 20.

[0121] Furthermore, we can use ε to represent the similarity threshold and W i Represents data point x i The similarity density of data point x i The number of data points in the similarity neighborhood determined by the similarity threshold ε. i The larger the value, the larger the table name data point x i The more data points there are in the corresponding similarity neighborhood, the more similar the data point x is. i The more similar data points there are, the more similar the data points are. Specifically, the similarity density W can be defined by the following formula (13): i ;

[0122]

[0123] In formula (13), when x ≥ 0, τ(x) = 1, otherwise τ(x) = 0, C ij Represents data point x i With x j similarity, where i=1, 2, 3, ..., m, j=1, 2, 3, ..., n.

[0124] Step 103: clustering the data set according to the similarity density and similarity relationship data of each data point in the data set.

[0125] As an implementation manner, clustering the data set according to the similarity density and similarity relationship data of each data point in the data set may be to determine the cluster center of the data set according to the similarity density and similarity relationship data of each data point in the data set, and cluster other data points outside the cluster center. For example, for data set X, cluster center x is determined according to the similarity density and similarity relationship number of each data point in data set X. i With x j , x i With x j belong to different data categories and have i With x j Other data outside of the classification.

[0126] In practical applications, steps 101 to 103 can be implemented using a processor in an electronic device, and the processor can be at least one of an application specific integrated circuit (ASIC), a digital signal processor (DSP), a digital signal processing device (DSPD), a programmable logic device (PLD), an FPGA, a central processing unit (CPU), a controller, a microcontroller, and a microprocessor.

[0127] It can be seen that the data processing method in the embodiment of the present invention first obtains similarity data of the data set, wherein the similarity data represents the similarity between data points in the data set, and then determines the similarity density and similarity relationship data of each data point in the data set based on the similarity data, wherein the similarity density represents the number of data points of each data point in the similarity neighborhood, and the similarity relationship data represents the size relationship between the similarities of each data point and at least two other data points; finally, the data set is clustered based on the similarity density and similarity relationship data of each data point in the data set. Since the data processing method performs clustering based on the similarity density and similarity relationship data of each data point in the data set determined based on the similarity data of the data set, and does not perform clustering based on the distance between the data, the method can be used to cluster high-dimensional data sets with relatively sparse data distribution, and the quality of the clustering results obtained is better.

[0128] In one embodiment, clustering the data set based on the similarity density and similarity relationship data of each data point in the data set may include: determining the clustering parameters of each data point based on the similarity data of the data set and the similarity relationship data of each data point in the data set; the each data point includes a first data point and a second data point, the first data point represents a data point in the data set with a non-maximum similarity density, and the second data point is a data point in the data set with a maximum similarity density; clustering the data set based on the similarity density of each data point in the data set and the clustering parameters of the each data point.

[0129] In one example, the clustering parameter ω of a data point i It can be a data point x i The similarity density is greater than the data point x iThe minimum value of the reciprocal of the similarity between the data points of the similarity density. For the first data point in the data set, the clustering parameter is as follows:

[0130]

[0131] In formula (14), C ij Represents data point x i With x j The similarity of W j Represents data point x j The similarity density, W i Represents data point x i The similarity density, clustering parameter ω i The larger the value, the greater the value of data point x. i The similarity density is greater than the data point x i The higher the similarity density, the smaller the similarity between the data points.

[0132] For the second data point in the data set, the clustering parameters can be found in formula (15):

[0133]

[0134] In one embodiment, determining the clustering parameters of each data point according to the similarity data of the data set and the similarity relationship data of each data point in the data set includes: selecting a data point with the smallest first calculated value from among the data points with a similarity density greater than that of the first data point, and using the first calculated value of the selected data point as the clustering parameter of the first data point, wherein the first calculated value is negatively correlated with the second calculated value, and the second calculated value represents the similarity between the data point with a similarity density greater than that of the first data point and the first data point;

[0135] Among the data points other than the second data point in the data set, the data point with the largest third calculated value is selected, and the selected data point is used as the clustering parameter of the second data point according to the third calculated value, the third calculated value is negatively correlated with the fourth calculated value, and the fourth calculated value represents the similarity between the data points other than the second data point in the data set and the second data point.

[0136] In one example, the first calculated value may refer to the reciprocal value of the similarity between the first data point and the data point whose similarity density is greater than the first data point. For example, the first calculated value may be Among them, C ij Represents data point x i With x jOf course, the first calculated value may also be other calculated values ​​that are negatively correlated with the calculated similarity value between the first data point and the data point whose similarity density is greater than the first data point. For example, the first calculated value may also be (C ij ) -2 , the third calculated value may refer to the reciprocal value of the similarity between the second data point and the data point whose similarity density is greater than the second data point. For example, the third calculated value may be Among them, C ij Represents data point x i With x j Of course, the third calculated value may also be other calculated values ​​that are negatively correlated with the calculated similarity value between the second data point and the data point whose similarity density is greater than the second data point. For example, the third calculated value may also be (C ij ) -2 Specifically, for data points B, C, and D in the data set whose similarity density is greater than that of the first data point A, where the calculated value corresponding to data point B is the smallest, the clustering parameter of the first data point A is the first calculated value of data point B; for data points other than the second data point E in the data set, where the third calculated value of data point H is the largest, the clustering parameter of the second data point E is the third calculated value of data point H.

[0137] In one embodiment, clustering the data set according to the similarity density and clustering parameters of each data point in the data set includes: determining the cluster center of the data set according to the clustering parameters and similarity density of each data point in the data set, and clustering other data points outside the cluster center.

[0138] In one example, the clustering center of the data set is determined based on the clustering parameters and similarity density of each data point in the data set. The method may be to determine, based on the clustering parameters and similarity density of each data point in the data set, a data point in the data set whose similarity density is greater than a first threshold and whose clustering parameters are greater than a second threshold as the clustering center of the data set; and cluster the other data points outside the clustering center in the data set based on the similarity density of each other data point outside the clustering center.

[0139] As an implementation method, for each data point x in the data set X i The corresponding similarity density W i With the clustering parameter ω i, when the similarity density W3 corresponding to a data point x3 in the data set X is greater than the first threshold, and the clustering parameter ω3 corresponding to the data point x3 is greater than the second threshold, the data point x3 is surrounded by more data points in the similarity neighborhood, and the similarity between the data point x3 and the data point with a similarity density greater than that of the data point x3 is small. Therefore, the similarity density W3 is selected. i is greater than the first threshold, and the clustering parameter ω i The data points corresponding to the second threshold are taken as cluster centers. Specifically, they can be expressed by the formula η i =W i ω i To calculate W i With ω i Comprehensive consideration value η i , take the first k largest η i The corresponding data points are used as cluster centers, where k represents the number of categories in the data set.

[0140] The clustering of other data points outside the cluster center according to the similarity density of each other data point outside the cluster center in the data set includes:

[0141] According to the similarity density of each other data point outside the cluster center, a data point set corresponding to each other data point outside the cluster center is determined among the other data points outside the cluster center in the data set, wherein the similarity density of the data points in the data point set is greater than the similarity density of each corresponding other data point;

[0142] In each of the data point sets, determining a third data point having the largest similarity density;

[0143] Determine that the category of each other data point outside the cluster center is the same as the category of the corresponding third data point.

[0144] As an implementation mode, clustering other data points outside the cluster center may be performed by arranging all other data points outside the cluster center in descending order according to similarity density, and classifying them starting from the data point with the largest similarity density. Specifically, for a data set X in which only cluster centers x3 and x5 exist, the similarity density of other data points except cluster centers x3 and x5 is determined. For example, clustering data points x1 and x2 in the data set may be performed by respectively determining a set of data points whose similarity density is greater than that of x1 and x2. Specifically, the set of data points whose similarity density is greater than that of x1 may include data point x1, x2, and x3. 11 、x 12 , x4 and x6, among which the data point x4 has the largest similarity density; the data point set with a similarity density greater than that of x2 can be data points x8, x7, x 10 and x15 , where the data point x 10 The similarity density is the largest, so it can be determined that the third data point corresponding to data point x1 is x4, and the third data point corresponding to data point x2 is x 10 , and then we can determine that data point x1 and data point x4 have the same category, and data point x2 and data point x 10 The same category.

[0145] In one implementation, it can be for the ORL face image dataset X=[x1, x2, ..., x n ]=[X1,X2,…,X k ]∈R m×n , find the optimal solution of formula (1) and determine the correlation matrix Z of data set X * , and construct the similarity matrix C through formula (12), and determine the similarity density ω of each data point in the data set according to the similarity matrix C i and clustering parameter W i , according to the similarity density ω of each data point determined i and clustering parameter W i Determine the cluster center of the data set and cluster other data points outside the cluster center.

[0146] The data processing method of the embodiment of the present invention includes the following three parameters: a balance parameter λ, a similarity threshold ε, and a γ value used to control the similarity difference when constructing a similarity matrix. Through analysis and comparison of multiple experimental results, the balance parameter λ can be set to 2. In order to determine the influence of the similarity threshold ε and the γ value used to control the similarity difference when constructing a similarity matrix on the quality of the clustering results, the first 10 categories of the Extended Yale B data set can be selected for experiments.

[0147] Figure 4 This is a schematic diagram of the change in the accuracy of the clustering results when the similarity difference control parameter γ value of the present application changes from 2 to 9, as shown in Figure 4 As shown in the figure, when the control parameter γ changes from 2 to 4, the accuracy of the clustering results increases from 57.03% to 95.46%. When the control parameter γ changes from 4 to 8, the clustering accuracy tends to be stable and remains at around 96.00%. When γ changes to 9, the accuracy of the clustering results drops to 86.41%. Therefore, in order to obtain a higher accuracy of the clustering results, the value of the similarity difference control parameter γ can be set between 4 and 8.

[0148] Figure 5 This is a schematic diagram of the change in the accuracy of the clustering results when the similarity threshold ε changes from 0.01 to 0.06 when γ=7 in this application, as shown in Figure 5As shown, when the similarity threshold ε is equal to 0.01 or 0.02, the accuracy of the clustering result is greater than 90%, which is in a higher accuracy range. When ε is greater than 0.02, the accuracy of the clustering result begins to decrease. In order to obtain higher clustering quality through the data processing algorithm, the similarity threshold can be set to 0.01.

[0149] Based on the data processing method proposed in the above-mentioned embodiment, an embodiment of the present invention proposes a data processing device.

[0150] Figure 6 The structure diagram of the data processing device of the present application is as follows: Figure 6 As shown, the device may include: an acquisition module 601, a determination module 602 and a clustering module 603, wherein:

[0151] An acquisition module 601 is used to acquire similarity data of a data set, wherein the similarity data represents the similarity between data points in the data set;

[0152] Determination module 602, used to determine the similarity density and similarity relationship data of each data point in the data set according to the similarity data; the similarity density represents the number of data points in the corresponding similarity neighborhood of each data point, and the similarity relationship data represents the magnitude relationship between the similarities of each data point and at least two other data points;

[0153] The clustering module 603 is used to cluster the data set according to the similarity density and similarity relationship data of each data point in the data set.

[0154] In one implementation, the determination module 602 is used to determine the similarity density of each data point in the data set according to the similarity data, including:

[0155] The similarity neighborhood of each data point is determined according to a preset similarity threshold, and the similarity density of each data point in the data set is determined according to the similarity neighborhood of each data point and the similarity between each data point in the data set.

[0156] In one implementation, the clustering module 603 is used to cluster the data set according to the similarity density and similarity relationship data of each data point in the data set, including:

[0157] Determine the clustering parameters of each data point according to the similarity data of the data set and the similarity relationship data of each data point in the data set; each data point includes a first data point and a second data point, the first data point represents a data point in the data set with a non-maximum similarity density, and the second data point is a data point in the data set with a maximum similarity density;

[0158] The data set is clustered according to the similarity density of each data point in the data set, the data points and the clustering parameters.

[0159] In one implementation, the clustering module 603 is used to determine the clustering parameters of each data point according to the similarity data of the data set and the similarity relationship data of each data point in the data set, including:

[0160] Selecting a data point with the smallest first calculated value from among the data points with a similarity density greater than the first data point, and using the first calculated value of the selected data point as a clustering parameter of the first data point, wherein the first calculated value is negatively correlated with the second calculated value, and the second calculated value represents the similarity between the data point with a similarity density greater than the first data point and the first data point;

[0161] Among the data points other than the second data point in the data set, the data point with the largest third calculated value is selected, and the third calculated value of the selected data point is used as the clustering parameter of the second data point. The third calculated value is negatively correlated with the fourth calculated value, and the fourth calculated value represents the similarity between the second data point and the data points other than the second data point in the data set.

[0162] In one implementation, the clustering module 603 is used to cluster the data set according to the similarity density and clustering parameters of each data point in the data set, including:

[0163] According to the clustering parameters and similarity density of each data point in the data set, determining the data point in the data set whose similarity density is greater than a first threshold and whose clustering parameter is greater than a second threshold as the clustering center of the data set;

[0164] According to the similarity density of each other data point outside the cluster center in the data set, clustering the other data points outside the cluster center. In one embodiment, the clustering module 603 is used to cluster the other data points outside the cluster center according to the similarity density of each other data point outside the cluster center in the data set, including:

[0165] According to the similarity density of each other data point outside the cluster center, a data point set corresponding to each other data point outside the cluster center is determined among the other data points outside the cluster center in the data set, wherein the similarity density of the data points in the data point set is greater than the similarity density of each corresponding other data point;

[0166] In each of the data point sets, determining a third data point having the largest similarity density;

[0167] Determine that the category of each other data point outside the cluster center is the same as the category of the corresponding third data point. In one embodiment, the acquisition module 601 is used to acquire similarity data of the data set, including:

[0168] Acquire relevant data of the data set, wherein the relevant data is a matrix;

[0169] The similarity data of the data set is determined according to the column vector of the relevant data of the data set.

[0170] In practical applications, the acquisition module 601, the determination module 602 and the clustering module 603 can be implemented using a processor in an electronic device, and the processor can be at least one of an ASIC, a DSP, a DSPD, a PLD, a FPGA, a CPU, a controller, a microcontroller, and a microprocessor.

[0171] In addition, each functional module in this embodiment can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The above integrated unit can be implemented in the form of hardware or software functional modules.

[0172] If the integrated unit is implemented in the form of a software function module and is not sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this embodiment is essentially or the part that contributes to the relevant technology or the whole or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions for a computer device (which can be a personal computer, server, or network device, etc.) or a processor to perform all or part of the steps of the method described in this embodiment. The aforementioned storage medium includes: U disk, mobile hard disk, read only memory (ROM), random access memory (RAM), disk or optical disk, etc. Various media that can store program codes.

[0173] Specifically, a computer program instruction corresponding to a data processing method in this embodiment can be stored on a storage medium such as a CD, a hard disk, or a USB flash drive. When the computer program instruction corresponding to a data processing method in the storage medium is read or executed by an electronic device, any data processing method in the aforementioned embodiments is implemented.

[0174] Based on the same technical concept as the above embodiments, see Figure 7, which shows an electronic device provided by the present invention, which may include: a memory 701 and a processor 702; wherein,

[0175] The memory 701 is used to store computer programs and data;

[0176] The processor 702 is used to execute the computer program stored in the memory to implement any one of the data processing methods in the foregoing embodiments.

[0177] In practical applications, the memory 701 may be a volatile memory, such as RAM; or a non-volatile memory, such as ROM, flash memory, hard disk drive (HDD) or solid-state drive (SSD); or a combination of the above types of memory, and provide instructions and data to the processor 702.

[0178] The processor 702 may be at least one of ASIC, DSP, DSPD, PLD, FPGA, CPU, controller, microcontroller, and microprocessor. It is understandable that for different augmented reality cloud platforms, the electronic device used to implement the processor function may also be other, which is not specifically limited in the embodiment of the present invention.

[0179] In some embodiments, the functions or modules included in the apparatus provided in the embodiments of the present invention can be used to execute the method described in the above method embodiment. The specific implementation can refer to the description of the above method embodiment. For the sake of brevity, it will not be repeated here.

[0180] The above description of the various embodiments tends to emphasize the differences between the various embodiments. The same or similar aspects can be referenced to each other. For the sake of brevity, this article will not repeat them.

[0181] The methods disclosed in the various method embodiments provided in this application can be arbitrarily combined without conflict to obtain new method embodiments.

[0182] The features disclosed in the various product embodiments provided in this application can be arbitrarily combined without conflict to obtain new product embodiments.

[0183] The features disclosed in the various method or device embodiments provided in this application can be arbitrarily combined without conflict to obtain new method embodiments or device embodiments.

[0184] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present invention, or the part that contributes to the relevant technology, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for a terminal (which can be a mobile phone, computer, server, air conditioner, or network equipment, etc.) to execute the methods described in each embodiment of the present invention.

[0185] The embodiments of the present invention are described above in conjunction with the accompanying drawings, but the present invention is not limited to the above-mentioned specific implementation modes, which are merely illustrative rather than restrictive. Under the guidance of the present invention, ordinary technicians in this field can also make many forms without departing from the scope of protection of the present invention and the claims, all of which are within the protection of the present invention.

Claims

1. A data processing method, characterized in that: The method is applied to image data and / or video data and / or text data; the method comprises: Acquire similarity data of a data set, wherein the similarity data represents similarities between data points in the data set; Determine the similarity density and similarity relationship data of each data point in the data set according to the similarity data; the similarity density represents the number of data points in the corresponding similarity neighborhood of each data point, and the similarity relationship data represents the magnitude relationship between the similarities of each data point and at least two other data points; Each data point in the data set includes a first data point and a second data point, wherein the first data point represents a data point in the data set with a non-maximum similarity density, and the second data point represents a data point in the data set with a maximum similarity density; Among the data points with a similarity density greater than the first data point, a data point with a minimum first calculated value is selected, and the first calculated value of the selected data point is used as a clustering parameter of the first data point; wherein the first calculated value is negatively correlated with the second calculated value, and the second calculated value indicates the similarity between the data point with a similarity density greater than the first data point and the first data point; Selecting a data point with the largest third calculated value from other data points in the data set except the second data point, and using the third calculated value of the selected data point as a clustering parameter of the second data point; wherein the third calculated value is negatively correlated with the fourth calculated value, and the fourth calculated value represents the similarity between other data points in the data set except the second data point and the second data point; According to the clustering parameter and similarity density of each data point in the data set, determining the data point in the data set whose similarity density is greater than a first threshold and whose clustering parameter is greater than a second threshold as the clustering center of the data set; The other data points outside the cluster center are clustered according to the similarity density of each other data point outside the cluster center in the data set.

2. The method according to claim 1, characterized in that Determining the similarity density of each data point in the data set according to the similarity data includes: Determine the similarity neighborhood of each data point according to a preset similarity threshold; The similarity density of each data point in the data set is determined according to the similarity neighborhood of each data point and the similarity between each data point in the data set.

3. The method according to claim 1, characterized in that The clustering of other data points outside the cluster center according to the similarity density of each other data point outside the cluster center in the data set includes: According to the similarity density of each other data point outside the cluster center, a data point set corresponding to each other data point outside the cluster center is determined among the other data points outside the cluster center in the data set, wherein the similarity density of the data points in the data point set is greater than the similarity density of each corresponding other data point; In each of the data point sets, determining a third data point having the largest similarity density; It is determined that the category of each other data point outside the cluster center is the same as the category of the corresponding third data point.

4. The method according to claim 1, characterized in that: The obtaining of similarity data of the data set includes: Acquire relevant data of the data set, wherein the relevant data is a matrix; The similarity data of the data set is determined according to the column vector of the relevant data of the data set.

5. A data processing device, characterized in that: The device is applied to image data and / or video data and / or text data; the device comprises: an acquisition module, a determination module and a clustering module, wherein: The acquisition module is used to acquire similarity data of the data set, wherein the similarity data represents the similarity between data points in the data set; The determination module is used to determine the similarity density and similarity relationship data of each data point in the data set according to the similarity data; the similarity density represents the number of data points in the corresponding similarity neighborhood of each data point, and the similarity relationship data represents the magnitude relationship between the similarities of each data point and at least two other data points; The clustering module is used for each data point in the data set, including a first data point and a second data point, wherein the first data point represents a data point with a non-maximum similarity density in the data set, and the second data point is a data point with the maximum similarity density in the data set; among the data points with a similarity density greater than the first data point, a data point with a minimum first calculated value is selected, and the first calculated value of the selected data point is used as a clustering parameter of the first data point; wherein the first calculated value is negatively correlated with the second calculated value, and the second calculated value represents the similarity between the data point with a similarity density greater than the first data point and the first data point; in the data set, except for the second data point, The data point with the largest third calculated value is selected from other data points in the data set, and the third calculated value of the selected data point is used as the clustering parameter of the second data point; wherein the third calculated value is negatively correlated with the fourth calculated value, and the fourth calculated value represents the similarity between the second data point and other data points in the data set except the second data point; according to the clustering parameter and similarity density of each data point in the data set, the data point in the data set whose similarity density is greater than the first threshold and whose clustering parameter is greater than the second threshold is determined as the clustering center of the data set; according to the similarity density of each other data point in the data set except the clustering center, the other data points outside the clustering center are clustered.

6. An electronic device, characterized in that: comprising a processor and a memory for storing a computer program capable of running on the processor; wherein, When the processor is used to run the computer program, it executes the data processing method according to any one of claims 1 to 4.

7. A computer storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the data processing method according to any one of claims 1 to 4 is implemented.

Citation Information

Patent Citations

  • Density peak clustering method based on local density and cluster center optimization

    CN108280472A

  • Hotspot path analysis method based on density clustering

    CN110135450A