Face recognition method, device and system based on improved density peak clustering
By improving the density peak clustering method, using support points and relative distances to construct a decision map, combining synergistic similarity and nearest neighbor information matrix for sample allocation, the problems of improper selection of clustering centers and difficult parameter selection in face recognition are solved, and higher recognition accuracy is achieved.
Patent Information
- Application Number
- CN202510254211.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-05
- Publication Date
- 2025-06-20
AI Technical Summary
The existing clustering methods have problems in the face recognition that the initial cluster center selection leads to incorrect clustering, and it is difficult to select appropriate parameter values in high overlapping cluster environments.
The improved density peak clustering method is adopted to calculate the number of support points and the relative distance, and a decision map is constructed, points with a larger product of relative distance and local density are selected as the clustering center, and samples are allocated using synergistic similarity and nearest neighbor information matrix.
Accurately identifying the clustering centers of different faces avoids the impact of density differences and chain reaction problems, and improves the accuracy of face recognition.
Smart Images

Figure CN120183012A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of face image processing, and particularly relates to a face recognition method, device and system based on improved density peak clustering. Background Technique
[0002] With the rapid development of information technology, face recognition technology has gradually become the focus of research and application. The application of clustering technology in face recognition can better manage and analyze a large amount of face data. For example, in security monitoring, by clustering the face images in a large number of surveillance videos, different images of the same person can be quickly classified, facilitating subsequent query and analysis. In intelligent photo album management, clustering technology can automatically aggregate photos of the same person together, providing users with a more convenient photo management and browsing experience. However, existing clustering methods still have some limitations in the application of face recognition.
[0003] K-means is a well-known clustering algorithm. It randomly initializes K clustering centers, repeatedly calculates the distances from data points to each center and assigns them to the corresponding clusters, and then updates the centers until convergence or the maximum number of iterations is reached to achieve the goal of dividing the data set into K clusters. However, it highly depends on the initialization of the cluster centers, and because points are always assigned to the nearest center, it may not be able to identify non-convex and nested-shaped clusters. Therefore, when clustering face images by different identities, if the initial clustering centers are not properly selected, face images of the same person may be assigned to different categories, or face images of different people may be wrongly clustered into one category. Density-based methods can effectively make up for the deficiencies of the K-means method. As a classic density-based algorithm, DBSCAN can identify clusters of any shape according to specific density-connectivity criteria. However, when facing highly overlapping clusters, it is crucial to select appropriate thresholds. A wide threshold may lead to the merging of different clusters, while a strict threshold may affect the ability to reconstruct clusters. Therefore, in the face recognition environment, due to the fact that the distribution of face images in the feature space is affected by various factors, resulting in uneven distribution densities of the feature vectors of face images, it is difficult to select appropriate parameter values, affecting the accuracy of face recognition.
[0004] Density Peak Clustering (DPC) effectively separates highly overlapping clusters by constructing a decision graph and searching for density peaks. After marking the selected cluster centers, each remaining point inherits the label of its nearest higher-density point to complete the clustering. This simple and effective implementation of locating cluster centers makes DPC one of the best-performing clustering methods, capable of clustering face images by different identities. Nevertheless, DPC still has some limitations. First, it cannot identify the correct cluster centers in clusters with large density differences, resulting in the possibility of selecting the wrong face image as a representative during face recognition. Second, when assigning non-central points, a chain reaction is likely to occur, leading to a large number of consecutive misclassifications of face images. Summary of the Invention
[0005] To solve the above problems, the present invention proposes a face recognition method, device, and system based on improved density peak clustering.
[0006] In the first aspect of the present invention, a face recognition method based on improved density peak clustering is proposed, including the following steps:
[0007] Obtain a face image dataset;
[0008] Based on the face image dataset, calculate the distance matrix and k-nearest neighbor density;
[0009] Based on the distance matrix and k-nearest neighbor density, calculate the number of support points and relative distance;
[0010] Based on the k-nearest neighbor density and the number of support points, calculate the local density;
[0011] Construct a decision graph of the face image dataset based on the relative distance and local density, and select the first several face feature vector points with a larger product of relative distance and local density as cluster centers;
[0012] Assign unassigned samples greater than the average collaborative similarity to the corresponding clusters; the collaborative similarity is determined by the Euclidean distance between the assigned samples and their k-nearest neighbor samples; the assigned samples are the face feature vector points that have been added to the corresponding cluster centers; the unassigned samples are the face feature vector points that have not been added to the corresponding cluster centers;
[0013] Assign the maximum value of the nearest neighbor information matrix and its corresponding unassigned samples to the corresponding clusters; the nearest neighbor information matrix is determined by the proximity between the unassigned samples and their assigned samples;
[0014] After adding all unassigned samples to the corresponding clusters, obtain the final recognition result of the face image dataset.
[0015] In the second aspect of the present invention, the present invention also provides a face recognition device based on improved density peak clustering, including a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete mutual communication through the communication bus;
[0016] The memory is used to store computer programs;
[0017] The processor is used to implement the method steps described in the first aspect of the present invention when executing the programs stored on the memory.
[0018] In the third aspect of the present invention, the present invention also provides a face recognition system based on improved density peak clustering, including a plurality of face recognition devices based on improved density peak clustering as described in the second aspect of the present invention. Each face recognition device is connected through a network, and the face feature libraries of each face recognition device form the total face feature library of the face recognition system.
[0019] The advantages and beneficial effects of the present invention are as follows:
[0020] 1. The present invention uses support points to recalculate the local density of samples. Support points reflect the dispersion degree of the area around the face feature vector points, so it can eliminate the influence of density differences on density measurement, accurately select the clustering centers, that is, accurately identify the clustering centers of different faces.
[0021] 2. The present invention adopts a new sample allocation strategy. First, according to the proposed collaborative similarity, the core points that are not easily misallocated are allocated, and then the remaining points are allocated by constructing a nearest neighbor information matrix according to the proximity. In addition, the spatial distance between face feature vector points and the connectivity between samples and surrounding points are fully considered in the proposed collaborative similarity; and the proximity is proposed to allocate the remaining points by using the nearest neighbor information, effectively using the information of the allocated points and fully considering the tightness of the association between the allocated nearest neighbors and the unallocated points, ensuring the accuracy of the allocation. This allocation strategy can solve the chain reaction problem when samples are misallocated, avoid the influence of a misallocated face image on the allocation of its surrounding similar feature face images in face image clustering, and improve the accuracy of face recognition. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 is the face recognition flowchart based on improved density peak clustering according to an embodiment of the present invention;
[0023] Figure 2 is the structural schematic diagram of the face recognition device based on improved density peak clustering according to an embodiment of the present invention;
[0024] Figure 3 is the schematic diagram of the recognition result using traditional density peak clustering;
[0025] Figure 4 It is the decision diagram for improving density peak clustering in the embodiments of the present invention;
[0026] Figure 5 It is the schematic diagram of the recognition result for improving density peak clustering in the embodiments of the present invention. Specific implementation manners
[0027] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0028] It should be noted that in the technical solutions of the present invention, operations such as the acquisition, storage, use, processing, transmission, provision, and disclosure of images containing human faces all comply with the provisions of relevant laws and regulations and do not violate public order and good customs. For example, the above operations are all executed under authorization.
[0029] A face recognition method based on improved density peak clustering provided by the embodiments of the present invention can be applied to an electronic device, which can be a terminal device or a server. The terminal device can include: mobile phones, tablet computers, etc. The present invention does not limit the specific form of the electronic device. A face recognition method based on improved density peak clustering provided by the embodiments of the present invention can be applied to any scenario with face recognition, such as: intelligent traffic monitoring, attendance scenarios, etc.; and the present invention does not limit the acquisition duration range, acquisition location, etc. of the images to be face recognized. For example: face recognition of multiple images collected at different locations on the same day, face recognition of multiple images collected for multiple days, etc.
[0030] In addition, the execution subject of a face recognition method based on improved density peak clustering provided by the embodiments of the present invention can be a face recognition device based on improved density peak clustering. Exemplarily, the face recognition device based on improved density peak clustering can be a functional software running on a terminal device, such as: functional software for face recognition. At this time, the face recognition device can perform face clustering and recognition on the input multiple images; of course, the face recognition device can also be a plug-in in an existing client, such as: a plug-in in a client for managing images containing human faces; in addition, the face recognition device can also be a functional module in a server-side program corresponding to a face recognition client. At this time, the face recognition client can upload multiple images to be face recognized to the face recognition device.
[0031] In the embodiment of the present invention, first, the k-nearest neighbor density of a sample is calculated using the k-nearest neighbor matrix of face feature vector points, then the local density of the sample is calculated based on the k-nearest neighbor density, a decision graph is constructed according to the local density and the relative distance, and samples with relatively large local density and relative distance are selected as clustering centers. For non-central points, they are first sorted from large to small according to the local density, then preliminarily assigned using the proposed collaborative similarity, and finally the remaining points are assigned using the k-nearest neighbor information of the data points to obtain the final face clustering result.
[0032] The face recognition method based on improved density peak clustering provided by the embodiment of the present invention, as Figure 1 shown, may include the following steps:
[0033] 101. Obtain a face image dataset;
[0034] In the embodiment of the present invention, the face image dataset may be multiple face images obtained by different acquisition devices at different times and different locations. The face images are composed of multiple face feature vector points, and the multiple face images may cover various age groups, genders, skin colors, and expression features. Some face images are taken in a bright indoor environment, clearly showing the detailed features of the face, while some face images are obtained in a dim outdoor scene and may be affected by factors such as light, resulting in a decline in image quality; precisely through the face recognition method of the present invention, different face images of the same object with different qualities can be classified, facilitating subsequent query and analysis.
[0035] Exemplarily, the face image may be a grayscale image of size m×n, with each pixel serving as a feature point. After reading in L face images, a feature matrix P i(m×n) is obtained, where i = 1, 2,..., L; then the feature matrix P i(m×n) is converted into a feature vector x ij , and x ij represents the value of the point x i on the jth feature in the original face image dataset.
[0036] 102. Based on the face image dataset, calculate a distance matrix and k-nearest neighbor density;
[0037] In the embodiment of the present invention, the feature vector set is input as the face image dataset S to be clustered, a distance matrix D is calculated, and in combination with a given positive integer k, the k-nearest neighbor matrix and snn matrix of each sample are obtained.
[0038] Before calculating the distance matrix for the face image dataset S to be clustered, the input variables need to be normalized first, and the normalization formula includes:
[0039]
[0040] Among them, max(x j ) and min(x j ) represents the maximum and minimum values of the corresponding columns in the original face image dataset, x ij Represents the jth feature point x in the original face image dataset i The value of x ij ' represents the jth feature point x in the normalized original face image dataset i This normalization method can stabilize facial images of different sizes on the same scale, and prevent certain features from dominating the subsequent face clustering recognition results due to their large value range. This embodiment calculates the Euclidean distance of each feature point in the normalized original face image data set, and uses the Euclidean distance calculation to obtain its distance matrix. This distance matrix can be used for subsequent clustering analysis to quickly determine the same or similar faces.
[0041] In some embodiments, the embodiments of the present invention can also calculate the Manhattan distance, Minkowski distance, Hamming distance, cosine distance, etc. for each feature point in the normalized original face image data set. Then use these distance calculations to obtain the corresponding distance matrix. Similarly, these distance matrices can also be used for subsequent clustering analysis to quickly determine the same or similar faces.
[0042] In an embodiment of the present invention, the calculation method of the k nearest neighbor density includes:
[0043] According to the Euclidean distance between the face feature vector point to be calculated and its k nearest neighbor face feature vector points, the normalized distance from the k nearest neighbor face feature vector point to the face feature vector point to be calculated is calculated;
[0044] According to the normalized distance from the k nearest neighbor facial feature vector point to the facial feature vector point to be calculated, the k nearest neighbor density of the facial feature vector point to be calculated is calculated. Among them, the k value in the k nearest neighbor set can be selected according to the specific application scenario or through cross-validation. For each facial feature vector point to be calculated, the top k nearest neighbor facial feature vector points can be determined by the Euclidean distance between it and other facial feature vector points; these k nearest neighbor facial feature vector points will be used as the k nearest neighbor set of the facial feature vector point to be calculated, and the k nearest neighbor density of the facial feature vector point to be calculated can be obtained by calculating the normalized distance from each facial feature vector point in the k nearest neighbor set to the facial feature vector point to be calculated.
[0045] In this embodiment, by considering the distance information of neighbors to calculate the density, the local density distribution around the face feature vector point to be calculated can be reflected. In this way, the feature distribution around the face feature vector point to be calculated can be described more comprehensively. For example, the feature distribution of facial features. The feature distribution of facial features presents relatively stable and unique rules. Based on this relatively stable and unique rule, the facial feature information can be accurately captured, and then a more accurate face clustering center can be selected.
[0046] Exemplarily, the calculation formula of the k-nearest neighbor density is expressed as:
[0047]
[0048] where density(x i ) represents the k-nearest neighbor density of the face feature vector point x i , d ij represents the Euclidean distance between the face feature vector point x i and the face feature vector point x j , knn(x i ) represents the k-nearest neighbor set of the face feature vector point x i , d max represents the maximum value of the Euclidean distance between face feature vector points, and k represents the number of neighbors.
[0049] 103. Based on the distance matrix and the k-nearest neighbor density, the number of support points and the relative distance are calculated;
[0050] In the embodiment of the present invention, the calculation method of the relative distance includes:
[0051] When there is a target face feature vector point with a k-nearest neighbor density greater than that of the face feature vector point to be calculated, the relative distance is the minimum value of the Euclidean distance between the face feature vector point to be calculated and the target face feature vector point with a higher density;
[0052] When the face feature vector point to be calculated is the sample with the largest k-nearest neighbor density in the face feature vector point dataset, its relative distance is the maximum value of the relative distance between the face feature vector point to be calculated and the target face feature vector point in the face feature vector point dataset.
[0053] Exemplarily, the calculation formula of the relative distance is expressed as:
[0054]
[0055] where δ i represents the relative distance of the i-th face feature vector point, density(x i ) represents the k-nearest neighbor density of the i-th face feature vector point, dij Represents x i and x j The Euclidean distance between them. When there is a face feature vector point x j with a higher density than the face feature vector point x i The relative distance is the minimum value of the distance between the face feature vector point x i and the point with a higher density. When the face feature vector point x i is the face feature vector point with the highest density in the dataset, its relative distance is the maximum value of the relative distances of other face feature vector points in the dataset.
[0056] In an embodiment of the present invention, the calculation method of the number of support points includes:
[0057] If the k-nearest neighbor density of the face feature vector point to be calculated is less than or equal to the k-nearest neighbor density of the target face feature vector point, then the minimum value of the Euclidean distance between the face feature vector point to be calculated and the target face feature vector point is used as the relative distance between the face feature vector point to be calculated and the target face feature vector point;
[0058] If the k-nearest neighbor density of the face feature vector point to be calculated is equal to the k-nearest neighbor density of the face feature vector point with the highest density, then the Euclidean distance between the face feature vector point to be calculated and the other face feature vector points is used as the relative distance between the face feature vector point to be calculated and the target face feature vector point;
[0059] If the relative distance between the face feature vector point to be calculated and the target face feature vector point is equal to its Euclidean distance, then the target face feature vector point is used as the support point of the face feature vector point to be calculated.
[0060] Exemplarily, the calculation formula of the number of support points is expressed as:
[0061]
[0062] Among them, Represents the support degree between object x i and x j x i represents the i-th object in the original face image dataset S, S = {x1, x2,..., x n}, n represents the total number of objects in the original face image dataset, d ij represents the Euclidean distance between object x i and x j sd(x i ) represents the support degree of the i-th object, which can reflect the dispersion degree of the area around the i-th object, and δ i represents the relative distance of the i-th object.
[0063] 104. Calculate the local density based on the k-nearest neighbor density and the number of support points.
[0064] In the traditional technology, the local density is determined by the Euclidean distance from the face feature vector point to be calculated to all other points and the corresponding truncation distance. However, this method cannot identify the correct clustering center in clusters with large density differences, resulting in the possibility of selecting the wrong face image as a representative during face recognition. In the embodiments of the present invention, the local density is determined by the sum of the k-nearest neighbor density of each face feature vector point and the number of support points of each face feature vector point.
[0065] Exemplarily, the calculation formula of the local density is expressed as:
[0066] ρ(x i ) = sd(x i ) + density(x i ) (5)
[0067] Where x i represents the i-th object in the face image dataset S, S = {x1, x2,..., x n}; n represents the total number of objects in the face image dataset; d ij represents the Euclidean distance between object x i and x j ; sd(x i ) represents the support degree of the i-th object; δ i represents the relative distance of the object; ρ(x i ) represents the local density of the object.
[0068] It can be understood that the embodiments of the present invention use support points to recalculate the local density of the samples. Since the support points reflect the dispersion degree of the area around the face feature vector point, the present invention can eliminate the influence of density differences on density measurement and accurately select the clustering center, that is, accurately identify the clustering centers of different faces.
[0069] 105. Construct a decision graph of the face image dataset based on the relative distance and the local density, and select the first several face feature vector points with a larger product of the relative distance and the local density as the clustering centers.
[0070] In the embodiments of the present invention, a decision graph is constructed based on two features, namely relative distance and local density; the relative distance reflects the relative distance relationship in spatial position between face feature vector points, and the local density reflects the degree of data density around the face feature vector points. The product of the relative distance and the local density combines the two features to form a comprehensive evaluation index. Samples with a larger product are relatively closer in spatial position and have a higher local density, so they are more representative clustering centers. The decision graph constructed by this comprehensive feature can better capture the internal relationship between face feature vector points and improve the accuracy and effectiveness of clustering. Among them, the clustering centers can be expressed as C = {c1, c2,..., c m}, that is, the clustering centers of m different faces. The selection of such clustering centers can help extract the key features of face data and reduce the dimension and complexity of the data.
[0071] 106. Assign unassigned samples with a collaborative similarity greater than the average collaborative similarity to the corresponding clusters; the collaborative similarity is determined by the Euclidean distance between the assigned samples and their k-nearest neighbor samples; the assigned samples are the face feature vector points that have been added to the corresponding clustering centers; the unassigned samples are the face feature vector points that have not been added to the corresponding clustering centers; in the embodiments of the present invention, the calculation method of the collaborative similarity includes:
[0072] Calculate the attraction coefficient according to the Euclidean distance between the face feature vector point to be calculated and the target face feature vector point and the maximum value of the Euclidean distances between all face feature vector points;
[0073] Determine a set of neighborhood points that meet the preset distance according to the Euclidean distance between the face feature vector point to be calculated and other face feature vector points;
[0074] Calculate the number of shared neighbors between the face feature vector point to be calculated and the target face feature vector point according to the intersection of the set of neighborhood points of the face feature vector point to be calculated and the target face feature vector point;
[0075] Calculate the collaborative similarity between the face feature vector point to be calculated and the target face feature vector point according to the product of the attraction coefficient and the number of shared neighbors between the face feature vector point to be calculated and the target face feature vector point.
[0076] Exemplarily, the calculation formula of the collaborative similarity is expressed as:
[0077]
[0078] Where represents the attraction coefficient, d(x i , x j ) represents x i and x jthe Euclidean distance between represents the maximum value of the distances between all points in the dataset X, |snn(x i , x j )| is the number of shared nearest neighbors between x i and x j . The left part of the collaborative similarity measures the spatial distance between samples through the Euclidean distance, and the right part measures the connectivity between two points through the number of shared neighbors. The closer the distance between two points, the greater the attraction coefficient, and the more the number of shared nearest neighbors, which will have a positive effect on establishing the relationship between the two points.
[0079] In the embodiments of the present invention, the core points that are not easily misallocated are assigned according to the collaborative similarity. The spatial distance between samples and the connectivity between samples and surrounding points are fully considered in the collaborative similarity; the information of the already assigned points is effectively utilized to ensure the accuracy of the assignment. This assignment strategy can solve the chain reaction problem when samples are misallocated, avoid the influence of a misassigned face image on the assignment of surrounding similar feature face images in face image clustering, and improve the accuracy of face recognition.
[0080] In some embodiments of the present invention, a queue Q can be created, and its initial value is m clustering centers. The elements in the queue Q are accessed in sequence and their k nearest neighbors are traversed. If the collaborative similarity between the elements in the k nearest neighbors and the already assigned samples is greater than the average collaborative similarity between the already assigned samples and their k nearest neighbors, then assign the same label to this nearest neighbor as the element in Q and add this nearest neighbor to the queue Q. Repeat this step until the queue Q is empty. In this embodiment, the unassigned samples are sequentially divided into the corresponding clusters through the queue, and the core points that are not easily misallocated are quickly assigned, ensuring the accuracy of the assignment.
[0081] 107. Assign the maximum value of the nearest neighbor information matrix and its corresponding unassigned sample to the corresponding cluster; the nearest neighbor information matrix is determined by the proximity between the unassigned sample and its already assigned sample;
[0082] In the embodiments of the present invention, the calculation method of the proximity includes:
[0083] According to the Euclidean distance between the unassigned face feature vector point and its already assigned sample, calculate the first parameter between the unassigned face feature vector point and its already assigned sample;
[0084] According to the sum of the Euclidean distances between the unassigned face feature vector point and the k nearest neighbor samples of its already assigned cluster, calculate the second parameter between the unassigned face feature vector point and its already assigned sample;
[0085] According to the ratio of the first parameter to the second parameter, calculate the proximity between the unassigned face feature vector point and its already assigned sample.
[0086] In the embodiment of the present invention, the calculation formula of the proximity is expressed as:
[0087]
[0088] where prox(x i , x j ) represents the proximity between x i and x j , d(x i , x j ) represents the Euclidean distance between x i and x j , and knn(x i ) represents the set of k nearest neighbors of the sample x i .
[0089] In the embodiment of the present invention, the proximity between the unassigned sample and its assigned k nearest neighbors is calculated, and the proximity is used as an element of the matrix to construct the nearest neighbor information matrix M, and the remaining points are assigned using the nearest neighbor information. The row vector of the matrix represents the unassigned sample i, and the column vector represents the cluster j. After creating the matrix, find the maximum value of the matrix, assign the corresponding sample i to the corresponding cluster j, and update the nearest neighbor information matrix.
[0090] 108. After adding all unassigned samples to the corresponding clusters, the final recognition result of the face image dataset is obtained.
[0091] Repeat the above steps until all samples are assigned, and obtain the final face clustering result CL(x), which represents the category of the face. If CL(x1) = CL(x2), it means that the two face images represented by x1 and x2 are of the same person.
[0092] The face recognition method of the embodiment of the present invention can correctly identify the face clustering center in the face image dataset with large inter-cluster density differences, and can process clusters of any shape. Compared with the traditional face recognition method based on the density peak clustering algorithm, it eliminates the influence of density differences and avoids the chain reaction problem in the process of assigning non-central points, improving the accuracy of face recognition.
[0093] For the above-mentioned face recognition device based on improved density peak clustering, the embodiment of the present invention also provides a face recognition device based on improved density peak clustering, as Figure 2 shown, including a processor 201, a communication interface 202, a memory 203, and a communication bus 204. Among them, the processor 201, the communication interface 202, and the memory 203 communicate with each other through the communication bus 204,
[0094] The memory 203 is used to store a computer program;
[0095] The processor 201, when executing the program stored in the memory 203, implements any face recognition method based on improved density peak clustering.
[0096] It should be noted that the above communication bus 204 can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus 204 can be divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity of representation, only a thick line is used in the figure, but it does not mean that there is only one bus or one type of bus.
[0097] The communication interface is used for communication between the above face recognition device and other devices.
[0098] The memory can include a Random Access Memory (RAM), and can also include a Non-Volatile Memory (NVM), such as at least one disk memory. Optionally, the memory can also be at least one storage device located far from the aforementioned processor.
[0099] The above-mentioned processor can be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it can also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0100] The DPC clustering result of the traditional Density Peaks Clustering (DPC) on a certain public dataset is as Figure 3 shown. It can be seen that the traditional DPC wrongly selects the points in the dense cluster as the clustering centers, while ignoring the clustering centers of the sparse clusters. This leads to the possibility of selecting the wrong face image as a representative during face recognition, and when allocating non-central points, it is easy to cause a chain reaction, resulting in a large number of consecutive wrong clusterings of face images.
[0101] Using the improved density peak clustering face recognition method of the present invention, when the parameter k is taken as 7 on a certain public dataset in the present invention, the obtained decision graph is Figure 4 as shown, where the pentagrams are the two selected clustering centers. Then, using the collaborative similarity, the core samples that meet the conditions are first assigned, and then the proximity is used to construct the nearest neighbor information matrix to assign the remaining samples. The clustering result is shown in Figure 5 .
[0102] Compared with the traditional technology, the present invention can correctly identify the face clustering centers in the dataset with large inter-cluster density differences, and can handle clusters of any shape, that is, it will select more correct face images as representatives. When assigning non-central points, the chain reaction is also avoided, so that the face images are clustered in the correct direction to identify the face images of the same object from a large number of face images.
[0103] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing the relevant hardware through a program, and the program can be stored in a computer-readable storage medium. The storage medium can include: ROM, RAM, disk, optical disc, etc.
[0104] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principle and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A face recognition method based on improved density peak clustering, characterized in that: The following steps are involved: Get a face image dataset; Based on the face image dataset, the distance matrix and k nearest neighbor density are calculated; Based on the distance matrix and k nearest neighbor density, the number of support points and relative distance are calculated; Based on the k nearest neighbor density and the number of support points, the local density is calculated; A decision graph of the face image dataset is constructed based on relative distance and local density, and the first several face feature vector points with the largest product of relative distance and local density are selected as cluster centers; Assigning unassigned samples with a co-similarity greater than the average to corresponding clusters; the co-similarity is determined by the Euclidean distance between the assigned samples and the k nearest neighbors and the number of shared neighbors; the assigned samples are the facial feature vector points that have been added to the corresponding cluster center; The unassigned samples are face feature vector points that are not added to the corresponding cluster center; Assign the maximum value of the neighbor information matrix and its corresponding unassigned sample to the corresponding cluster; the neighbor information matrix is determined by the proximity between the unassigned sample and its assigned sample; After adding all unassigned samples to the corresponding clusters, the final recognition results of the face image dataset are obtained.
2. A face recognition method based on improved density peak clustering according to claim 1, characterized in that: The calculation method of the k nearest neighbor density includes: According to the Euclidean distance between the face feature vector point to be calculated and its k nearest neighbor face feature vector points, the normalized distance from the k nearest neighbor face feature vector point to the face feature vector point to be calculated is calculated; According to the normalized distance between the k nearest neighbor facial feature vector point and the facial feature vector point to be calculated, the k nearest neighbor density of the facial feature vector point to be calculated is calculated.
3. A face recognition method based on improved density peak clustering according to claim 2, characterized in that: The calculation formula of the k nearest neighbor density is expressed as: Among them, density(x i ) represents the facial feature vector point x i The k nearest neighbor density, d ij Represents the face feature vector point xi and the face feature vector point x j The Euclidean distance, knn(x i ) represents the k nearest neighbor set of the facial feature vector point xi, d max It represents the maximum value of the Euclidean distance between facial feature vector points, and k represents the number of nearest neighbors.
4. The face recognition method based on improved density peak clustering according to claim 1, characterized in that: The calculation method of the number of support points includes: If the k-nearest-neighbor density of the face feature vector point to be calculated is less than or equal to the k-nearest-neighbor density of the target face feature vector point, the minimum value of the Euclidean distance between the face feature vector point to be calculated and the target face feature vector point is taken as the relative distance between the face feature vector point to be calculated and the target face feature vector point; If the k-nearest-neighbor density of the face feature vector point to be calculated is equal to the k-nearest-neighbor density of the largest face feature vector point, the Euclidean distance between the face feature vector point to be calculated and the other face feature vector points is used as the relative distance between the face feature vector point to be calculated and the target face feature vector point; If the relative distance between the face feature vector point to be calculated and the target face feature vector point is equal to their Euclidean distance, the target face feature vector point is used as the support point of the face feature vector point to be calculated.
5. The face recognition method based on improved density peak clustering according to claim 1, characterized in that: The relative distance is calculated by: When there is a target face feature vector point with a higher k-nearest neighbor density than the face feature vector point to be calculated, the relative distance is the minimum value of the Euclidean distance between the face feature vector point to be calculated and the target face feature vector point with a higher density; When the face feature vector point to be calculated is the sample with the largest k nearest neighbor density in the face feature vector point data set, its relative distance is the maximum value of the relative distance between the face feature vector point to be calculated and the target face feature vector point in the face feature vector point data set.
6. A face recognition method based on improved density peak clustering according to any one of claims 1 to 5, characterized in that: The local density is determined by the sum of the k-nearest neighbor density of each face feature vector point and the number of supporting points of each face feature vector point.
7. The face recognition method based on improved density peak clustering according to claim 1, characterized in that: The calculation method of the collaborative similarity includes: The attraction coefficient is calculated based on the Euclidean distance between the face feature vector point to be calculated and the target face feature vector point and the maximum value of the Euclidean distance between the face feature vector point to be calculated and all the face feature vector points; Determine a set of neighborhood points that meet a preset distance based on the Euclidean distance between the face feature vector point to be calculated and other face feature vector points; According to the intersection of the neighborhood point set of the face feature vector point to be calculated and the target face feature vector point, the number of shared neighbors between the face feature vector point to be calculated and the target face feature vector point is calculated; The collaborative similarity between the face feature vector point to be calculated and the target face feature vector point is calculated according to the product of the attraction coefficient and the number of shared neighbors between the face feature vector point to be calculated and the target face feature vector point.
8. The face recognition method based on improved density peak clustering according to claim 1, characterized in that: The calculation method of the proximity includes: Calculate the first parameter between the unassigned face feature vector point and its assigned sample according to the Euclidean distance between the unassigned face feature vector point and its assigned sample; The second parameter between the unassigned face feature vector point and its assigned samples is calculated according to the sum of the Euclidean distances between the unassigned face feature vector point and its assigned k-nearest neighbor samples of the cluster; According to the ratio of the first parameter to the second parameter, the proximity between the unassigned face feature vector point and its assigned sample is calculated.
9. A face recognition device based on improved density peak clustering, characterized in that: It includes a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other through the communication bus; Memory, used to store computer programs; A processor, for implementing the method steps described in any one of claims 1 to 8 when executing a program stored in a memory.
10. A face recognition system based on improved density peak clustering, characterized in that: It comprises a plurality of face recognition devices based on improved density peak clustering as described in any one of claims 9, each face recognition device is connected via a network, and the face feature libraries of each face recognition device form a total face feature library of the face recognition system.