A Massive Face Comparison Acceleration Method Based on Deep Learning

By adding branches of the fully connected layer to the face feature extraction model of the convolutional neural network, outputting two-dimensional feature values ​​and dividing areas, the problems of slow comparison speed and resource consumption caused by large data volume of face database are solved, and efficient face feature comparison is achieved.

CN115249376BActive Publication Date: 2025-06-27ZHEJIANG UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210909516.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-29
Publication Date
2025-06-27
Estimated Expiration
2042-07-29

AI Technical Summary

Technical Problem

As the data volume of face base database is too large, traditional face feature value comparison methods lead to slower comparison speed, high resource consumption, and may even cause memory overflow and system crash.

Method used

A massive face comparison acceleration method based on deep learning is proposed. By adding branches of a fully connected layer to the face feature extraction model of a convolutional neural network, two-dimensional feature values ​​are output, and regions are divided according to the two-dimensional feature values, reducing the range of 512-dimensional feature values.

Benefits of technology

It effectively improves the speed of facial feature comparison, avoids excessive resource consumption and memory overflow problems, and does not affect recognition accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115249376B_ABST
    Figure CN115249376B_ABST
Patent Text Reader

Abstract

A method for accelerating the comparison of massive faces based on deep learning, which adds a new fully connected layer branch to the original output face feature extraction network, so that face data can simultaneously output two-dimensional feature values ​​through the feature extraction network and be visualized in a two-dimensional plane; select a suitable center point, and divide N features of the base library into m areas from the center point according to the angle; when comparing the feature values ​​of the face to be tested, first determine the feature area where it is located through the two-dimensional feature value, and then compare it one by one with the multi-dimensional feature value in the area, avoiding comparing the multi-dimensional face feature value with all the face feature values ​​in the base library, thereby achieving acceleration, and realizing the acceleration of face feature value comparison of massive images.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and particularly to a method for accelerating massive face comparison based on deep learning. Background Art

[0002] Currently, face recognition technology is mainly applied to the situation where the number of face databases is not large. The face image data to be recognized is input into a face feature extraction model based on a convolutional neural network to extract the feature values of the face, and then all registered data in the database is compared in full, that is, the similarity between two feature values is compared (the cosine distance of the feature vector is calculated). If the similarity is the highest and exceeds the set threshold, it can be regarded as the same person. As the technology becomes more mature, face recognition technology is applied to more and more fields, such as community management, public security management, etc. On some larger platforms, the face database will reach a massive level (hundreds of thousands to millions). If the traditional method of comparing face feature values is still used, it will inevitably cause high latency and large resource consumption, and even cause the problem of memory overflow and crash the system. Therefore, proposing a method to improve the speed of face feature value comparison is the key to solving the problem of too large face database data at present. Summary of the Invention

[0003] In order to overcome the deficiencies of slow comparison speed and large resource consumption caused by too large face database at present, the present invention proposes a method for accelerating massive face comparison based on deep learning, which can well handle the problem of too slow face feature value comparison speed and too large resource consumption when the database data volume is too large.

[0004] The technical solution adopted by the present invention to solve its technical problems is:

[0005] A method for accelerating massive face comparison based on deep learning, the method comprising the following steps:

[0006] 1) Modify the structure of the face feature extraction model based on the convolutional neural network: Add a new fully connected layer branch to the face feature extraction network originally with an output of 1×512, increase the output of 1×2. Each face data can obtain a 1×512 face feature value and a 1×2 two-dimensional feature value through the feature extraction network. The 512-dimensional face feature value is still used for face feature comparison later, and the two-dimensional feature value is visualized in the plane for subsequent feature division regions;

[0007] 2) Model training, design the loss function as shown in formula (1):

[0008] Loss_total = w1×Loss_softmax + w2×Loss_arcface (1)

[0009] where w1 and w2 are the weights of the loss functions Loss_softmax and Loss_arcface respectively;

[0010] where Loss_softmax is as shown in Equation (2):

[0011]

[0012] where N and n are the batch size and class number respectively, x i ∈R d , d is the feature dimension, set to 512, representing the depth feature of the i-th sample, belonging to class y i ; W j ∈R d represents the j-th column of the weight W ∈ R d×n , b j ∈R n is the bias term;

[0013] where Loss_arcface is as shown in Equation (3):

[0014]

[0015] where s represents a hypersphere with radius s, θ is the angle between the weight W and the feature x. m is an additional angular margin penalty added between x i and W yi ;

[0016] where Loss_arcface is used to extract multi-dimensional features, and Loss_softmax is used to extract two-dimensional features. During the model training process, the extraction of multi-dimensional face features is the main body of the model. Therefore, it is designed that before the i-th training iteration, w1 is 0 and w2 is 1, that is, Loss_total = Loss_arcface, and the weight parameters of the convolutional neural network and the multi-dimensional output fully connected layer are trained;

[0017] When the face feature extraction training reaches stability, take w2 > w1 > 0, and at this time, train the weight parameters of the two-dimensional output fully connected layer;

[0018] 3) Divide the region according to the two-dimensional plane feature distribution. Set the number of different registered faces in the gallery as N, select the center point, and divide the N features registered in the gallery into m regions according to the angles radiating from the center point, and save the region number where the feature is located together with the 512-dimensional feature value in the gallery for subsequent comparison processes;

[0019] 4) Comparison process: Input the face to be tested into the trained model to obtain a two-dimensional feature vector and a 512-dimensional feature vector. First, determine the eigenvalue region where the face to be tested is located according to the two-dimensional eigenvalue, and then retrieve the features in the same region in the database according to the region number to narrow the range of comparison of the 512-dimensional eigenvalues. Then, calculate the similarity of each 512-dimensional eigenvalue included in this region one by one;

[0020] Among them, the calculation of similarity: Calculate the cosine distance between the feature to be detected and the features in the database according to the 512-dimensional feature vector, as shown in Equation (4):

[0021]

[0022] Let the two 512-dimensional vectors to be compared be (x1, x2,......, x 512 ) and (y1, y2,......, y 512 );

[0023] 5) Result output and decision: Let the comparison threshold be threshold. Screen the face data with a cosine value greater than threshold in the selected eigenvalue region, sort them from largest to smallest, output the face data with the largest cosine value calculation result, and the recognition result is the same person, and output the id of this data.

[0024] Furthermore, in step 5), if no face data with a cosine value greater than threshold is found in the selected eigenvalue region, according to the actual requirements for the single recognition rate and recognition speed, the following decision methods can be selected: (1) Continue to compare all the features in two adjacent regions. If no result is still obtained, continue with this decision method; (2) Output recognition failure, that is, there is no such face data in the database.

[0025] When no result exceeding the threshold threshold is obtained in the region, choosing to perform method (1) more times will increase the face comparison time, but the single recognition rate will increase. On the contrary, performing method (1) less and choosing method (2) will accelerate the single comparison time, but the corresponding recognition rate will decrease; The specific number of times of performing method (1) needs to be obtained through some tests according to the actual requirements for the single recognition rate and recognition speed to get the optimal result.

[0026] The technical concept of the present invention is as follows: aiming at the problem of excessive resource consumption in one-to-one eigenvalue comparison due to the large size of the face data base, a method for accelerating massive face comparison based on deep learning is proposed. By adding a branch of a fully connected layer to the original face feature extraction network that outputs 1×512, the face data can output a 1×2 two-dimensional eigenvalue through the feature extraction network and be visualized in the two-dimensional plane. According to the characteristics of Softmax Loss, the two-dimensional features of N face data in the base are radially distributed in the two-dimensional plane. The center point is selected and the N features are divided into m regions according to the angle. When comparing the eigenvalues of the face to be tested, first determine the eigenvalue region where it is located through the two-dimensional eigenvalue, and then compare the multi-dimensional eigenvalues in the region one by one, avoiding comparing the multi-dimensional face eigenvalues with all the face eigenvalues in the base, realizing the acceleration from 1:N comparison to Acceleration is achieved, and the acceleration of face eigenvalue comparison for a large number of pictures is realized.

[0027] The beneficial effects of the present invention are mainly manifested in: as the number of features N in the face data base increases over time, the problem of excessive resource consumption in 1:N eigenvalue comparison caused by the large size of the base data is solved. Without affecting the recognition accuracy, the speed of face feature comparison can be effectively improved, and problems such as memory overflow are avoided. Description of the Drawings

[0028] Figure 1 It is a schematic diagram of a face feature extraction network.

[0029] Figure 2 It is a flowchart of face feature comparison.

[0030] Figure 3 It is a structural diagram of face feature comparison. Specific Embodiments

[0031] The method of the present invention will be further described in detail below with reference to the drawings.

[0032] Refer to Figures 1 to 3 , a method for accelerating massive face comparison based on deep learning, the method includes the following steps:

[0033] 1) Modify the structure of the face feature extraction model based on the convolutional neural network: Add a new branch of the fully connected layer to the original face feature extraction network with an output of 1×512 to increase the output of 1×2. In this way, each face data can obtain a 1×512 face eigenvalue and a 1×2 two-dimensional eigenvalue through the feature extraction network. The 512-dimensional face eigenvalue is still used for feature comparison, and the two-dimensional eigenvalue is visualized in the plane for subsequent feature block division;

[0034] 2) Model training, design the loss function as shown in Equation (1):

[0035] Loss_total = w1 × Loss_softmax + w2 × Loss_arcface (1)

[0036] Where w1 and w2 are the weights of the loss functions Loss_softmax and Loss_arcface respectively,

[0037] Where Loss_softmax is as shown in Equation (2):

[0038]

[0039] Where N and n are the batch size and class number respectively. x i ∈R d (d is the feature dimension, set to 512) represents the deep feature of the i-th sample, belonging to class y i class. W j ∈R d represents the j-th column of the weight W ∈ R d×n of, b j ∈R n is the bias term,

[0040] Where Loss_arcface is as shown in Equation (3):

[0041]

[0042] Where, s represents a hypersphere with a radius of s, θ is the angle between the weight W and the feature x. m is an additional angular margin penalty added between x i and W yi ;

[0043] During the process of model training, the multi-dimensional face feature extraction is the main body of the model. Therefore, before the i-th training iteration, w1 is set to 0 and w2 is set to 1, that is, Loss_total = Loss_arcface. The weight parameters of the convolutional neural network and the multi-dimensional output fully connected layer are mainly trained,

[0044] When the face feature extraction training reaches stability, take w2 > w1 > 0. At this time, the weight parameters of the two-dimensional output fully connected layer are mainly trained;

[0045] 3) Divide the region according to the two-dimensional plane feature distribution: Since Softmax Loss is introduced, according to its characteristics, the two-dimensional features of the N face data in the database will be radially distributed on the two-dimensional plane after the model is trained. Select a suitable center point, divide the N features into m regions according to the angles radiating from the center point, and save the region number where each feature is located in the database;

[0046] 4) Comparison process: In the face comparison stage, the face to be tested is input into the trained model to obtain a two-dimensional feature vector and a 512-dimensional feature vector. First, determine the eigenvalue region where it is located according to the two-dimensional eigenvalue, and then retrieve the features in the same region in the database according to the region number, and compare the multi-dimensional eigenvalues contained in this region one by one;

[0047] Calculate the cosine distance between the feature to be detected and the features in the database according to the 512-dimensional feature vector, as shown in Equation (4):

[0048]

[0049] The two vectors to be compared are (x1, x2,......, x 512 ) and (y1, y2,......, y 512 ). If there is a feature with a cosine value greater than the set threshold and the largest, output the information of this face, and the recognition result is the same person. If no target face feature is found, for the consideration of recognition rate, all features in two adjacent regions can be continued to be compared or the recognition failure can be directly output;

[0050] Such a comparison method avoids comparing the multi-dimensional face eigenvalues with all the face eigenvalues in the database in a 1:N manner, realizing the acceleration from 1:N comparison to and achieving the acceleration of face eigenvalue comparison for a large number of pictures.

[0051] Figure 1 It is a schematic diagram of the face feature extraction network. As shown in the figure: The input face data passes through the convolutional neural network and passes through the fully connected layer 1 on the left and the fully connected layer 2 on the right respectively. The fully connected layer 1 on the left outputs a 512-dimensional feature vector for the similarity comparison between two features; the fully connected layer 2 on the right outputs a two-dimensional feature vector for dividing the features by angle.

[0052] Figure 2 It is a face feature comparison flow chart. As shown in the figure: Input the face data to be recognized into the network model, extract the multi-dimensional features and two-dimensional features of this face data, determine the feature region where the candidate face is located according to the two-dimensional features, and compare the face to be recognized with the candidate faces in the region according to the multi-dimensional features, and finally obtain the comparison result and output it.

[0053] Figure 3It is a structural diagram for face feature comparison, as shown in the figure: The two-dimensional features obtained from N faces in the database through the face feature extraction model are visualized in a plane according to the two-dimensional features, and the N features are divided into m regions. The face data to be recognized undergoes face feature extraction to obtain multi-dimensional features and two-dimensional features. The two-dimensional features are used to determine the k-th feature region where it is located, and a cosine distance comparison is made with the 512-dimensional features in this region to find the feature in the k-th database with a value less than the threshold and the smallest distance, and then the information of k is output.

Claims

1. A method for accelerating massive face comparison based on deep learning, characterized in that, The method includes the following steps: 1) Modify the structure of the face feature extraction model based on the convolutional neural network: Add a new fully connected layer branch to the face feature extraction network with an original output of 1×512, increasing the output by 1×2. Each face data can obtain a 1×512 face feature value and a 1×2 two-dimensional feature value through the feature extraction network. Among them, the 512-dimensional face feature value is still used for face feature comparison later, and the two-dimensional feature value is visualized in the plane for subsequent feature division regions; 2) Model training, design the loss function as shown in Equation (1): Loss_total = w1×Loss_softmax + w2×Loss_arcface (1) where w1 and w2 are the weights of the loss functions Loss_softmax and Loss_arcface respectively; where Loss_softmax is as shown in Equation (2): where N and n are the batch size and class number respectively, and x i ∈R d , d is the feature dimension, set to 512, representing the depth feature of the i-th sample, belonging to y i class; W j ∈R d represents the j-th column of the weight W ∈ R d×n , and b j ∈R n is the bias term; where Loss_arcface is as shown in Equation (3): where s represents a hypersphere with radius s, θ is the angle between the weight W and the feature x, and m is an additional angular margin penalty added between x i and ; where Loss_arcface is used to extract multi-dimensional features, and Loss_softmax is used to extract two-dimensional features. During the model training process, the extraction of multi-dimensional face features is the main body of the model. Therefore, it is designed that before the I-th training iteration, w1 is 0 and w2 is 1, that is, Loss_total = Loss_arcface, and the weight parameters of the convolutional neural network and the multi-dimensional output fully connected layer are trained; When the face feature extraction training reaches stability, take w2 > w1 > 0, and at this time, train the weight parameters of the two-dimensional output fully connected layer; 3) Divide regions according to the two-dimensional plane feature distribution. Set the number of different registered faces in the database as N, select the center point, and divide the N features registered in the database into m regions according to the angles radiating from the center point. And save the region number where the feature is located together with the 512-dimensional feature value in the database for subsequent comparison processes; 4) Comparison process: Input the face to be tested into the trained model to obtain a two-dimensional feature vector and a 512-dimensional feature vector. First, determine the feature value region where the face to be tested is located according to the two-dimensional feature value, then retrieve the features in the same region in the database according to the region number to narrow the range of 512-dimensional feature value comparison, and then calculate the similarity of each 512-dimensional feature value included in this region one by one; Among them, the calculation of similarity: Calculate the cosine distance between the feature to be detected and the feature in the database according to the 512-dimensional feature vector, as shown in Equation (4): Let the two 512-dimensional vectors to be compared be (x1, x2,......, x 512 ) and (y1, y2,......, y 512 ); 5) Result output and decision: Set the comparison threshold as threshold, screen the face data with a cosine value greater than threshold in the selected feature value region, and sort them from large to small, output the face data with the largest cosine value calculation result, the recognition result is the same person, and output the id of this data.

2. The mass face comparison acceleration method based on deep learning according to claim 1, wherein In step 5), if no face data with a cosine value greater than threshold is found in the selected eigenvalue region, the following decision-making methods can be selected according to the actual requirements for single recognition rate and recognition speed: (1) continue to compare all features in two adjacent regions, and if no result is obtained, continue with this decision-making method; (2) output recognition failure, that is, the face data does not exist in the database.

Citation Information

Patent Citations

  • Human face comparison method and device

    CN107330359A

  • Face recognition method and device based on deep learning

    WO2021218060A1