Method for customizing face data set
Through the combination of the face tracking model and the face recognition model and artificial calibration, the problems of insufficient data storage and insufficient recognition of difficult scenes in the face data classification method in the prior art are solved, and higher recognition accuracy is achieved.
Patent Information
- Application Number
- CN202410117769.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-29
- Publication Date
- 2025-07-29
AI Technical Summary
The prior art face data classification method results in the final saved face data and the inability to effectively identify faces in difficult scenarios, resulting in high recognition error rates and error rates, and it is impossible to effectively save and classify face data in complex scenarios.
The face tracking model is used for preliminary classification, and the phased classification is carried out in combination with different thresholds of the face recognition model. Through manual calibration, the collected face data is finally saved as much as possible, especially in difficult scenarios.
By customizing the method of creating face datasets, the recognition accuracy on low-quality face datasets is improved, the recognition accuracy on the IJBB dataset is improved from 95.6 to 97.7, and the recognition accuracy is achieved in ordinary video monitoring scenarios.
Smart Images

Figure CN120388403A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of face recognition, and particularly relates to a method for customizing a face data set. Background Art
[0002] In the prior art, with the development of technology, deep learning has become an important field of artificial intelligence. Many AI technical problems have been solved through deep learning, such as playing an important role in face recognition applications, object detection applications, and audio applications. In the field of face recognition, face ID annotation is usually performed on the faces in the collected data using an open-source face recognition model on the Internet.
[0003] In the prior art, the general process of face data classification is as follows:
[0004] 1) Collect face data in different scenarios and collect the registration pictures of each person.
[0005] 2) Perform a bottom database registration operation on the collected registration pictures of each person.
[0006] 3) Use a face recognition model to compare the collected data with the bottom database. If the recognition is passed, it is considered to be the same person and the data is saved. All the data is compared to classify the face data.
[0007] However, the defects in the prior art are as follows:
[0008] 1. The general method will result in fewer face data saved in the end. No matter which face recognition model is used, there will be a certain error rate and misrecognition rate. This will cause the faces that fail to pass the recognition to be unable to be distinguished and thus unable to be saved. Among them, some faces are of the same person, but the recognition fails due to the poor ability of the face recognition model.
[0009] 2. The general method can only classify face data in simple scenarios. In some scenarios, the face quality is very poor, and the face recognition model cannot recognize the faces in difficult scenarios.
[0010] In addition, the commonly used technical terms include:
[0011] 1. Face data classification: The process of labeling the pictures of faces that are of the same person in the data set with the same id.
[0012] 2. Face tracking model: A model that can implement the function of tracking the position of faces in a video.
[0013] 3. Trackid: The output data of the face tracking model. The model will output the trackid of each face, and the trackid of the same face in a video is the same.
[0014] 4. Bottom database: A collection of face recognition registration pictures.
[0015] 5. Difficult scenarios: Faces with different lighting conditions, different angles, and different degrees of blur.
[0016] 6. Face angle model: Can judge the size of the face angle.
[0017] 7. Face similarity threshold: A threshold used to determine whether two faces belong to the same person.
[0018] 8. Face alignment: Crop the face part from the original image and scale it to a certain size.
[0019] 9. Id: Equivalent to a person's name. Summary of the Invention
[0020] To solve the above problems, the purpose of this application is to: preliminarily classify face data using face tracking model technology; perform stage classification on face data using different thresholds of face recognition models (manual participation is required in some stages). Furthermore, save as much of the collected face data as possible; and classify face data for difficult scenarios.
[0021] Specifically, the present invention provides a method for customizing a face dataset, and the method includes the following steps:
[0022] S1. Collect face data in different scenarios and collect the registration pictures of each person in the recorded faces:
[0023] First, collect the registration pictures of the collectors. Use a mobile phone or a camera to take the ID photos of each collector, and then collect the videos of the faces of the collectors;
[0024] S2. Cut the collected videos into picture formats and label the picture names according to the actual recording time to facilitate subsequent face picture classification;
[0025] S3. Use the face tracking model and face alignment technology to crop the face parts in the pictures in step S2, and place the faces with the same trackid value obtained by tracking in the same folder;
[0026] S4. Perform face angle screening on the faces aligned by the face angle model, and filter out the pictures with too large face angles, that is, the pictures without face information or pictures of the back of the head;
[0027] S5. Perform a face comparison operation on the registration pictures collected in step S1 and the face data screened in step S4, obtain a face similarity, and save it as a txt text;
[0028] S6. First, select a threshold that is slightly higher than the empirical threshold of the face recognition model; use this slightly higher threshold to determine whether the screened face and the registration image belong to the same person. If, based on the similarity threshold, it is determined that there are 5 or more face images in a trackid that belong to the same person as the registration image, then combine all the images under this trackid and the recognized registration image into a single face data;
[0029] S7. Using the slightly higher threshold will leave some trackids that do not meet the conditions. Then, lower the similarity threshold within the range of 0.6 - 0.3, and repeat the operations in step S6;
[0030] S8. Perform id annotation for the screened faces. The folder named after the name of the person in the registration image is the face recognition dataset required for the face recognition network.
[0031] In step S1, it includes: placing a camera in places where the collecting personnel often pass by, and turning on the recording function to save the video.
[0032] In step S3, for the face tracking: First, extract the face part in the image in step S2 according to the face alignment technology to form a small image, and use the face tracking model to assign a trackid value to each extracted face. Then, place the faces with the same trackid in the same folder. Assume that the face images with trackid 1 are stored in the folder named 1, and the face images with trackid 2 are stored in the folder named 2. For the face alignment: Transform the facial features of the face into a standard face through affine transformation. The standard face means that the facial features are regular, specifically, the positions of the two eyes, one nose, and two mouth corners in the 112 * 112 image are [[38.2946 + 8.0000, 51.6963], [73.5318 + 8.0000, 51.6963], [56.0252 + 8.0000, 71.7366], [41.5493 + 8.0000, 92.3655], [70.7299 + 8.0000, 92.3655]]. For the affine transformation: Use the cv2.warpAffine function.
[0033] In step S4, the face angle model is a face angle model trained using the tensorflow framework, named face_att. This model can obtain a value between 0 and 1 for the face angle. Based on the face angle model, the faces processed in step S3 are screened for face angles. Using 0.4 as the threshold, faces with an angle less than 0.4 are pictures where half of the face is not visible. Since it is difficult for the face recognition model to recognize such faces, faces with an angle value below 0.4 (i.e., pictures with no face information or the back of the head) are screened out.
[0034] In step S5, the registration pictures collected in step S1 and the face data screened in step S4 are used to obtain the feature values of each face picture through the face recognition network. The face recognition network is a face recognition network trained using the resnet101 framework, named facerec. And use the formula np.dot(capmat,regmat_T): Calculate the similarity between each face registration picture in step S1 and each face screened in step S4 through a comparison operation and save it as a txt text. Among them, capmat: the face feature numerical value obtained by the face recognition network for the registration picture collected in step S1; regmat_T: the transpose of the face feature numerical value obtained by the face recognition network for the face pictures screened in step S4.
[0035] In step S6, it further includes:
[0036] Set the threshold slightly higher than the empirical threshold of the face recognition model to 0.6. A normal similarity value greater than 0.4 is considered to be the same person, but there is a one in ten thousand error rate, so it is set to 0.6 here. Then use this threshold as the judgment criterion to determine whether the faces screened in S4 and the registration pictures collected in S1 are the same person.
[0037] Compare the similarity value obtained in S5 with the set similarity threshold. If the similarity value obtained in S5 is greater than the similarity threshold, it is considered to be the same person.
[0038] Then, when it is determined that there are 5 or more face pictures and registration pictures in a trackid folder that are the same person (i.e., the similarity is greater than the threshold), all the pictures in this trackid and the recognized registration picture are combined into a face data folder, and the folder is named after the name of this registration picture.
[0039] In step S7, it should be noted that lowering the threshold may cause inaccurate model classification. Therefore, manual calibration may be required here. Use the human eye to judge whether the registered image and the face in the trackid folder classified in step S3 are of the same person. If they are of the same person, then combine all the images under this trackid and the recognized registered image into a face data folder, and name the folder with the name of the registered image.
[0040] Therefore, the advantages of this application are as follows: Ordinary face dataset production can only obtain relatively clear face data, while in real application scenarios, it is more desirable to be able to recognize relatively difficult faces (for example, relatively blurred faces captured by cameras). Because there are relatively few blurred and real faces close to those captured by real cameras in the datasets downloaded from the Internet, the face recognition network cannot learn the features of such relatively difficult faces. And to solve this problem, it is necessary to provide relatively difficult faces for the face recognition network to learn. Then, using this method, relatively difficult faces can be obtained. The resnet101 face recognition network trained with the dataset made by this method has an accuracy improvement of 2 points (from 95.6 to 97.7) on the IJBB dataset (low-quality face dataset), and reaches an identification accuracy of 96% in ordinary video surveillance scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] The drawings described herein are used to provide a further understanding of the present invention, form a part of this application, and do not limit the present invention.
[0042] Figure 1 It is a schematic flowchart of the method of this application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0043] In order to more clearly understand the technical content and advantages of the present invention, the present invention will be further described in detail below with reference to the drawings.
[0044] As Figure 1 shown, the present invention provides a method for customizing the production of a face dataset, including the following steps:
[0045] S1. Collect face data in different scenarios and collect the registered images of each person whose face is recorded.
[0046] First, collect the registered images of the collectors. Use a mobile phone or a camera to take the ID photos of each collector, and then collect the videos of the collectors' faces.
[0047] For example, place a camera in a place where the collectors often pass by and turn on the recording function to save the video.
[0048] S2. Crop the collected video into picture format and label the picture names according to the actual recording time to facilitate subsequent face picture classification;
[0049] S3. Use the face tracking model and face alignment technology to extract the face parts in the pictures in step S2, and place the faces with the same trackid value obtained by tracking in the same folder;
[0050] The face tracking: First, use the face alignment technology to extract the face parts in the pictures in step S2 to form a small picture, and use the face tracking model to assign a trackid value to each extracted face. Then, place the faces with the same trackid in the same folder (for example, the face pictures with trackid 1 are stored in the folder named 1, and the face pictures with trackid 2 are stored in the folder named 2);
[0051] The face alignment: Transform the facial features of the face into a standard face through affine transformation (the standard face means that the facial features are regular. Specifically, the two eyes, one nose, and two mouth corners are at ([38.2946 + 8.0000, 51.6963], [73.5318 + 8.0000, 51.6963], [56.0252 + 8.0000, 71.7366], [41.5493 + 8.0000, 92.3655], [70.7299 + 8.0000, 92.3655]) in a 112*112 picture
[0052] Affine transformation: cv2.warpAffine function;
[0053] S4. Perform face angle screening on the faces aligned by the face angle model, and filter out the pictures with too large face angles, that is, the pictures without face information or the back of the head;
[0054] The face angle model is a face angle model trained with the tensorflow framework, named face_att. This model can obtain the value of the face angle between (0-1); according to the face angle model, perform face angle screening on the faces processed in step S3 (using 0.4 as the threshold. Generally, the faces with a value less than 0.4 are pictures with less than half a face visible, because it is difficult for the face recognition model to recognize such faces), and filter out the faces with too large angles (the angle values are below 0.4), that is, the pictures without face information or the back of the head;
[0055] S5. Perform a face comparison operation on the registered pictures collected in step S1 and the face data screened in step S4, obtain a face similarity, and save it as a txt text;
[0056] The registered images collected in step S1 and the face data filtered in step S4 are passed through a face recognition network (the face recognition network trained with the resnet101 framework is named facerec) to obtain the feature values of each face image; and the formula np.dot(capmat,regmat_T) is used: Calculate the similarity between each face registration image in step S1 and each face filtered in step S4, and save it as a txt text.
[0057] Among them, capmat: the face feature value obtained by passing the registered image collected in step S1 through the face recognition network;
[0058] regmat_T: the transpose of the face feature value obtained by passing the face image filtered in step S4 through the face recognition network;
[0059] S6. First, select a threshold that is a little higher than the empirical threshold of the face recognition model; use this higher threshold to judge whether the filtered face and the registered image are of the same person; if it is judged according to the similarity threshold that there are 5 or more face images in a trackid that are of the same person as the registered image, then combine all the images under this trackid and the recognized registered image into a face data; further:
[0060] First, select a relatively high threshold of 0.6 for the face recognition model. The normal similarity value greater than 0.4 is considered to be the same person, but there is a one in ten thousand error rate, so 0.6 is adopted. Then use this threshold as the judgment criterion to judge whether the face filtered in S4 and the registered image collected in S1 are of the same person.
[0061] Compare the similarity value obtained in S5 with the set similarity threshold. If the similarity value obtained in S5 is greater than the similarity threshold, it is regarded as the same person;
[0062] Then, when it is judged that there are 5 or more face images in a trackid folder that are of the same person as the registered image (the similarity greater than the threshold means the same person), combine all the images under this trackid and the recognized registered image into a face data folder, and name the folder with the name of the person in the registered image.
[0063] S7. Using the higher threshold will leave some trackids that do not meet the conditions. Then lower the similarity threshold and repeat the operation in step S6. Note that lowering the threshold will cause inaccurate model classification, so manual calibration may be required here; further:
[0064] Using the higher threshold will leave some trackids that do not meet the conditions. Then, lower the similarity threshold, and the reduction range is between (0.6 - 0.3). Repeat the operation in step S6. It should be noted here that lowering the threshold will cause inaccurate model classification. Therefore, manual participation may be required for calibration here. Use the human eye to judge whether the registered image and the face in the trackid folder classified in S3 are the same person. If they are the same person, then combine all the images under this trackid and the recognized registered image into a face data folder, and name the folder with the name of the registered image.
[0065] S8. Do id annotation for the selected faces. The folder named with the name of the registered image is the face recognition dataset required for the face recognition network produced.
[0066] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, various changes and modifications can be made to the embodiments of the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for customizing and making a face dataset, characterized in that, The method comprises the following steps: S1. Collect facial data from different scenes and collect registration images of each person in the recorded face: First, collect the registration pictures of the collectors. Use a mobile phone or camera to take a photo of each collector's ID. Then, collect a video of the collected person's face. S2. Cut the collected video into image format and label the image names according to the actual recording time to facilitate the subsequent face image classification; S3. Cut out the face part in the picture in step S2 according to the face tracking model and face alignment technology, and put the faces with the same trackid value obtained by tracking in the same folder; S4. Perform facial angle screening based on the aligned faces using the facial angle model, filtering out images with excessively large facial angles, i.e., images without facial information or images of the back of the head; S5. Perform a face comparison operation on the registration image collected in step S1 and the face data filtered in step S4 to obtain a face similarity and save it as a txt text; S6. First, select a threshold slightly higher than the empirical threshold of the face recognition model; use this higher threshold to determine whether the screened faces and the registration image are the same person; if the similarity threshold determines that five or more face images and the registration image for a trackID are the same person, then combine all images under this trackID and the recognized registration image into a single face data; S7. If a higher threshold is used, some trackids will not meet the conditions. Then, the similarity threshold is lowered to between 0.6 and 0.3, and step S6 is repeated. S8. Label the selected faces with IDs, and the folders named after the registered people are the face recognition datasets required by the face recognition network.
2. The method for customizing and creating a face dataset according to claim 1, wherein, The step S1 includes placing a camera at a place where the collectors often pass by, turning on the recording function, and saving the video.
3. The method for customizing a face dataset according to claim 1, wherein: In the step S3, The face tracking method is as follows: first, the face part of the image in step S2 is cut out to form a small image according to the face alignment technology, and a trackid value is assigned to each cut out face using the face tracking model. Then, faces with the same trackid are placed in the same folder. For example, the face image with trackid 1 is stored in folder 1, and the face image with trackid 2 is stored in folder 2. The face alignment is as follows: the facial features of the face are converted into a standard face through affine transformation, wherein the standard face has regular facial features, specifically two eyes, one nose and two corners of the mouth in the 112*112 picture [[38.2946+8.0000,51.6963], [73.5318+8.0000,51.6963], [56.0252+8.0000,71.7366], [41.5493+8.0000,92.3655], [70.7299+8.0000,92.3655]]; the affine transformation is performed using the cv2.warpAffine function.
4. A method for customizing and producing a face dataset according to claim 1, characterized in that, In step S4, the face angle model is a face angle model trained using the tensorflow framework and is named face_att. The model can obtain a value between the angle of the face (0-1). The face processed in step S3 is screened for face angle according to the face angle model, with 0.4 as the threshold. Faces with a value less than 0.4 are pictures in which half of the face is not visible, because it is difficult for the face recognition model to recognize such faces. Faces with angle values below 0.4 and too large angles, that is, pictures without face information or back of the head, are screened out.
5. A method for customizing and producing a face data set according to claim 1, characterized in that In step S5, the registration image collected in step S1 and the face data screened in step S4 are passed through a face recognition network to obtain the feature value of each face image, and the face recognition network is a face recognition network trained by the resnet101 framework and named facerec; and Using the formula np.dot(capmat,regmat_T): Calculate the similarity between each face registration image in step S1 and each face screened in step S4, and save it as a txt text; where capmat is the face feature value obtained from the registration image collected in step S1 through the face recognition network; regmat_T is the transpose of the face feature value obtained from the face image screened in step S4 through the face recognition network.
6. A method for customizing and creating a face dataset according to claim 1, characterized in that The step S6 further includes: The threshold value, which is slightly higher than the empirical threshold value of the face recognition model, is set to 0.
6. Normally, a similarity value greater than 0.4 is considered to be the same person, but there is a one in ten thousand error rate, so it is set to 0.6 here. This threshold value is then used as a criterion to determine whether the face screened by S4 and the registration image collected in S1 are the same person. The similarity value obtained in S5 is compared with the set similarity threshold. If the similarity value obtained in S5 is greater than the similarity threshold, it is considered that they are the same person; Then, if there are 5 or more face images in a trackid folder and the registered image are of the same person, that is, if the similarity is greater than the threshold, then all the images under this trackid and the recognized registered image are combined into a face data folder, and the folder is named after the person in the registered image.
7. A method for customizing a face dataset according to claim 1, characterized in that In step S7, it should be noted that lowering the threshold will cause inaccurate model classification, so manual calibration may be required here. The human eye is used to determine whether the face in the registration image and the face in the trackid folder classified in step S3 are the same person. If they are the same person, all the images under this trackid and the recognized registration image are combined into a face data folder, and the folder is named after the person in the registration image.