A method, device, equipment and storage medium for generating multi-pose image data for face recognition
By adjusting the face image generation model and using a potential vector generator, combining face detection and pose estimation, SVM is used for linear classification to generate multi-pose face images, which solves the demand for multi-pose images of the face recognition system in the prior art, and improves the recognition rate and efficiency.
Patent Information
- Application Number
- CN202411421185.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-12
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2044-10-12
AI Technical Summary
The existing face image generation methods cannot effectively generate multi-faceted face images, resulting in low recognition rate of face recognition in actual applications and cannot meet the multi-faceted requirements of face recognition for generated images.
By obtaining the face data set of a specific scene, adjusting the face image generation model, generating a new face data set, using the latent vector generator to capture image details and global structure, performing face detection and pose estimation, filtering and classifying latent vectors, using SVM for linear classification, obtaining decision boundaries, and generating multiple face images with different poses through linear interpolation.
It realizes the generation of face images covering different pose angles, improves the multi-pose adaptability of the face recognition system, improves the recognition rate and efficiency, and reduces unnecessary computing resource consumption.
Smart Images

Figure CN119445625B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of face image generation, and particularly relates to a method, device, equipment and storage medium for generating multi-pose image data for face recognition. Background Art
[0002] With the development of computer vision and deep learning technologies, face image generation has become a research hotspot. However, existing face image generation methods often only focus on the generation of random data. For the training of face recognition, a single image does not have the ability to improve the recognition rate of face recognition in practical applications, and the requirement of face recognition for single-object multi-pose generated images is ignored. Therefore, it is of great significance to develop a method that can generate multiple face images with different face postures according to different pose boundaries for a single face image, so as to improve the recognition rate of face recognition in practical applications. Summary of the Invention
[0003] The present invention provides a method, device, equipment and storage medium for generating multi-pose image data for face recognition, so as to realize active warning braking for the side blind area of a semi-trailer truck train and improve the active safety of the semi-trailer truck train.
[0004] In a first aspect, an embodiment of the present invention provides a method for generating multi-pose image data for face recognition. This method is applied to face image generation, and the method process is as follows:
[0005] S1. Obtain the face dataset information of a specific scenario, fine-tune using a face image generation model, map the model space to a new face data sample space, and obtain a new generation model;
[0006] S2. Use the new generation model to generate a new face dataset, capture the details and global structure of the input image through a set potential vector generator, extract the direction features to form a feature vector representing the face image, and form a set of potential vectors;
[0007] S3. Perform face detection on the face image data in the generated new face dataset, screen according to the result score of the face detection, further perform threshold screening and classification on the qualified faces in combination with the pose estimation angle in the face detection result, divide the qualified face image potential vectors into multiple groups of face image potential vectors according to different pose angle direction threshold ranges, and use the corresponding positive and negative face angle values as the labels of each group of potential vectors;
[0008] S4. Use the potential vector label and the potential vector of the face image as training data, and perform a linear classification operation with the largest margin in the feature space using SVM to obtain the decision boundary of different face pose angles;
[0009] S5. Use the decision boundary to perform linear interpolation on the latent vectors of the face images generated by the generative model to form new latent vectors of face images, and then use the generative model in S1 to generate multiple photos of a person at different pose angles.
[0010] Preferably, in step S1, the face dataset of a specific scenario includes various face pose images of visible light and various face pose images of near-infrared collected from various cameras;
[0011] The face image generation model includes an image generator and an image discriminator;
[0012] On the basis of adjusting the model to the model pre-trained on a large dataset, use the face dataset of a specific scenario to further adjust the parameters of the model, map the model space to a new face data sample space, and obtain a new generative model.
[0013] Preferably, in step S2, the latent vector set is that in the generated images by the generative model in S1, the details and global structure of the input images are captured through a set latent vector generator to form feature vectors representing face images, and a latent vector set Z of face images is formed.
[0014] Preferably, the specific steps of step S3 are as follows:
[0015] S31. Perform face detection on the face image data in the newly generated face dataset, and the obtained face scores should meet the following conditions:
[0016]
[0017] where, I c is the set of face images that meet the requirements; I is the face dataset to be face-detected; score is the face score in the face detection result; threshold is the set score threshold, and the corresponding latent vector set of face images is
[0018] S32. Further perform face pose estimation angle screening and classification on the set of face images that meet the requirements. To obtain the expected number of faces with different poses, the face pose angles include the pitch, yaw, and roll directions of the face pose estimation, and the range of angles included in each direction meets the following conditions:
[0019] pitch angle , |angle| ∈ {0, 5, 10, 15, 20, 25};
[0020] yaw angle , |angle| ∈ {0, 5, 10, 15, 20, 25};
[0021] roll angle , |angle| ∈ {0, 5, 10, 15, 20, 25};
[0022] Among them, pitch angle , yaw angle , roll angle are the angle ranges included in the three directions of pitch, yaw, and roll respectively; the multiple groups of face latent vectors corresponding to the three directions of pitch, yaw, and roll are:
[0023]
[0024] Among them, the angle value is the set angle range; dir represents the three directions of pitch, yaw, and roll;
[0025] S33. Use the positive and negative values of the corresponding face angles as the labels of each group of latent vectors, that is, in the process of screening and classifying the face pose estimation angles in S32, use the positive and negative values of the current face pose angle as the basis for the classification labels, and the following conditions should be met:
[0026]
[0027] Among them, are the labels for screening and classifying the three directions of pitch, yaw, and roll of the face pose estimation angle; the angle value is the set angle range; dir represents the three directions of pitch, yaw, and roll.
[0028] Preferably, the latent vector labels and the latent vectors of the face images in step S4 are used as training data, and their composition is as follows:
[0029]
[0030] Among them, T dir is the training set that needs to be further trained in the three pose estimation directions; the angle value is the set angle range; dir represents the three directions of pitch, yaw, and roll;
[0031] It involves processing the latent vector labels and the latent vectors of the face images. Specifically, step S4 uses a support vector machine to perform a linear classification operation with the largest margin in the feature space to obtain the decision boundaries of the face at different pose angles. These decision boundaries contain the direction features of the current face in the three directions of pitch, yaw, and roll of the pose estimation angle, as follows:
[0032]
[0033] Among them, the angle value is the set angle range; dir represents the three directions of pitch, yaw, and roll.
[0034] Preferably, in step S5, the decision boundary is used to linearly interpolate the latent vectors of the face images generated by the generation model to form new latent vectors of face images, and then the generation model in S1 is used to generate multiple photos of a person at different pose angles. In this process, first, the latent vectors of face images are generated by the generation model, and then the decision boundary learned by the SVM algorithm is used to linearly interpolate these latent vectors. Among them, the process of interpolating to find new latent vectors of face images is as follows:
[0035]
[0036] Among them; L is an arithmetic sequence from start_d to end_d, with a total of steps elements; is the transpose of the decision boundary;
[0037] Through interpolation, new latent vectors are generated in the latent vector space. These latent vectors correspond to face images at different pose angles. Finally, these new latent vectors are input into the generation model in S1 to generate corresponding face images.
[0038] In a second aspect, an embodiment of the present invention provides a multi-pose image data generation device for face recognition, which is applied to store a multi-pose image data generation method for face recognition; the device includes a data collection module, a model adjustment module, an image generation module, a detection and screening module, a classification module, and an interpolation image generation module;
[0039] The data collection module is used to collect and prepare a face data set for a specific scenario;
[0040] The model adjustment module is used to adjust and train the generation model;
[0041] The image generation module is used to generate new face images;
[0042] The detection and screening module is used to implement face detection and pose estimation functions;
[0043] The classification module is used to perform feature classification using SVM;
[0044] The interpolation image generation module is used to complete linear interpolation and the generation of new images.
[0045] In a third aspect, an embodiment of the present invention provides a multi-pose image data generation device for face recognition, which is applied to execute a multi-pose image data generation method for face recognition.
[0046] Fourthly, an embodiment of the present invention further provides a storage medium for generating multi-pose image data for face recognition, which is applied to store a method for generating multi-pose image data for face recognition.
[0047] In summary, due to the adoption of the above technical solutions, the beneficial effects of the present invention are as follows:
[0048] 1. The present invention can meet the face recognition requirements in specific scenarios. By pre-training on a large dataset and then fine-tuning on the dataset of a specific scenario, the adaptability and generalization ability of the model to the data of the specific scenario are improved; by using a generative adversarial network (GANs) to generate high-quality face images, the latent vector generator ensures the consistency of the details and global structure of the generated images;
[0049] 2. Through pose estimation and classification, the present invention ensures that the generated face images cover different pose angles, improving the multi-pose adaptability of the face recognition system; through threshold screening of face detection and pose estimation, the quality and diversity of the generated images are ensured, while reducing unnecessary consumption of computing resources;
[0050] 3. By using the scores in the face detection results and pose angle estimation, the present invention greatly reduces many unfriendly face images, ensuring the quality of the face images. At the same time, using the positive and negative nature of the angles as the labels of the SVM training samples also greatly reduces the difficulty of manual intervention in labeling; due to the linear division of the decision boundary of the SVM, the convergence time of SVM training can be greatly reduced, and at the same time, it can ensure that the obtained decision boundary has better robustness. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0052] Figure 1 The flowchart of a method for generating multi-pose face image data for face recognition according to the present invention is shown. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0053] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0054] Example 1
[0055] With the development of computer vision and deep learning technologies, face image generation has become a research hotspot. However, existing face image generation methods often only focus on the generation of random data. For the training of face recognition, a single image does not have the ability to improve the recognition rate of face recognition in practical applications, and ignores the requirements of face recognition for single-object multi-pose generated images. Therefore, this embodiment develops a multi-pose face image data generation method for face recognition, so as to achieve more accurate and robust face image generation, which is of great significance for improving the accuracy and efficiency of multi-pose face recognition.
[0056] Refer to Figure 1 As shown, a multi-pose face image data generation method for face recognition in this embodiment is as follows:
[0057] S1. Obtain the face dataset information of a specific scenario, fine-tune it using a face image generation model, and map the model space to a new face data sample space to obtain a new generation model.
[0058] S2. Use the new generation model to generate a new face dataset, capture the details and global structure of the input image through a set potential vector generator, extract the directional features to form a feature vector representing the face image, and form a potential vector set.
[0059] S3. Perform face detection on the face image data in the generated new face dataset, screen according to the result scores of the face detection, and further perform threshold screening and classification on the qualified faces in combination with the pose estimation angles in the face detection results. Divide the potential vectors of the qualified face images into multiple groups of face image potential vectors according to different pose angle direction threshold ranges, and use the corresponding positive and negative face angles as the labels of each group of potential vectors.
[0060] S4. Use the potential vector labels and the potential vectors of the face images as training data, and perform a linear classification operation with the maximum margin in the feature space using SVM. Obtain the decision boundaries for different face pose angles.
[0061] S5. Perform linear interpolation on the face image potential vectors generated by the generation model using the decision boundary to obtain multiple photos of a person with different pose angles.
[0062] In step S1, various face pose images of visible light and various near-infrared face pose images collected from various types of cameras in this example are respectively set as subsets of the target face database. A face image generation model of StyleGAN3 including an image generator and an image discriminator is adopted. According to the requirements, the collected face data set is preprocessed to a specified size of 1024*124, and the preprocessed face data set is fine-tuned on the basis of the StyleGAN3 model pre-trained on a large data set, and the model space is mapped to a new face data sample space to obtain a new generation model StyleGAN3_s. This approach can reduce the training time of the model while obtaining a model with the same characteristics as the custom data.
[0063] In step S2, a new face image data set is generated using the new image generator model StyleGAN3_s obtained in step S1. In this example, in order to ensure good results in subsequent steps, it is set that the number of generated face image data is 50k, which constitutes a new face image data set I. At the same time, the latent vector generator set by StyleGAN3_s is used to capture the details and global structure of the input image, form a high-dimensional vector representing the face image, and form a latent vector set Z of the face image.
[0064] With the new face image data set I and the corresponding face image latent vector set Z obtained through step S1 and step S2, further screening is required to filter out unqualified faces and the corresponding latent vectors. In step S3 of this example, the following steps are executed:
[0065] S31 performs face detection on the face image data in the generated new face data set, and the obtained face score should meet the following conditions:
[0066]
[0067] where, I c is a set of face images that meet the requirements; I is the face data set to be face-detected; score is the face score in the face detection result; threshold is the set score threshold, and the corresponding latent vector set of the face image is
[0068] In this example, the value of threshold is recommended to be set to 0.9. The finally obtained I c is a set of face images that meet the requirements.
[0069] Furthermore, in step S32, for the set of face images I that meet the requirements cFurther, perform face pose estimation angle screening and classification. To obtain more faces with different poses, the face pose angles should include the pitch, yaw, and roll directions of face pose estimation. The angle ranges included in each direction should meet the following conditions:
[0070] pitch angle , |angle| ∈ {0, 5, 10, 15, 20, 25};
[0071] yaw angle , |angle| ∈ {0, 5, 10, 15, 20, 25};
[0072] roll angle , |angle| ∈ {0, 5, 10, 15, 20, 25};
[0073] Among them, pitch angle , yaw angle , roll angle are the angle ranges included in the three directions of pitch, yaw, and roll of pose estimation. Since the true values of the angles have the property of positive and negative values, the angle ranges mentioned in the above conditions formula refer to the angle ranges after taking the absolute value. Further, while obtaining the angle ranges included in the three directions of pitch, yaw, and roll of pose estimation, multiple groups of face latent vectors corresponding to each angle range are also screened, extracted, and saved. The multiple groups of face latent space sets obtained are as follows:
[0074]
[0075] Among them, the value of angle is the set angle range, and dir represents the three directions of pitch, yaw, and roll.
[0076] Then further, in step S33, according to the positive and negative value attributes of the true face angles, in this example, the positive and negative values of each face angle are used as the classification label basis for each group of latent vectors. The labels should meet the following conditions:
[0077]
[0078] Among them, are the labels for screening and classification of the three directions of pitch, yaw, and roll of face pose estimation; the value of angle is the set angle range; dir represents the three directions of pitch, yaw, and roll.
[0079] Further, for the face latent space set obtained in step S3: and the corresponding classification label set: Form a training data set, which is composed as follows:
[0080]
[0081] Among them, T dir is the training set that needs to be further trained in three pose estimation directions. The angle value is the set angle range, and dir represents the three directions of pitch, yaw, and roll. Further, this step uses a support vector machine (SVM) to perform a linear classification operation with the maximum margin in the feature space. Through this operation process, the decision boundaries of the face at different pose angles are obtained. These decision boundaries contain the directional features of the current face in the three directions of pitch, yaw, and roll in the pose estimation angle.
[0082]
[0083] Among them, the angle value is the set angle range, and dir represents the three directions of pitch, yaw, and roll. Thus, more accurate and robust face image generation is achieved, and this step is of great significance for improving the accuracy and efficiency of multi-pose face recognition.
[0084] Finally, use the described decision boundary to perform linear interpolation on the face image latent vectors generated by the generation model to form new face image latent vectors. It should be noted that the process of interpolating to obtain the new face image latent vectors is as follows:
[0085]
[0086] Among them, L is an arithmetic progression from start_d to end_d, with a total of steps elements. is the transpose of the decision boundary. In this embodiment, the values of start_d and end_d in the arithmetic progression L are -3 and 3 respectively, and the value of steps is 10. According to the formula, first calculate the dot product of the face latent vector and the transpose of the decision boundary of the face at different pose angles , and subtract this result from the arithmetic progression L. The purpose of this step is to adjust the arithmetic progression L so that it is offset relative to the subspace defined by and . Finally, perform an interpolation operation through the face latent vector and the adjusted arithmetic progression L and . Here, each value in the adjusted arithmetic progression L will be multiplied by the corresponding column in , and then the result will be added to . In this way, each sample in the adjusted arithmetic progression L will move along Perform linear interpolation in the defined direction, and the distance or step size of the interpolation is determined by the values in the adjusted arithmetic progression L. Through interpolation, new latent vectors can be generated in the latent vector space. This vector contains these latent vectors corresponding to face images at different pose angles. Finally, input these new latent vectors into the stylegan3_s model described in S1, and corresponding face images can be generated. In this embodiment, since the value of steps is 10, 10 face photos in this direction will be generated, and these photos have different direction angles. This method can generate a series of face images with continuously changing pose angles, providing rich training data for applications such as face recognition.
[0087] A multi-pose image data generation device for face recognition, including a data collection module, a model adjustment module, an image generation module, a detection and screening module, a classification module, and an interpolation image generation module;
[0088] The data collection module is used to collect and prepare a face dataset for a specific scenario;
[0089] The model adjustment module is used to adjust and train a generation model;
[0090] The image generation module is used to generate new face images;
[0091] The detection and screening module is used to implement face detection and pose estimation functions;
[0092] The classification module is used to perform feature classification using SVM;
[0093] The interpolation image generation module is used to complete linear interpolation and generate new images.
[0094] The beneficial effects of this embodiment are as follows: By pre-training on a large dataset and then fine-tuning on the dataset for a specific scenario, the adaptability and generalization ability of the model to the data of the specific scenario are improved; Using a generative adversarial network to generate high-quality face images, the latent vector generator ensures the consistency of the details and global structure of the generated images; Through pose estimation and classification, it is ensured that the generated face images cover different pose angles, improving the multi-pose adaptability of the face recognition system; Through threshold screening of face detection and pose estimation, the quality and diversity of the generated images are ensured, while reducing unnecessary consumption of computing resources; Using a support vector machine for classification can obtain an accurate decision boundary; By generating new images in the latent vector space through linear interpolation, the image dataset is effectively expanded, enabling the model to also perform well at pose angles not seen before.
[0095] As described above, only the preferred specific embodiments of the present invention are provided, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, making equivalent substitutions or changes, should be covered within the protection scope of the present invention.
[0096] The preferred embodiments of the present invention disclosed above are only used to help illustrate the present invention. The preferred embodiments do not describe all the details in detail, nor do they limit the present invention to only the specific embodiments. Obviously, many modifications and changes can be made according to the content of this specification. These embodiments are selected and specifically described in this specification to better explain the principles and practical applications of the present invention, so that those skilled in the art can well understand and utilize the present invention. The present invention is only limited by the claims and their full scope and equivalents.
Claims
1. A method for generating multi-pose image data for face recognition, characterized in that: The process is as follows: S1. Obtain face dataset information for a specific scene, use the face image generation model to perform fine-tuning, map the model space to the new face data sample space, and obtain a new generation model; S2, using the new generative model to generate a new face data set, capturing the details and global structure of the input image through the set latent vector generator, extracting directional features to form feature vectors that characterize the face image, and forming a latent vector set; S3, performing face detection on the face image data in the generated new face data set, screening according to the result score of the face detection, further screening and classifying the faces that meet the requirements based on the threshold value of the posture estimation angle in the face detection result, dividing the latent vectors of the face images that meet the requirements into multiple groups of face image latent vectors according to different posture angle direction threshold ranges, and taking the corresponding positive and negative values of the face angle as the label of each group of latent vectors; S4, using the latent vector label and the latent vector of the face image as training data, using SVM to perform a linear classification operation with the largest interval in the feature space to obtain decision boundaries at different face posture angles; S5, using the decision boundary to linearly interpolate the face image latent vector generated by the generative model to form a new face image latent vector, and then using the generative model described in S1 to generate multiple photos of a person at different postures and angles; In step S1, the face data set of the specific scene includes multiple face posture images of visible light and multiple face posture images of near infrared collected from various cameras; The face image generation model includes an image generator and an image discriminator; The model is adjusted based on a model pre-trained on a large dataset, and a face dataset of a specific scene is used to further adjust the parameters of the model, and the model space is mapped to a new face data sample space to obtain a new generative model; The latent vector label and the latent vector of the face image described in step S4 are used as training data, which are composed as follows: ; in, The training set for the three poses to estimate the direction needs to be trained in the next step; The value is the angle range of the setting; represent , , Three directions; It involves processing the latent vector labels and the latent vectors of the face image. Specifically, step S4 uses a support vector machine to perform a linear classification operation with the largest interval in the feature space to obtain the decision boundaries of the face at different posture angles. These decision boundaries include the directional features of the current face in the three directions of pitch, yaw, and roll at the posture estimation angle, as shown below: ; in, The value is the angle range of the setting; represent , , Three directions.
2. The method for generating multi-pose image data for face recognition according to claim 1, characterized in that: In step S2, the latent vector set is the feature vector that represents the face image, and forms the latent vector set Z of the face image. The generative model in S1 captures the details and global structure of the input image through the set latent vector generator in the generated image to form a latent vector set Z of the face image.
3. The method for generating multi-pose image data for face recognition according to claim 1, characterized in that: The specific steps of step S3 are as follows: S31, perform face detection on the face image data in the generated new face data set, and the obtained face score meets the following conditions: ; in, A collection of face images that meet the requirements; A face dataset for face detection; is the face score in the face detection result; is the score threshold set, and the corresponding potential vector set of the face image is ; S32, further screening and classifying the face pose estimation angles for the face image set that meets the requirements, in order to obtain the expected number of faces with different poses, the face pose angles include three directions of pitch, yaw, and roll for face pose estimation, and the range of the angles in each direction meets the following conditions: ; ; ; in, , , They are , , The angle range included in the three directions; , , The multiple groups of face potential vectors corresponding to the three directions are: ; in, The value is the angle range of the setting; represent , , Three directions; S33, the positive and negative values of the corresponding face angles are used as labels for each group of potential vectors, that is, in the process of face posture estimation angle screening and classification in S32, the positive and negative values of the current face posture angles are used as the classification label basis, which should meet the following conditions: ; in, Angle for face pose estimation , , Filter the classification labels in three directions; The value is the angle range of the setting; represent , , Three directions.
4. The method for generating multi-pose image data for face recognition according to claim 1, characterized in that: Step S5 uses the decision boundary to perform linear interpolation on the latent vector of the face image generated by the generative model to form a new latent vector of the face image, and then uses the generative model described in S1 to generate multiple photos of a person at different postures and angles. In this process, the latent vector of the face image is first generated by the generative model, and then the decision boundary learned by the SVM algorithm is used to perform linear interpolation on these latent vectors. The process of interpolating the new latent vector of the face image is as follows: ; in; is an arithmetic progression from start_d to end_d, with a total of steps elements; is the transpose of the decision boundary; Through interpolation, new latent vectors are generated in the latent vector space, which correspond to face images at different posture angles. Finally, these new latent vectors are input into the generative model described in S1 to generate the corresponding face images.
5. A multi-pose image data generation device for face recognition, characterized in that: A method for generating multi-pose image data for face recognition according to any one of claims 1 to 4 is used to store the data; the device comprises a data collection module, a model adjustment module, an image generation module, a detection and screening module, a classification module and an interpolation image generation module; The data collection module is used to collect and prepare face datasets for specific scenarios; The model adjustment module is used to adjust and train the generated model; The image generation module is used to generate new face images; The detection and screening module is used to realize face detection and posture estimation functions; The classification module is used to classify features using SVM; The interpolation image generation module is used to complete linear interpolation and generate new images.
6. A multi-pose image data generation device for face recognition, characterized in that: Applicable to executing a method for generating multi-pose image data for face recognition as described in any one of claims 1 to 4.
7. A storage medium for generating multi-pose image data for face recognition, characterized in that: Applicable to storing a method for generating multi-pose image data for face recognition as described in any one of claims 1-4.
Citation Information
Patent Citations
Face recognition method based on head posture
CN110096965A
Method and system for classifying images
GB202401994D0