Method and System for Constructing a Facial Image Recognition Feature Library
By constructing a facial image recognition feature library and fine-tuning the face recognition model, the problem of high error detection rate in the existing technology is solved, and the accuracy and security of facial recognition are improved.
Patent Information
- Application Number
- CN202411126794.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-16
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2044-08-16
AI Technical Summary
The existing facial recognition technology has a high false detection rate in the safety management of public places and crowded areas, resulting in safety hazards.
By constructing a facial image recognition feature library, obtaining personnel information of the monitoring object, extracting and aligning the face images, data augmentation, extracting feature vectors, and generating noise features through random noise sampling and reverse engineering, updating the feature library, and fine-tuning the face recognition model.
It improves the accuracy and robustness of the face recognition model, reduces the false detection rate, enhances the recognition ability of specific personnel, and reduces security risks.
Smart Images

Figure CN119107683B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent security, in particular to identity recognition technology, and more specifically to a method and system for constructing a face image recognition feature library. Background Art
[0002] With the continuous optimization of computer image vision algorithms and machine learning algorithms, the use of face image recognition technology for identity recognition has been widely applied, covering multiple aspects from traditional access control systems, security and monitoring, financial services, intelligent devices (smart phones, computers, etc.), retail and customer service, traffic management, social media, healthcare, education and training, tourism and hotels, smart homes, etc., covering many aspects from security monitoring to personalized services, and face recognition technology has demonstrated its diversity and application potential.
[0003] In security, monitoring, and tracking systems based on face image recognition, for special identification and tracking of specific personnel, sensitive personnel, etc., on the one hand, the 1:N method is used for identification comparison and tracking. At this time, the identity information and face information of the monitored objects are pre-entered in the database of the security monitoring system. After the system captures a photo of "me", it finds the image that matches the face data of the current user from the massive portrait database and performs matching to find out "who I am", achieving high-precision identity recognition and security management; on the other hand, the system can also judge suspicious personnel according to specific algorithms and track the trajectories of the identified suspicious personnel to achieve security warning and alarm.
[0004] Face recognition technology acquires video images through collection, extracts facial images from them, and then uses face recognition algorithms to calculate and analyze features such as the positions of facial features, face shapes, and angles of the face, and then compares them with the existing feature library in its own database to determine the true identity of the user. Face recognition algorithms mainly complete the extraction of face features, compare them with the known faces in the inventory, and complete the final classification output. The accuracy and robustness of face recognition algorithms are the premise for accurate recognition, and are obtained through training with a specific algorithm model based on a massive collection and processing of face feature libraries.
[0005] Building a face recognition database is a crucial step, which is the core part of security and tracking systems based on face image recognition and directly affects the accuracy and reliability of the system. The face recognition database contains a vast amount of feature information of different individuals. By training to improve the performance and capabilities of the face recognition model, it can enhance the accurate recognition rate of the system for targets and reduce false alarms and missed detections. For identified individuals, such as key security personnel, sensitive individuals, fugitives, etc., although the recognition algorithms trained using general face feature databases and algorithm models can achieve basic identity recognition, such as a recognition accuracy rate of over 90% for commercial use, there is still a certain false detection rate, posing a potential threat to the security of public places and certain areas with public attributes and crowded populations. Summary of the Invention
[0006] In view of the problems existing in the prior art, according to the first aspect of the object of the present invention, a method for constructing a face image recognition feature library is proposed, including:
[0007] Obtain information of multiple monitored object persons, and each monitored object person information includes personal information and personal images;
[0008] Extract face images from the personal images in the first database using a face detection model to construct a face image library;
[0009] Align the extracted face images so that the faces in the face images are at a unified angle and size;
[0010] Adjust the face images using at least one of image rotation, flipping, scaling, and brightness, and add each adjusted face image to the face image library;
[0011] Extract face features from each face image in the face image library, generate a feature vector of a fixed length for each face image, and construct a first face feature library;
[0012] For the feature vector corresponding to each face image in the first face feature library, randomly sample one of Gaussian noise, salt-and-pepper noise, and Poisson noise as the noise feature, reverse the noise feature, and generate a random noise vector feature corresponding to the noise signal;
[0013] Discriminate between the random noise vector feature and the feature vector, and use the discrimination result as an increment to supplement it to the first face feature library and update the first face feature library;
[0014] Fine-tune the trained face recognition model using the updated first face feature library to obtain an updated face recognition model.
[0015] As an alternative embodiment, for each feature vector corresponding to a face image in the first face feature library, randomly sample one of Gaussian noise, salt-and-pepper noise, and Poisson noise as the noise feature, and reverse the noise feature to generate a random noise vector feature corresponding to the noise signal, including:
[0016] For each feature vector corresponding to a face image in the first face feature library, construct Gaussian noise, salt-and-pepper noise, and Poisson noise functions, randomly sample one of the three types of Gaussian noise, salt-and-pepper noise, and Poisson noise to construct a random noise vector feature, and perform reverse engineering through the diffusion model DM to generate a random noise vector feature corresponding to the noise signal.
[0017] As an alternative embodiment, the diffusion model DM is implemented by neural network training to enable it to effectively reverse the noise process during the reverse process. During the training process, the root mean square error RMSE and the MSE objective function are used to make the model converge, minimizing the difference between the original feature vector and the data recovered by reverse.
[0018] As an alternative embodiment, discriminating the random noise vector feature and the feature vector, and using the discrimination result as an increment to supplement and update the first face feature library, including:
[0019] Use an adversarial domain adaptation model for discrimination, with the random noise vector feature as the target domain and the feature vector as the source domain, and the two have a unified feature dimension; the adversarial domain adaptation model includes a domain classifier and a feature generator, where the feature generator attempts to generate features that cannot be distinguished by the domain classifier as coming from the source domain or the target domain, thereby forcing the generated features to have domain invariance; use adversarial training to encourage the feature generator to generate features that can deceive the domain classifier, and use the random noise vector features discriminated as the source domain as an increment to supplement the aforementioned first face feature library.
[0020] According to the second aspect of the object of the present invention, a method for constructing a face image recognition feature library is also proposed, including:
[0021] Obtain information of multiple monitored object persons, and each monitored object person information includes personal information and personal images;
[0022] Extract face images from the personal images in the first database using a face detection model to construct a face image library;
[0023] Align the extracted face images so that the faces in the face images are at a unified angle and size;
[0024] Adjust the face images using at least one of image rotation, flipping, scaling, and brightness, and add each adjusted face image to the face image library;
[0025] Extract face features from each face image in the face image library. Each face image generates a feature vector of a fixed length, and a first face feature library is constructed.
[0026] For the feature vector corresponding to each face image in the first face feature library, randomly sample one of Gaussian noise, salt-and-pepper noise, and Poisson noise as the noise feature, reverse the noise feature, and generate a random noise vector feature corresponding to the noise signal.
[0027] Discriminate the random noise vector feature and the feature vector, and use the discrimination result as an increment to supplement it to the first face feature library and update the first face feature library until the number of samples in the first face feature library reaches the expected number.
[0028] Obtain a public face dataset, which contains diverse face images, including face images of different ages, genders, expressions, and lighting conditions.
[0029] Adjust the face images in the public face dataset using at least one of image rotation, flipping, scaling, and brightness, and add each adjusted face image to the face image library.
[0030] Extract face features from each face image in the face image library. Each face image generates a feature vector of a fixed length and adds it to the first face feature library.
[0031] Use the first face feature library to train a deep learning-based face recognition model for subsequent face recognition tasks.
[0032] As an optional method, for the feature vector corresponding to each face image in the first face feature library, randomly sample one of Gaussian noise, salt-and-pepper noise, and Poisson noise as the noise feature, reverse the noise feature, and generate a random noise vector feature corresponding to the noise signal, including:
[0033] For the feature vector corresponding to each face image in the first face feature library, construct Gaussian noise, salt-and-pepper noise, and Poisson noise functions, randomly sample one of the three types of Gaussian noise, salt-and-pepper noise, and Poisson noise to construct a random noise vector feature, and perform reverse engineering through the diffusion model DM to generate a random noise vector feature corresponding to the noise signal.
[0034] Among them, the diffusion model DM is implemented by neural network training so that it can effectively reverse the noise process during the reverse process. During the training process, the root mean square error RMSE and the MSE objective function are used to make the model converge, minimizing the difference between the original feature vector and the data recovered by reverse.
[0035] As an alternative method, the discrimination of the random noise vector feature and the feature vector, and using the discrimination result as an increment to supplement and update the first face feature library includes:
[0036] Use an adversarial domain adaptation model for discrimination, with the random noise vector feature as the target domain and the feature vector as the source domain, and both having a unified feature dimension; the adversarial domain adaptation model includes a domain classifier and a feature generator, where the feature generator attempts to generate features that cannot be distinguished by the domain classifier as being from the source domain or the target domain, thereby forcing the generated features to have domain invariance; use adversarial training to encourage the feature generator to produce features that can deceive the domain classifier, and use the random noise vector features discriminated as the source domain as an increment to supplement the aforementioned first face feature library.
[0037] According to the third aspect of the purpose of the present invention, a computer system is also proposed, including:
[0038] One or more processors;
[0039] A memory for storing operable instructions;
[0040] Wherein, when the instructions are executed by one or more processors, the one or more processors are caused to perform operations, and the operations include executing the process of the aforementioned method for constructing a face image recognition feature library.
[0041] Combining the method and system for constructing a face image recognition feature library proposed in the above embodiments can improve the quality and data diversity of the feature library for identified security monitoring personnel, such as key security personnel, sensitive personnel, etc., and accordingly construct a face recognition model or fine-tune an existing face recognition model. It is particularly applicable to security and monitoring systems in key industries, key areas, and public management areas, improving the performance and recognition accuracy of the system's face recognition model, achieving precise and efficient recognition of specific personnel, and reducing security risks and hidden dangers.
[0042] It should be understood that all combinations of the foregoing concepts and additional concepts described in more detail below can be regarded as part of the inventive subject matter of the present disclosure as long as such concepts do not conflict with each other. Additionally, all combinations of the claimed subject matter are regarded as part of the inventive subject matter of the present disclosure.
[0043] The foregoing and other aspects, embodiments, and features of the teachings of the present invention can be more fully understood from the following description in conjunction with the accompanying drawings. Other additional aspects of the present invention, such as features and / or beneficial effects of exemplary embodiments, will be apparent in the following description or will be learned through practice of specific embodiments in accordance with the teachings of the present invention. Description of the Drawings
[0044] The accompanying drawings are not intended to be drawn to scale. In the accompanying drawings, each identical or nearly identical component shown in each figure may be denoted by the same reference numeral. For the sake of clarity, not every component is labeled in each figure. Now, embodiments of various aspects of the present invention will be described by way of example and with reference to the accompanying drawings.
[0045] Figure 1 is a schematic diagram of a method for constructing a face image recognition feature library according to a first embodiment of the present invention.
[0046] Figure 2 is a schematic diagram of a method for constructing a face image recognition feature library according to a second embodiment of the present invention. Detailed Embodiments
[0047] To better understand the technical content of the present invention, specific embodiments are hereby given and described in conjunction with the accompanying drawings as follows.
[0048] In the present disclosure, aspects of the present invention are described with reference to the accompanying drawings, in which many illustrative embodiments are shown. The embodiments of the present disclosure are not necessarily intended to cover all aspects of the present invention. It should be understood that the various concepts and embodiments introduced above, as well as those concepts and embodiments described in more detail below, can be implemented in any of a number of ways, because the concepts and embodiments disclosed in the present invention are not limited to any one embodiment. Additionally, some aspects of the present invention can be used alone, or in any suitable combination with any other aspects of the present invention.
[0049] {Embodiment 1}
[0050] Combined with Figure 1 the method for constructing a face image recognition feature library shown in the example, includes the following steps:
[0051] S101. Obtain information of multiple monitored object persons, and each monitored object person information includes personal information and personal images;
[0052] S102. Use a face detection model to extract face images from the personal images in the first database, and construct a face image library;
[0053] S103. Align the extracted face images so that the faces in the face images are at a unified angle and size;
[0054] S104. Adjust the face images using at least one of image rotation, flipping, scaling, and brightness, and add each adjusted face image to the face image library;
[0055] S105. Extract face features from each face image in the face image library. Generate a feature vector of a fixed length for each face image and construct a first face feature library.
[0056] S106. For the feature vector corresponding to each face image in the first face feature library, randomly sample one of Gaussian noise, salt-and-pepper noise, and Poisson noise as the noise feature. Reverse the noise feature to generate a random noise vector feature corresponding to the noise signal.
[0057] S107. Discriminate between the random noise vector feature and the feature vector, and use the discrimination result as an increment to supplement it to the first face feature library and update the first face feature library; and
[0058] S108. Use the updated first face feature library to fine-tune the trained face recognition model to obtain an updated face recognition model.
[0059] It should be understood that in the embodiments of the present invention, the monitored object personnel especially refer to the personnel who are under key attention and monitoring, including but not limited to the regulated personnel identified by the financial system, enterprise public management system, tax management system, judicial execution management system, credit management system, social security management system, etc., such as fugitives, major suspects, persons under execution, persons restricted from high consumption, fee evasion persons, and persons with bad credit. The information of these personnel is collected and aggregated into a unified key security monitoring library to achieve key deployment, especially in the video monitoring systems of airports, railway stations, subways, banks, financial institutions, and other places and venues involving public security and public management.
[0060] The aforementioned personal information includes personal name and at least one of the following information: gender information, identity number information, contact address information, and contact phone number.
[0061] The aforementioned personal image includes the personal image on the identity document or the personal image collected by other image acquisition devices, and the personal image contains an image of the face part.
[0062] As an optional embodiment, in the face detection step of the aforementioned step S102, use one of the Haar cascade classification detection algorithm, LBP feature classifier detection algorithm, FisherFace classification algorithm based on linear discriminant analysis, and MTCNN multi-task convolutional neural network detection algorithm to detect the face area and extract the face image in the image. Of course, other detection algorithms based on deep learning network models can also be used, such as the Yolo deep learning detection method, to extract the face area from the image and obtain the face image.
[0063] Further, on the basis of the extracted face image, in order to reduce the influence of pose changes and angles, especially the influence of face images extracted from personal images other than identity documents, in step S103 of the embodiment of the present invention, the extracted face image is aligned so that the faces in the face images are at a unified angle and size, such as a frontal image and the size is 128*128, 64*64, etc. In the present invention, the size of 64*64 is adopted.
[0064] As an optional method, the foregoing face image is aligned by means of affine transformation, which includes:
[0065] The detected face region is further processed to extract the key points of the face, including the positions of the eyes, nose, and corners of the mouth;
[0066] According to the positions of the key points, an affine transformation matrix is calculated to align the face to a standard position. In the implementation of the present invention, two eyes and the nose are selected as the reference points for alignment;
[0067] The face image is rotated and scaled so that the horizontal positions of the two eyes are aligned and the positions of the nose and mouth are the same, so as to standardize the pose, angle, and scale of the face;
[0068] Finally, the aligned face region is cropped to a fixed size (such as 64*64 as described above), and pixel value normalization is performed to eliminate brightness differences to ensure that the image size, angle, and brightness input to the subsequent model are consistent.
[0069] As an optional embodiment, in step S104 of the foregoing method, at least one of image rotation, flipping, scaling, and brightness is used to adjust the face image, and each adjusted face image is added to the face image library. Through data augmentation processing such as rotation, flipping, scaling, and brightness adjustment, the model can meet the requirements of sample diversity and improve the robustness of the subsequent training model.
[0070] As an optional embodiment, in the foregoing step S105, face features are extracted from each face image in the face image library, and each face image generates a feature vector of a fixed length to construct a first face feature library, including:
[0071] A pre-trained model (such as FaceNet, VGGFace) in a deep learning framework (such as TensorFlow, PyTorch) is used to extract face features to obtain a feature vector of a fixed length. For example, in the embodiment of the present invention, 128 dimensions are selected for subsequent training, matching, and recognition.
[0072] Among them, the length of the feature vector of the fixed length is pre-configured and is greater than or equal to 108.
[0073] It should be understood that in specific implementations, for the constructed first face feature library, a suitable database management system (such as MySQL, MongoDB) can be selected to store feature vectors and related information, such as names, IDs, and other personal information data, to form sample data.
[0074] Combined with Figure 1 Taking the example shown, in the aforementioned step 106, based on the face feature library, for each feature vector corresponding to a face image in the first face feature library, one of Gaussian noise, salt-and-pepper noise, and Poisson noise is randomly sampled as the noise feature, and the noise feature is reversed to generate a random noise vector feature corresponding to the noise signal.
[0075] Specifically, for each feature vector corresponding to a face image in the first face feature library, Gaussian noise, salt-and-pepper noise, and Poisson noise functions are constructed, and one of the three types of Gaussian noise, salt-and-pepper noise, and Poisson noise is randomly sampled to construct a random noise vector feature. Through the diffusion model DM, reverse engineering is performed to generate a random noise vector feature corresponding to the noise signal.
[0076] In the embodiment of the present invention, the diffusion model DM has two aspects of functions: on the one hand, in the forward stage, starting from the feature vector, noise (salt-and-pepper noise, Poisson noise, or Gaussian noise) is continuously introduced on its basis, so that the data is transformed into pure noise through multiple time steps. This process is realized through a series of small conditional probability steps, and more and more noise is added to the data in each step until the data becomes pure noise; on the other hand, in the reverse generation stage, it is the reverse process of forward diffusion, and the purpose is to gradually reconstruct the original data from the pure noise state. In this stage, the model learns to remove noise in each step and gradually restores the original structure of the data. This process can be parameterized by a neural network to remove noise.
[0077] As an optional implementation manner, the aforementioned diffusion model DM is implemented by neural network training, so that it can effectively reverse the noise process in the reverse process. During the training process, the root mean square error RMSE and the MSE objective function are used to make the model converge, so as to minimize the difference between the original feature vector and the data reversely restored.
[0078] It should be understood that Gaussian noise, salt-and-pepper noise, and Poisson noise are typical noises in images. In the embodiments of the present invention, by randomly sampling Gaussian noise, salt-and-pepper noise, and Poisson noise, the diversity of the face features of the determined range samples (key monitored personnel) is improved, and the scale and quality of the samples are improved by the introduced noise, thereby improving the accuracy of subsequent model training. In particular, after simulating the addition of Poisson noise to the image and expanding the data samples, and then training and enhancing the model, it is particularly beneficial to process image enhancement and face recognition effects in indoor environments, low-light conditions, or poor lighting conditions (not in the natural outdoor sunlight environment), and improve the accuracy of recognition.
[0079] It should be understood that in the embodiments of the present invention, after simulating the addition of salt-and-pepper noise to the face image features, the diversity of the training data is increased, and the robustness of the model when processing poor collected image data input is improved.
[0080] As an optional implementation manner, in the foregoing step S107, the discriminant of the random noise vector feature and the feature vector is performed, and the result is used as an increment according to the discriminant result and supplemented to the first face feature library to update the first face feature library, including:
[0081] Use an adversarial domain adaptation model for discrimination, with the random noise vector feature as the target domain and the feature vector as the source domain, and the two have a unified feature dimension; the adversarial domain adaptation model includes a domain classifier and a feature generator, where the feature generator attempts to generate features that cannot be distinguished by the domain classifier as coming from the source domain or the target domain, thereby forcing the generated features to have domain invariance.
[0082] In the embodiments of the present invention, adversarial training is used to encourage the feature generator to generate features that can deceive the domain classifier, and the random noise vector features discriminated as the source domain are used as an increment and supplemented to the foregoing first face feature library.
[0083] As an example, the foregoing adversarial domain adaptation model adopts a DANN (Domain-Adversarial Neural Network) deep adversarial neural network model.
[0084] Thus, by randomly sampling noise features and continuously performing reverse and domain discrimination, the high-quality face feature library is expanded and supplemented, the signal-to-noise ratio of the samples is improved, and it is determined whether to continue randomly sampling noise features by judging whether the number of samples in the first face feature library reaches the expected number. If the expected number is reached, the sampling of noise features is stopped, otherwise, the random sampling of noise signals and the reverse generation of new incremental data are continued and expanded to the first face feature library.
[0085] As an example of the present invention, the number of samples in the first face feature library is set to be determined according to the number m of key monitored persons, and is configured to be 10 to 50 times the number m.
[0086] In an embodiment of the present invention, in the foregoing step S108, the trained face recognition model is fine-tuned using the updated first face feature library to obtain an updated face recognition model.
[0087] It should be understood that the foregoing trained face recognition model refers to a pre-trained face recognition model, especially a recognition model based on deep learning, including but not limited to a face detection model based on OpenCV, a face detection model based on Dlib, a face recognition model based on CNNs (such as a face recognition model based on VGG-16, a face recognition model based on FaceNet developed by Google, a face recognition model based on DeepFace developed by Facebook, etc.), a face recognition model based on RNNs (such as a face recognition model developed based on RNN, LSTM, GRU, etc.), a face recognition model based on ensemble learning, etc.
[0088] On the one hand, these face recognition models pre-trained based on the public face feature library have the characteristics of universality and robustness. On the other hand, their performance and accuracy need to be improved for the recognition of specific ranges of personnel. On this basis, the present invention fine-tunes (trains) through the updated first face feature library in the method of the above embodiments to obtain a face recognition model with better performance, so as to achieve efficient and highly accurate detection and recognition of specific ranges of personnel.
[0089] As an optional example, below we exemplarily elaborate on the methods and processes of randomly sampling Gaussian noise, salt-and-pepper noise, and Poisson noise.
[0090] Gaussian noise, also known as normal noise, is a random noise whose statistical characteristics conform to the normal distribution. Its characteristic is that the noise value added to each pixel point in the image is randomly drawn from the same normal distribution. The mathematical expression of Gaussian noise is N(mu, sigma^2), where mu is the mean (usually configured to 0), and sigma is the standard deviation, which is used to control the intensity of the noise. By adding Gaussian noise to the training data, the diversity of the samples is increased, so that the subsequent training and learning can be robust to small difference changes and interferences, and the generalization of the model is improved.
[0091] The steps of randomly sampling Gaussian noise include:
[0092] Determine the noise parameters: Select the mean and standard deviation. The mean can be set to 0, and the standard deviation is adjusted according to the required noise intensity;
[0093] Generate a noise matrix: For each pixel of the image, draw a value from a Gaussian distribution N(0, sigma^2) to generate a noise matrix with the same size as the image;
[0094] Apply the noise to the image: Add the noise matrix to the original image features.
[0095] Salt-and-pepper noise, also known as impulse noise, appears as randomly setting the values of some pixels to the highest or lowest values in the image, simulating data errors that may occur during the transmission of digital images. This kind of noise usually appears as randomly distributed black and white dots, hence the name "salt and pepper".
[0096] When constructing salt-and-pepper noise, first determine the noise ratio, that is, set what proportion of pixel points will be modified into noise points. Usually, there are two parameters: p - used to control the proportion of pixels becoming white (the maximum value), q - used to control the proportion of pixels becoming black (the minimum value); then determine the generated noise pattern, that is, randomly decide whether to add noise at each pixel position of the image and the type of noise (black or white), and finally modify the pixel values of the original image according to the generated noise pattern.
[0097] Poisson noise is a type of noise commonly found in image and signal processing, especially in imaging systems under low-light conditions. Especially when the lighting conditions are poor, such as in foggy or cloudy days, the randomness of photons reaching the image sensor causes noise errors. The characteristic of Poisson noise is that its intensity depends on the signal intensity, that is, the variance of the noise is proportional to the average value of the signal.
[0098] During random sampling, it specifically includes the following steps:
[0099] Determine the signal intensity: The value of each pixel represents the signal intensity at that place;
[0100] Generate noise: For each pixel in the image, use its pixel value as the parameter lambda of the Poisson distribution to generate noise;
[0101] Apply the noise: The generated noise directly represents the new value of this pixel because Poisson noise reflects the random fluctuations of actual photon counts.
[0102] {Embodiment 2}
[0103] Combined Figure 2 As shown, according to the present invention, a method for constructing a face image recognition feature library is also proposed, including:
[0104] Step S201, obtain information of multiple monitored object persons, and each monitored object person information includes personal information and personal images;
[0105] Step S202: Extract face images from personal images in the first database using a face detection model to construct a face image library;
[0106] Step S203: Align the extracted face images so that the faces in the face images are at a unified angle and size;
[0107] Step S204: Adjust the face images using at least one of image rotation, flipping, scaling, and brightness, and add each adjusted face image to the face image library;
[0108] Step S205: Extract face features from each face image in the face image library. Each face image generates a feature vector of a fixed length to construct a first face feature library;
[0109] Step S206: For the feature vector corresponding to each face image in the first face feature library, randomly sample one of Gaussian noise, salt-and-pepper noise, and Poisson noise as the noise feature, reverse the noise feature, and generate a random noise vector feature corresponding to the noise signal;
[0110] Step S207: Discriminate between the random noise vector feature and the feature vector, and use the discrimination result as an increment to supplement it to the first face feature library and update the first face feature library until the number of samples in the first face feature library reaches the expected number;
[0111] Step S208: Obtain a public face dataset, which contains diverse face images, including face images of different ages, genders, expressions, and lighting conditions;
[0112] Step S209: Adjust the face images in the public face dataset using at least one of image rotation, flipping, scaling, and brightness, and add each adjusted face image to the face image library;
[0113] Step S210: Extract face features from each face image in the face image library. Each face image generates a feature vector of a fixed length and adds it to the first face feature library;
[0114] Step S211: Use the first face feature library to train a deep learning-based face recognition model for subsequent face recognition tasks.
[0115] It should be understood that in the method for constructing a face image recognition feature library proposed in Embodiment 2 of the present invention, on the basis of Embodiment 1, instead of using the fine-tuning training method, on the basis of expanding the face feature library of the determined personnel range, the public face database is further combined to jointly construct the face feature library. On this basis, a deep learning-based face recognition model is trained to achieve the purpose of model training.
[0116] It should be understood that the processing of the foregoing steps S201 to S207 is the same as the corresponding processing method in the foregoing Embodiment 1.
[0117] In step S208, the foregoing public face dataset refers to publicly available face image datasets such as LFW, VGGFace, CASIA, and OpenCV, which are available for downloading and use.
[0118] In step S209, at least one of the adjustment methods of image rotation, flipping, scaling, and brightness used can be implemented by using classical image data augmentation operations corresponding to the methods in Embodiment 1.
[0119] It should be understood that in step S210, the face features can be extracted in the same way as in step S205, and the lengths of the obtained feature vectors are the same.
[0120] As an optional embodiment, in step S211, the face recognition model based on deep learning includes but is not limited to a face detection model based on OpenCV, a face detection model based on Dlib, a face recognition model based on CNNs (such as a face recognition model based on VGG-16, a face recognition model based on FaceNet developed by Google, a face recognition model based on DeepFace developed by Facebook, etc.), a face recognition model based on RNNs (such as a face recognition model developed based on RNN, LSTM, GRU, etc.), a face recognition model based on ensemble learning, etc.
[0121] It should be understood that whether in the foregoing Embodiment 1 or Embodiment 2, a security key monitoring library for uniformly storing the personnel to be focused on and monitored supports periodically updating the content of the database, that is, dynamically adding the information of the monitored personnel.
[0122] Based on the dynamically updated security key monitoring library, for the monitoring personnel information in the added part, the corresponding face feature library can be dynamically updated for the incremental data, and the model can be further fine-tuned based on the dynamic incremental part to achieve dynamic update.
[0123] It should be understood that for security systems deployed in different ways, the face recognition model can be updated by means of online upgrade or offline upgrade, etc.
[0124] {Embodiment 3}
[0125] Combined with the method for constructing a face image recognition feature library in the above embodiments, according to the present invention, a computer system is further proposed, including: one or more processors; and a memory that stores operable instructions.
[0126] Wherein, when the foregoing instructions are executed by one or more processors, the one or more processors are caused to perform operations including the process of the foregoing method for constructing a face image recognition feature library.
[0127] {Embodiment 4}
[0128] Combined with the method for constructing a face image recognition feature library in the above embodiments, according to the present invention, a computer-readable medium storing software is further provided. The foregoing software includes instructions that can be executed by one or more computers. When these instructions are executed by one or more computers, they perform the process of the method for constructing a face image recognition feature library in the foregoing Embodiment 1, Embodiment 2, and the embodiments deformed therefrom.
[0129] Although the present invention has been disclosed above with preferred embodiments, it is not intended to limit the present invention. Those of ordinary skill in the technical field to which the present invention pertains can make various modifications and refinements without departing from the spirit and scope of the present invention. Therefore, the protection scope of the present invention shall be determined by the scope defined by the claims.
Claims
1. A method for constructing a facial image recognition feature library, characterized in that: include: Acquire multiple monitored person information, each monitored person information including personal information and personal image; Extracting facial images from personal images in the first database using a face detection model to construct a facial image library; Align the extracted face images so that the faces in the face images are at a uniform angle and size; Adjust the facial image using at least one of image rotation, flipping, scaling, and brightness, and add each adjusted facial image to a facial image library; Extracting facial features from each facial image in the facial image library, generating a feature vector of a fixed length for each facial image, and constructing a first facial feature library; For the feature vector corresponding to each face image in the first face feature library, randomly sample one of Gaussian noise, salt and pepper noise, and Poisson noise as a noise feature, perform inverse engineering on the noise feature, and generate a random noise vector feature corresponding to the noise signal; Discriminate the random noise vector feature and the feature vector, and use the result as an increment according to the discrimination result, add it to the first face feature library and update the first face feature library; Using the updated first facial feature library to fine-tune the trained face recognition model to obtain an updated face recognition model; The step of discriminating the random noise vector feature from the feature vector and using the discriminating result as an increment to supplement the first face feature library and update the first face feature library includes: An adversarial domain adaptation model is used for discrimination, with random noise vector features as the target domain and feature vectors as the source domain, both of which have a unified feature dimension; the adversarial domain adaptation model includes a domain classifier and a feature generator, wherein the feature generator attempts to generate features that cannot be distinguished by the domain classifier as being from the source domain or the target domain, thereby forcing the generated features to have domain invariance; adversarial training is used to encourage the feature generator to generate features that can deceive the domain classifier, and the random noise vector features discriminated as the source domain are used as increments to supplement the first face feature library.
2. The method for constructing a facial image recognition feature library according to claim 1, characterized in that: The length of the fixed-length feature vector is preconfigured and is greater than or equal to 108.
3. The method for constructing a facial image recognition feature library according to claim 1, characterized in that: The step of randomly sampling one of Gaussian noise, salt and pepper noise, and Poisson noise as a noise feature for the feature vector corresponding to each face image in the first face feature library, and reversing the noise feature to generate a random noise vector feature corresponding to the noise signal includes: For the feature vector corresponding to each face image in the first face feature library, Gaussian noise, salt and pepper noise, and Poisson noise functions are constructed, and one of the three types of Gaussian noise, salt and pepper noise, and Poisson noise is randomly sampled to construct a random noise vector feature. Reverse engineering is performed through the diffusion model DM to generate a random noise vector feature corresponding to the noise signal.
4. The method for constructing a facial image recognition feature library according to claim 3, characterized in that: The diffusion model DM is implemented by neural network training so that it can effectively reverse the noise process in the reverse process. The root mean square error RMSE and MSE objective functions are used in the training process to make the model converge and minimize the difference between the original feature vector and the reverse restored data.
5. The method for constructing a facial image recognition feature library according to claim 1, characterized in that: The adversarial domain adaptation model adopts a DANN deep adversarial neural network model.
6. The method for constructing a facial image recognition feature library according to claim 1, characterized in that: In the method, whether to continue randomly sampling noise features is determined by judging whether the number of samples in the first facial feature library has reached the expected number. If the expected number is reached, sampling noise features is stopped; otherwise, random sampling of noise signals is continued and new incremental data is generated in reverse to expand the first facial feature library.
7. A method for constructing a facial image recognition feature library, characterized in that: include: Acquire multiple monitored person information, each monitored person information including personal information and personal image; Extracting facial images from personal images in the first database using a face detection model to construct a facial image library; Align the extracted face images so that the faces in the face images are at a uniform angle and size; Adjust the facial image using at least one of image rotation, flipping, scaling, and brightness, and add each adjusted facial image to a facial image library; Extracting facial features from each facial image in the facial image library, generating a feature vector of a fixed length for each facial image, and constructing a first facial feature library; For the feature vector corresponding to each face image in the first face feature library, randomly sample one of Gaussian noise, salt and pepper noise, and Poisson noise as a noise feature, perform inverse engineering on the noise feature, and generate a random noise vector feature corresponding to the noise signal; The random noise vector feature is discriminated from the feature vector, and the result is used as an increment according to the discrimination result, and the first face feature library is added and updated until the samples in the first face feature library reach the expected number; Obtain a public face dataset, wherein the public face dataset contains diverse face images, including face images of different ages, genders, expressions, and lighting conditions; Adjusting the face images in the public face dataset using at least one of image rotation, flipping, scaling, and brightness, and adding each adjusted face image to the face image library; Extracting facial features from each facial image in the facial image library, generating a feature vector of a fixed length for each facial image and adding the feature vector to the first facial feature library; Using the first facial feature library to train a deep learning-based face recognition model for subsequent face recognition task execution; The step of discriminating the random noise vector feature from the feature vector and using the discriminating result as an increment to supplement the first face feature library and update the first face feature library includes: An adversarial domain adaptation model is used for discrimination, with random noise vector features as the target domain and feature vectors as the source domain, both of which have a unified feature dimension; the adversarial domain adaptation model includes a domain classifier and a feature generator, wherein the feature generator attempts to generate features that cannot be distinguished by the domain classifier as being from the source domain or the target domain, thereby forcing the generated features to have domain invariance; adversarial training is used to encourage the feature generator to generate features that can deceive the domain classifier, and the random noise vector features discriminated as the source domain are used as increments to supplement the aforementioned first face feature library.
8. The method for constructing a facial image recognition feature library according to claim 7, characterized in that: The step of randomly sampling one of Gaussian noise, salt and pepper noise, and Poisson noise as a noise feature for the feature vector corresponding to each face image in the first face feature library, and reversing the noise feature to generate a random noise vector feature corresponding to the noise signal includes: For the feature vector corresponding to each face image in the first face feature library, Gaussian noise, salt and pepper noise, and Poisson noise functions are constructed, and one of the three types of Gaussian noise, salt and pepper noise, and Poisson noise is randomly sampled to construct a random noise vector feature. Reverse engineering is performed through the diffusion model DM to generate a random noise vector feature corresponding to the noise signal.
9. The method for constructing a facial image recognition feature library according to claim 8, characterized in that: The diffusion model DM is implemented by neural network training so that it can effectively reverse the noise process in the reverse process. The root mean square error RMSE and MSE objective functions are used in the training process to make the model converge and minimize the difference between the original feature vector and the reverse restored data.
10. A computer system, characterized in that: include: one or more processors; Memory, storing instructions that can be operated; Wherein, when the instruction is executed by one or more processors, the aforementioned one or more processors perform an operation, and the operation includes the process of executing the method for constructing a facial image recognition feature library as described in any one of claims 1-9.