Training Method, Device, Electronic Device and Storage Medium of Face Recognition Model
In the deep learning face recognition model, the source domain face features and target domain face image samples are used for incremental training and model parameters are adjusted, and the problem of poor performance in the target domain is solved, and efficient face recognition under limited storage space is achieved.
Patent Information
- Application Number
- CN202111637930.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-29
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2041-12-29
AI Technical Summary
In the prior art, the deep learning-based face recognition model has poor performance in the target domain, and it is difficult to maintain the recognition performance of the source domain and improve the recognition ability of the target domain when there is only a small amount of storage space.
By acquiring the face features of the source domain and initializing the recognition model, the target face image samples of the target domain are used to adjust some model parameters until the model converges, forming a face recognition model suitable for the source domain and the target domain. The method includes filtering the full amount of face image samples, fixing some model parameters, using the target face image samples and the source domain face features for incremental training, and combining classification loss, constraint loss and pseudo-label for model adjustment.
When the face data of the source domain and the target domain cannot be trained at the same time, the recognition ability of the source domain is maintained, and the recognition ability of the target domain is improved, and the overall recognition ability and accuracy of the face recognition model are improved.
Smart Images

Figure CN114333013B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of deep learning, and particularly to a method, apparatus, electronic device, and storage medium for training a face recognition model. Background Art
[0002] With the rapid development of computer and deep learning technologies, deep learning models are increasingly widely used in the field of face recognition. Currently, face recognition models based on deep learning need to be trained using all face data. After training, the face recognition model can perform identity recognition based on the facial feature information of a person. However, the performance of a face recognition model trained using source domain data is poor in the target domain.
[0003] Due to the catastrophic forgetting characteristic of model training, a face recognition model will forget its recognition performance in the source domain after being trained solely using target domain data.
[0004] In the case where only a small amount of storage space is available to store a small amount of source domain data, how to train a face recognition model in the target domain so that the trained face recognition model can not only maintain its performance in the source domain but also improve its performance in the target domain to enhance the recognition ability and accuracy of the face recognition model is an urgent problem to be solved. Summary of the Invention
[0005] The purpose of the embodiments of the present invention is to provide a method, apparatus, electronic device, and storage medium for training a face recognition model, which can maintain the performance of the face recognition model in the source domain and improve its performance in the target domain. The specific technical solutions are as follows:
[0006] In a first aspect, the embodiments of the present invention provide a method for training a face recognition model, the method comprising:
[0007] Obtain source domain face features and initialize a recognition model, wherein the initialized recognition model is trained based on all face image samples in the source domain, and the source domain face features are the face features of the all face image samples obtained through the initialized recognition model;
[0008] Obtain target face image samples in the target domain, wherein the identity labels corresponding to the target face image samples are unknown;
[0009] Based on the target face image samples and the source domain face features, adjust some model parameters of the initialized recognition model until the initialized recognition model converges, to obtain a face recognition model for the source domain and the target domain.
[0010] Optionally, the step of obtaining source domain face features includes:
[0011] According to a preset screening strategy, screen the full set of face image samples to obtain the screened full set of face image samples. Among them, the preset screening strategy ensures that the number of identity information corresponding to the screened full set of face image samples is not less than a preset number while the number of the screened full set of face image samples remains unchanged;
[0012] Input the screened full set of face image samples into the initialized recognition model to obtain the face features output by the intermediate layer of the initialized recognition model;
[0013] Determine the source domain face features based on the face features output by the intermediate layer.
[0014] Optionally, the initialized recognition model includes a part with fixed parameters and a part to be trained;
[0015] The step of adjusting part of the model parameters of the initialized recognition model based on the target face image samples and the source domain face features includes:
[0016] Input the target face image samples into the part with fixed parameters and the part to be trained to obtain the first predicted labels, and determine the first classification loss based on the first predicted labels and the pseudo-labels corresponding to the target face image samples;
[0017] Input the source domain face features into the part to be trained to obtain the second predicted labels, and determine the second classification loss based on the second predicted labels and the identity labels corresponding to the source domain face features;
[0018] Input the source domain face features into the part to be trained and the corresponding initial part of the part to be trained respectively to obtain the estimated features and the initial features, and determine the constraint loss based on the estimated features and the initial features, where the initial part is the model part corresponding to when the model parameters of the part to be trained are fixed to the model parameters trained based on the full set of face image samples;
[0019] Adjust the model parameters of the part to be trained based on the first classification loss, the second classification loss and the constraint loss.
[0020] Optionally, before the step of inputting the target face image samples into the part with fixed parameters and the part to be trained to obtain the first predicted labels, the method further includes:
[0021] Cluster the target face image samples to determine the pseudo-labels corresponding to each target face image sample, where the pseudo-labels are used to identify the personnel identities to which the corresponding target face image samples belong.
[0022] Optionally, the step of adjusting the model parameters of the part to be trained based on the first classification loss, the second classification loss, and the constraint loss includes:
[0023] Calculate the loss function value L based on the first classification loss, the second classification loss, and the constraint loss according to the following formula:
[0024] L = L c1 + L c2 + λL kd
[0025] where L c1 is the first classification loss, L c2 is the second classification loss, L kd is the constraint loss, and λ is a preset parameter;
[0026] Adjust the model parameters of the part to be trained based on the loss function value.
[0027] Optionally, the step of determining the constraint loss based on the predicted feature and the initial feature includes,
[0028] Calculate the constraint loss L based on the predicted feature and the initial feature according to the following formula kd :
[0029]
[0030] where n is the number of source domain face features, F i is the initial feature corresponding to the i-th source domain face feature, is the predicted feature corresponding to the i-th source domain face feature.
[0031] Optionally, the step of determining the source domain face features based on the face features output by the intermediate layer includes:
[0032] Perform dimensionality reduction on the face features output by the intermediate layer to obtain the dimensionality-reduced face features as the source domain face features;
[0033] Before the step of adjusting the partial model parameters of the initialized recognition model based on the target face image sample and the source domain face features, the method further includes:
[0034] Perform dimensionality restoration on the source domain face features to obtain the restored source domain face features.
[0035] Optionally, the method further includes:
[0036] Obtain the face image to be recognized in the target domain;
[0037] Identify the to-be-identified face image based on the face recognition model to determine the identity corresponding to the to-be-identified face image.
[0038] In a second aspect, an embodiment of the present invention provides a training device for a face recognition model, the device includes:
[0039] An initialization training module, configured to obtain source domain face features and initialize a recognition model, wherein the initialized recognition model is trained based on the full amount of face image samples in the source domain, and the source domain face features are the face features of the full amount of face image samples obtained through the initialized recognition model;
[0040] A target domain sample acquisition module, configured to obtain target face image samples in the target domain, wherein the identity labels corresponding to the target face image samples are unknown;
[0041] An incremental training module, configured to adjust some model parameters of the initialized recognition model based on the target face image samples and the source domain face features until the initialized recognition model converges, to obtain a face recognition model for the source domain and the target domain.
[0042] Optionally, the initialization training module includes:
[0043] A sample screening unit, configured to screen the full amount of face image samples according to a preset screening strategy to obtain the screened full amount of face image samples, wherein the preset screening strategy enables the number of identity information corresponding to the screened full amount of face image samples to be not less than a preset number while the number of the screened full amount of face image samples remains unchanged;
[0044] A feature acquisition unit, configured to input the screened full amount of face image samples into the initialized recognition model to obtain the face features output by the middle layer of the initialized recognition model;
[0045] A feature determination unit, configured to determine source domain face features based on the face features output by the middle layer.
[0046] Optionally, the initialized recognition model includes a parameter fixed part and a to-be-trained part;
[0047] The incremental training module includes:
[0048] A first input unit, configured to input the target face image samples into the parameter fixed part and the to-be-trained part to obtain a first predicted label, and determine a first classification loss based on the first predicted label and the pseudo label corresponding to the target face image samples;
[0049] A second input unit, configured to input the source domain face features into the part to be trained, obtain a second prediction label, and determine a second classification loss based on the second prediction label and the identity label corresponding to the source domain face features;
[0050] A third input unit, configured to input the source domain face features into the part to be trained and the corresponding initial part of the part to be trained respectively, obtain a predicted feature and an initial feature, and determine a constraint loss based on the predicted feature and the initial feature, where the initial part is the model part corresponding to when the model parameters of the part to be trained are fixed to the model parameters after being trained based on the full amount of face image samples;
[0051] A parameter adjustment unit, configured to adjust the model parameters of the part to be trained based on the first classification loss, the second classification loss, and the constraint loss.
[0052] Optionally, the apparatus further includes:
[0053] A target domain sample clustering module, configured to cluster the target face image samples before the step of inputting the target face image samples into the parameter fixed part and the part to be trained to obtain a first prediction label, and determine a pseudo label corresponding to each target face image sample, where the pseudo label is used to identify the personnel identity to which the corresponding target face image sample belongs.
[0054] Optionally, the parameter adjustment unit includes:
[0055] A loss function value calculation sub-unit, configured to calculate a loss function value L based on the first classification loss, the second classification loss, and the constraint loss according to the following formula:
[0056] L = L c1 + L c2 + λL kd
[0057] where L c1 is the first classification loss, L c2 is the second classification loss, L kd is the constraint loss, and λ is a preset parameter;
[0058] A parameter adjustment sub-unit, configured to adjust the model parameters of the part to be trained based on the loss function value.
[0059] Optionally, the third input unit includes:
[0060] A constraint loss calculation sub-unit, configured to calculate the constraint loss L based on the predicted feature and the initial feature according to the following formula kd :
[0061]
[0062] wherein, n is the number of the source domain face features, and F i is the initial feature corresponding to the i-th source domain face feature, and is the predicted feature corresponding to the i-th source domain face feature.
[0063] Optionally, the feature determination unit includes:
[0064] a feature dimensionality reduction sub-unit, configured to perform dimensionality reduction processing on the face features output by the intermediate layer to obtain the face features after dimensionality reduction as the source domain face features;
[0065] The apparatus further includes:
[0066] a feature restoration module, configured to perform dimensionality restoration processing on the source domain face features to obtain the restored source domain face features before the step of adjusting some model parameters of the initialization recognition model based on the target face image sample and the source domain face features.
[0067] Optionally, the apparatus further includes:
[0068] a face image to be recognized acquisition module, configured to acquire the face image to be recognized in the target domain;
[0069] an identity determination module, configured to recognize the face image to be recognized based on the face recognition model and determine the identity corresponding to the face image to be recognized.
[0070] In a third aspect, an embodiment of the present invention provides an electronic device, including a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory complete communication with each other through the communication bus;
[0071] The memory is used to store a computer program;
[0072] The processor, when executing the program stored in the memory, implements the method steps of any one of the above first aspects.
[0073] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium, in which a computer program is stored, and when the computer program is executed by a processor, the method steps of any one of the above first aspects are implemented.
[0074] Advantageous effects of the embodiments of the present invention:
[0075] In the solution provided by the embodiments of the present invention, an electronic device can obtain source domain face features and initialize a recognition model. Among them, the initialized recognition model is trained based on all face image samples in the source domain, and the source domain face features are the face features of all face image samples obtained through the initialized recognition model; obtain target face image samples in the target domain, where the identity labels corresponding to the target face image samples are unknown; based on the target face image samples and the source domain face features, adjust some model parameters of the initialized recognition model until the initialized recognition model converges, and obtain a face recognition model for the source domain and the target domain. After the initialized model is trained using all face image samples in the source domain, some source domain face features are saved, and some parameters of the initialized model are fixed. Furthermore, after further training the initialized model using the target face image samples in the target domain and the source domain face features, a face recognition model for the source domain and the target domain is obtained. This face recognition model not only maintains the recognition ability for all face images in the source domain, but also can accurately recognize the target face images in the target domain. In the case where the face data of the source domain and the target domain cannot be used simultaneously to train the face recognition model, it can accurately recognize face images, improving the recognition ability and accuracy of the face recognition model. Of course, any product or method implementing the present invention does not necessarily need to achieve all the above-mentioned advantages at the same time. BRIEF DESCRIPTION OF THE DRAWINGS
[0076] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention, and those of ordinary skill in the art can also obtain other embodiments based on these drawings.
[0077] Figure 1 It is a flowchart of a method for training a face recognition model provided by an embodiment of the present invention;
[0078] Figure 2 Based on Figure 1 It is a specific flowchart of obtaining source domain face features in step S101 in the shown embodiment;
[0079] Figure 3 Based on Figure 1 It is a specific flowchart of adjusting some model parameters of the initialized recognition model in step S103 in the shown embodiment;
[0080] Figure 4 Based on Figure 1 It is a schematic diagram of training the initialized recognition model using target face image samples and source domain face features in the shown embodiment;
[0081] Figure 5 Based onFigure 1 A schematic diagram of feature extraction and feature dimensionality reduction processing of the illustrated embodiment;
[0082] Figure 6 Based on Figure 1 A specific flowchart for determining the identity of a face image based on the illustrated embodiment;
[0083] Figure 7 A schematic structural diagram of a training device for a face recognition model provided by an embodiment of the present invention;
[0084] Figure 8 Based on Figure 7 A specific schematic structural diagram of the incremental training module of the illustrated embodiment;
[0085] Figure 9 A schematic structural diagram of an electronic device provided by an embodiment of the present invention. Specific embodiments
[0086] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art based on the present invention belong to the scope of protection of the present invention.
[0087] In order to enable the trained face recognition model to maintain the performance of the source domain and improve its performance in the target domain, and improve the recognition ability and accuracy of the face recognition model in the scenario where it is impossible to train the face recognition model using the face data of the source domain and the target domain at the same time, the embodiments of the present invention provide a training method, device, electronic device, computer-readable storage medium, and computer program product for a face recognition model. First, a training method for a face recognition model provided by the embodiments of the present invention will be introduced below.
[0088] The training method for the face recognition model provided by the embodiments of the present invention can be applied to any electronic device capable of training a face recognition model. For example, it can be various computing devices for model training, the servers corresponding to the entrance and exit gates of various parks, the servers of the face recognition devices in the security system, etc., which are not specifically limited here. For the sake of clear description, it will be referred to as an electronic device hereinafter.
[0089] As Figure 1 shown, a training method for a face recognition model, the method includes:
[0090] S101, obtaining source domain face features and initializing the recognition model;
[0091] Among them, the initialized recognition model is trained based on the full-face image samples of the source domain, and the source-domain face features are the face features of the full-face image samples obtained through the initialized recognition model.
[0092] S102. Obtain the target face image samples of the target domain;
[0093] Among them, the identity labels corresponding to the target face image samples are unknown.
[0094] S103. Based on the target face image samples and the source-domain face features, adjust some model parameters of the initialized recognition model until the initialized recognition model converges, and obtain a face recognition model for the source domain and the target domain.
[0095] It can be seen that in the solution provided by the embodiment of the present invention, the electronic device can obtain the source-domain face features and the initialized recognition model. Among them, the initialized recognition model is trained based on the full-face image samples of the source domain, and the source-domain face features are the face features of the full-face image samples obtained through the initialized recognition model; obtain the target face image samples of the target domain, where the identity labels corresponding to the target face image samples are unknown; based on the target face image samples and the source-domain face features, adjust some model parameters of the initialized recognition model until the initialized recognition model converges, and obtain a face recognition model for the source domain and the target domain. After the initialized model is trained with the full-face image samples of the source domain, it saves some source-domain face features and fixes some parameters of the initialized model. Furthermore, after further training the initialized model with the target face image samples of the target domain and the source-domain face features, a face recognition model for the source domain and the target domain is obtained. This face recognition model not only maintains the recognition ability for the full-face image of the source domain, but also can accurately recognize the target face image of the target domain. In the case where the face data of the source domain and the target domain cannot be used to train the face recognition model at the same time, it can accurately recognize the face image, improving the recognition ability and accuracy of the face recognition model.
[0096] The face recognition performance of the deep learning-based face recognition model in conventional scenarios is already relatively excellent. However, in some special scenarios, such as child face recognition, low-quality face recognition, and face recognition with masks, etc., there is still room for improvement in face recognition performance. For example, the convolutional neural network is a commonly used deep learning network for face recognition. When the convolutional neural network is trained, due to issues such as data privacy or training resources, face images in different scenarios cannot be trained simultaneously. The training of the convolutional neural network also has the characteristic of catastrophic forgetting. Catastrophic forgetting means that a face recognition model that has already obtained some face recognition capabilities through training forgets or loses some of the face recognition capabilities obtained previously when learning to recognize new face images.
[0097] In order for a face recognition model to accurately recognize face images, it is necessary to train the model using all face data. However, due to reasons such as data privacy, it may not be possible to obtain all face data. Using such incomplete face data to train the model, the face recognition performance of the obtained face recognition model will be very poor. For example, if the face recognition model is a model based on a convolutional neural network and the face data is face images of people wearing masks, due to the characteristic of catastrophic forgetting, if only these face images of people wearing masks are used for conventional training of the face recognition model, the obtained face recognition model can hardly recognize face images of people not wearing masks, and the recognition ability and accuracy are very poor.
[0098] In the case where it is not possible to train the face recognition model using face data from both the source domain and the target domain at the same time, in order to accurately recognize face images in the target domain, an electronic device can use the method of domain adaptation incremental learning to train the face recognition model. In the embodiments of the present invention, the face data used to train the face recognition model can be divided into full-face image samples in the source domain and target face image samples in the target domain. Among them, the full-face image samples are face images with known identity information; the target face image samples in the target domain are face image samples of the person to be recognized, and the identity labels corresponding to the target face image samples are unknown.
[0099] Domain adaptation is a transfer learning method where the corresponding data distributions in the source domain and the target domain are different but the tasks are the same. For example, in this embodiment, although the full-face image samples and the target face image samples have different distributions, the tasks are both used to train the face recognition model. Incremental learning refers to a learning method that can continuously learn new knowledge from new samples and can retain most of the knowledge that has been learned. For example, in this embodiment, after training an initial model using the full-face image samples in the source domain, further training the initial model using the target face image samples in the target domain to obtain a face recognition model, the face recognition model can retain the recognition ability for face images in the source domain and also enhance the recognition ability for target face images in the target domain.
[0100] In the above step S101, the electronic device can obtain the source domain face features and initialize the recognition model. Among them, the initialized recognition model is trained based on the full set of face image samples in the source domain. The full set of face image samples can be CASIA, VGGFace2, MS1MV2, etc., which are not limited here. The identity information of the full set of face image samples in the source domain is known, that is, the identity labels of the full set of face image samples are known. When training the initialized model, after inputting the full set of face image samples in the source domain into the initialized model, the initialized model can obtain the predicted labels of the full set of face image samples. Based on the predicted labels and the identity labels of the full set of face image samples, the electronic device can calculate the classification loss according to the following formula:
[0101]
[0102] where L c is the cross-entropy loss used, m represents the size of the margin, s represents the size of the scale value, θ represents the angle between the weight and the feature, i represents the index of the input sample, and y i represents the label corresponding to the input sample with index i, and N represents the number of samples.
[0103] Furthermore, based on the classification loss calculated by the above formula, the electronic device can continuously reduce the classification loss by adjusting the model parameters of the face recognition model until the number of iterations of the full set of face image samples in the source domain reaches the preset number, determine that the initialized model converges, and obtain the initialized recognition model. Of course, it is also reasonable to determine that the initialized model converges and obtain the initialized recognition model based on the convergence of the loss function of the initialized model. In this way, the trained initialized recognition model has the ability to recognize face images in the source domain, that is, it has the ability to recognize the faces of ordinary people.
[0104] In order to enable the face recognition model to maintain the ability to recognize face images in the source domain after incremental training, the face features of the full set of face image samples in the source domain can be extracted, and the initialized recognition model can be trained with the source domain face features during the incremental training process, so that the initialized recognition model can better retain the ability to recognize face images in the source domain. In one implementation, the electronic device can only store a small amount of source domain face features. For example, in the case of only a small amount of storage space, the electronic device can input a randomly selected part of the full set of face image samples into the initialized recognition model, and obtain the face features output by the initialized recognition model as the source domain face features.
[0105] The initialized recognition model obtained through initialized training already has the ability to recognize source domain face images. To enhance the ability to recognize target domain face images, the initialized recognition model can be incrementally trained using target face image samples of the target domain. After the electronic device obtains the source domain face features and the initialized recognition model, it can obtain the target face image samples of the target domain, that is, execute the above step S102.
[0106] The target face image samples of the target domain can be collected by the electronic device or input into the electronic device by an external device, which is not limited here. The identity labels corresponding to the target face image samples are unknown, that is, the specific identity information corresponding to the target face image samples of the target domain cannot be determined. For example, the target face image samples of the target domain obtained by the electronic device are the face images of multiple employees in an industrial park. Due to privacy protection, the true identities of these face images cannot be determined. The electronic device can perform clustering operations on these face images and record the labels of each category of face images as "Employee A", "Employee B", etc. respectively, so as to make judgments on the identities of personnel during the face recognition process.
[0107] Furthermore, in the above step S103, the electronic device can adjust some model parameters of the initialized recognition model based on the target face image samples and the source domain face features until the initialized recognition model converges, and obtain a face recognition model for the source domain and the target domain.
[0108] To maintain the ability of the initialized recognition model to recognize source domain face images and enhance the ability to recognize target face images of the target domain, the initialized recognition model can include a parameter-fixed part and a parameter-adjustable part. The model parameters of the parameter-fixed part remain unchanged, so that the ability of the initialized recognition model to recognize source domain face images can be maintained.
[0109] Furthermore, the electronic device can input the target face image samples of the target domain and the source domain face features into the initialized recognition model and train the initialized recognition model. The electronic device can adjust some model parameters of the initialized recognition model, that is, the model parameters of the parameter-adjustable part, until the initialized recognition model converges, and obtain a face recognition model for the source domain and the target domain. The face recognition model for the source domain and the target domain will then have the ability to recognize target face images of the target domain.
[0110] By adopting the solution provided by the embodiment of the present invention, the electronic device can use the full-face image samples of the source domain to train the face recognition model, and the obtained initialized recognition model has the ability to recognize the face images of the source domain. During the process of training the initialized recognition model with the target face image samples of the target domain, by fixing some parameters of the initialized recognition model and retraining the initialized recognition model with the source domain face features, the obtained face recognition model for the source domain and the target domain not only maintains the ability to recognize the face images of the source domain, but also can accurately recognize the target face images of the target domain. In the case where it is impossible to train the face recognition model with the face data of both the source domain and the target domain at the same time, only a small number of source domain face features can be used to obtain the face recognition model for the source domain and the target domain. The trained face recognition model can not only maintain the performance of the source domain, but also improve its performance in the target domain, can accurately recognize face images, and improves the recognition ability and accuracy of the face recognition model.
[0111] As an implementation manner of the embodiment of the present invention, as Figure 2 shown, the above step of obtaining the source domain face features may include:
[0112] S201, screening the full-face image samples according to a preset screening strategy to obtain the screened full-face image samples;
[0113] Among them, the preset screening strategy can make the number of identity information corresponding to the screened full-face image samples not less than the preset number when the number of the screened full-face image samples remains unchanged.
[0114] During the initialization training process, the data scale of the full-face image samples of the source domain for training the face recognition model is usually very large, which is not conducive to storage. And in the step of adjusting some model parameters of the initialized recognition model based on the target face image samples and the source domain face features, it is not necessary to use the source domain face features corresponding to all the full-face image samples of the source domain, and the ability of the initialized recognition model to recognize the face images of the source domain can also be maintained. Therefore, in order to reduce the storage space required for data storage and the computational amount required for face recognition model training, the electronic device can screen the full-face image samples of the source domain.
[0115] For example, for the full-face image samples of the source domain corresponding to each identity information, only three source domain face features included in the full-face image samples need to be extracted to maintain the ability of the initialized recognition model to recognize the face images of the source domain. Then, three full-face image samples corresponding to each identity information can be screened out for subsequent face feature extraction.
[0116] The electronic device can screen the full set of face image samples according to a preset screening strategy to obtain the screened full set of face image samples. When the storage space remains unchanged, that is, the number of the screened full set of face image samples remains unchanged, the higher the inter-class richness of the full set of face image samples, the better the initialized recognition model can maintain the recognition ability for the source domain face images. That is to say, it is possible to obtain as many full set of face image samples corresponding to multiple identity information as possible, and the number of full set of face image samples in different states corresponding to each identity information does not have to be too large, that is, the inter-class richness can be appropriately reduced, so as to better maintain the recognition ability of the face recognition model for the target domain to the source domain face images.
[0117] Therefore, the above preset screening strategy can be a strategy that when the number of the screened full set of face image samples remains unchanged, the number of identity information corresponding to the screened full set of face image samples is not less than a preset number, where the preset number can be set according to requirements such as storage space and computing power, and is not limited here.
[0118] In one implementation, the electronic device can randomly select a certain number of full set of face image samples corresponding to identity information from all the full set of face image samples, and then for each full set of face image samples corresponding to an identity information, randomly select a certain number of full set of face image samples from them. For example, the electronic device can first randomly select 1000 full set of face image samples corresponding to identity information, and then for each full set of face image samples corresponding to an identity information, randomly select 10 full set of face image samples from them.
[0119] In another implementation, for each full set of face image samples corresponding to an identity information, the electronic device can select the full set of face image samples according to the distance between the face features of each full set of face image samples and the feature center. For example, a certain number of full set of face image samples with the closest distance can be selected. Among them, the feature center is the classifier vector corresponding to the identity information to which these full set of face image samples belong. The closer the distance to the feature center, the higher the accuracy of the true face characteristics of the person whose face feature identification of the full set of face image samples belongs to the identity information. Therefore, using such full set of face image samples for subsequent feature extraction can be more beneficial to the recognition ability of the trained face recognition model.
[0120] S202, input the screened full set of face image samples into the initialized recognition model, and obtain the face features output by the middle layer of the initialized recognition model.
[0121] After determining the screened full set of face image samples, the electronic device can input the screened full set of face image samples into the initialized recognition model to obtain the face features output by the middle layer of the initialized recognition model. The formula F can be usedi = f(x i ) is used to represent the feature extraction operation, where F i represents the extracted face features, and x i represents the full-face image sample input to the initialized recognition model. f(x) represents the function based on which feature extraction is performed in the initialized recognition model.
[0122] As an implementation, the initialized recognition model may include multiple layers. For example, the initialized recognition model is a residual network, which includes four residual blocks. During the incremental training of the initialized training, the model parameters of the first three residual blocks can be fixed, and only the model parameters of the fourth residual block are adjusted. Then, the electronic device can obtain the face features output by the third residual block of the initialized recognition model.
[0123] S203. Based on the face features output by the intermediate layer, determine the source-domain face features.
[0124] After the electronic device obtains the face features output by the intermediate layer of the initialized recognition model, it can determine the source-domain face features. In one implementation, the electronic device can use the face features output by this intermediate layer as the source-domain face features. In another implementation, the electronic device can perform dimensionality reduction processing on the face features output by this intermediate layer to obtain the face features after dimensionality reduction processing as the source-domain face features for storage convenience.
[0125] In this embodiment, the electronic device can screen the full-face image samples and determine the source-domain face features based on the screened full-face image samples. The source-domain face features are beneficial to maintaining the recognition ability of the face recognition model for the source-domain face images. By screening the full-face image samples of the source domain, on the basis of not reducing the recognition ability of the face recognition model for the source-domain face images, the data storage space and calculation amount required for training the face recognition model are reduced.
[0126] As an implementation of the embodiment of the present invention, the above-mentioned initialized recognition model may include a parameter-fixed part and a part to be trained.
[0127] In order to maintain the recognition ability of the face recognition model for the source-domain face images, after training the initialized recognition model based on the full-face image samples of the source domain, the initialized recognition model can be divided into a parameter-fixed part and a part to be trained. During the process of training the initialized recognition model using the target face image samples and the source-domain face features, the model parameters of the parameter-fixed part are no longer adjusted to maintain the recognition ability of the face recognition model for the source-domain face images. And the model parameters of the part to be trained are adjusted so that the recognition ability of the trained face recognition model for the target-domain face images is also strong.
[0128] In one embodiment, the source domain face features are the face features output by the intermediate layer of the initialized recognition model. Then, the fixed part of the parameters of the initialized recognition model can be consistent with the source domain face features, that is, the model part for processing the source domain face features. For example, the initialized recognition model is a residual network, which includes four residual blocks, and the source domain face features are output by the third residual block. Then, the fixed part of the parameters of the initialized recognition model can include the first three residual blocks, and the part to be trained includes the fourth residual block and the classifier.
[0129] Correspondingly, as Figure 3 shown, the step of adjusting part of the model parameters of the initialized recognition model based on the target face image sample and the source domain face features may include:
[0130] S301, input the target face image sample into the fixed part of the parameters and the part to be trained to obtain a first predicted label, and determine a first classification loss based on the first predicted label and the pseudo-label corresponding to the target face image sample.
[0131] Based on the target face image sample and the source domain face features, perform incremental training on the initialized recognition model. Specifically, the teacher-student network method can be used for training. The student network, that is, the face recognition model, can use a small amount of stored features, that is, the above-mentioned source domain face features, to perform incremental training on the teacher network, that is, the initialized recognition model, so that the student network obtains a performance similar to that of the teacher network, that is, maintains the recognition ability for the source domain face images. During the incremental training process, use the target face image samples in the target domain to further train the student network so that it can have the recognition ability for the target face images in the target domain.
[0132] Since the identity labels corresponding to the target face image samples in the target domain are unknown, the electronic device can determine the pseudo-label corresponding to each target face image sample after obtaining the target face image samples in the target domain. The pseudo-label is used to identify the personnel identity to which the corresponding target face image sample belongs, but it is not the true identity of the person to whom the target face image sample belongs. For example, the pseudo-labels can be A, B, C, or 11, 12, 13, etc., which are not limited here.
[0133] To improve the recognition ability of the face recognition model for the target domain face images, the electronic device can input each target face image sample into the fixed part of the parameters and the part to be trained to obtain a first predicted label. Furthermore, based on the difference between the first predicted label and the pseudo-label corresponding to the target face image sample, the electronic device can calculate the classification loss corresponding to the target face image sample, that is, the first classification loss. The first classification loss can represent the difference between the recognition result and the true result of the current face recognition model for the face images in the target domain.
[0134] In S302, input the source domain face feature into the part to be trained to obtain a second prediction label, and determine a second classification loss based on the second prediction label and the identity label corresponding to the source domain face feature.
[0135] To maintain the recognition ability of the face recognition model for source domain face images, the electronic device can input the source domain face feature into the part to be trained to obtain a second prediction label output by the part to be trained. Since the identity label of the full set of face image samples corresponding to the source domain face feature is known, the electronic device can calculate the classification loss corresponding to the source domain face feature, that is, the second classification loss, based on the difference between the second prediction label and the identity label corresponding to the source domain face feature. This second classification loss can represent the difference between the recognition result of the current face recognition model for source domain face images and the true result.
[0136] In S303, input the source domain face feature into the part to be trained and the initial part corresponding to the part to be trained respectively to obtain a predicted feature and an initial feature, and determine a constraint loss based on the predicted feature and the initial feature;
[0137] Wherein, the initial part is the model part corresponding when the model parameters of the part to be trained are fixed to the model parameters after being trained based on the full set of face image samples.
[0138] The electronic device can fix the model parameters of the part to be trained. The fixed model parameters are the model parameters of the part to be trained of the initialized recognition model after being trained based on the full set of face image samples. After the model parameters of the part to be trained of the initialized recognition model are fixed, it is the initial part. For example, the initialized recognition model is a residual network, which includes four residual blocks. Fix the model parameters of the first three residual blocks and only adjust the model parameters of the fourth residual block. The part to be trained includes the fourth residual block. Fix the model parameters of the fourth residual block to the parameters of this part corresponding to the initialized model, which is the initial part.
[0139] The electronic device can input the source domain face feature into the part to be trained. The part to be trained can perform further feature extraction on the source domain face feature based on the current model parameters to obtain a predicted feature. Input the source domain face feature into the initial part corresponding to the part to be trained. The initial part can perform further feature extraction on the source domain face feature based on the fixed model parameters to obtain an initial feature.
[0140] Furthermore, the electronic device can calculate a constraint loss based on the difference between the predicted features and the initial features. This constraint loss can represent the difference between the face features extracted from the part to be trained and the face features extracted from the corresponding part with fixed model parameters in the initialization model, and can be used as a constraint condition to supervise the training of the face recognition model.
[0141] S304. Based on the first classification loss, the second classification loss, and the constraint loss, adjust the model parameters of the part to be trained.
[0142] After obtaining the above first classification loss, second classification loss, and constraint loss, the electronic device can adjust the model parameters of the part to be trained based on the first classification loss, the second classification loss, and the constraint loss to train the initialization recognition model until the number of iterations of the target face image sample and the source domain face features reaches a preset number, and determine that the initialization model converges.
[0143] In another implementation, the loss function value of the face recognition model can be calculated based on the first classification loss, the second classification loss, and the constraint loss, and the model parameters of the part to be trained can be adjusted based on this loss function value until the loss function of the face recognition model converges, and it is determined that the initialization recognition model converges.
[0144] Among them, the specific method of adjusting the model parameters can adopt algorithms such as gradient descent algorithm and stochastic gradient descent algorithm, which are not specifically limited and described here.
[0145] Since the first classification loss can represent the difference between the recognition result of the current face recognition model for the face image in the target domain and the true result, the second classification loss can represent the difference between the recognition result of the current face recognition model for the face image in the source domain and the true result, and the constraint loss can represent the difference between the face features extracted from the part to be trained and the face features extracted from the corresponding part with fixed model parameters in the initialization model, so adjusting the model parameters of the part to be trained based on the first classification loss, the second classification loss, and the constraint loss can make the difference between the recognition result of the face recognition model for the face image in the target domain and the true result smaller and smaller, and maintain the accuracy of the recognition result of the face image in the source domain.
[0146] In this embodiment, the electronic device can use the source domain face features to train the initialized recognition model to maintain the recognition ability of the face recognition model for source domain face images. The electronic device can use the target face image samples in the target domain to train the initialized recognition model to improve the recognition ability of the face recognition model for target domain face images. By fixing some model parameters of the initialized recognition model, the recognition ability of the face recognition model for source domain face images is better maintained, and the recognition ability of the face recognition model for the target face images in the target domain is improved.
[0147] The following combines Figure 4 This paper gives an example of the process of adjusting some model parameters of the initialized recognition model based on the target face image samples and source domain face features. Among them, conv is the convolutional layer of the initialized recognition model, bn represents batch normalization of data, relu and tanh are the activation functions used in the activation function layer, residual block is the residual block of the initialized recognition model, the parameter fixed part includes the first three residual blocks, the part to be trained includes the fourth residual block and the classifier, and the initial part is the fourth residual block with fixed model parameters.
[0148] The electronic device can input the target face image samples into the parameter fixed part and the part to be trained to obtain the first predicted label, and determine the first classification loss based on the first predicted label and the pseudo label corresponding to the target face image sample.
[0149] The electronic device can perform dimensionality recovery processing on the source domain face features obtained by dimensionality reduction to obtain the recovered source domain face features, input the recovered source domain face features into the part to be trained to obtain the second predicted label, and determine the second classification loss based on the second predicted label and the identity label corresponding to the source domain face features.
[0150] The electronic device can input the recovered source domain face features into the part to be trained and the corresponding initial part of the part to be trained respectively to obtain the estimated features and the initial features, and determine the constraint loss based on the estimated features and the initial features.
[0151] Furthermore, the electronic device can calculate the loss function of the face recognition model based on the first classification loss, the second classification loss and the constraint loss, adjust the model parameters of the part to be trained until the initialized recognition model converges, and obtain the face recognition model for the source domain and the target domain.
[0152] Since the specific implementation manners of the above-mentioned various processes have been introduced in the above-mentioned embodiments, they will not be elaborated herein. In this embodiment, by fixing some parameters of the initialization recognition model and retraining the initialization recognition model using the source domain face features, the obtained face recognition model for the target domain maintains the recognition ability for the source domain face images and also enhances the recognition ability for the target domain face images. In the case where it is impossible to train the face recognition model using the face data of both the source domain and the target domain at the same time, the face images can be accurately recognized, and the recognition ability and accuracy of the face recognition model are improved.
[0153] As an implementation manner of an embodiment of the present invention, before the step of inputting the target face image sample into the parameter fixing part and the part to be trained to obtain the first prediction label, the above method may further include:
[0154] Clustering the target face image samples to determine the pseudo labels corresponding to each target face image sample.
[0155] Since the identity labels corresponding to the target face image samples in the target domain are unknown, the initialization recognition model cannot use such target face image samples for training. The electronic device can perform clustering processing on the target face image samples in the target domain to determine the pseudo labels corresponding to each target face image sample.
[0156] The clustering processing can classify the target face image samples according to the similarity of face features according to the identities to which they belong, so as to obtain multiple groups of target face image samples. Among them, each group of target face image samples belongs to the same person, and the electronic device can assign pseudo labels to each group of target face image samples to identify the identity of the person to which the corresponding target face image sample belongs.
[0157] In one implementation manner, the k-means++ (K-means) clustering algorithm can be used to cluster the target face image samples to obtain multiple categories and determine the pseudo labels of each target face image sample included in each category. The electronic device can also use methods such as the maximum expectation clustering of the Gaussian mixture model, agglomerative hierarchical clustering, and mean shift clustering to cluster the target face image samples, which are not limited herein.
[0158] For example, the target face image samples in the target domain are the face images of the staff in a certain factory. The target face image samples are clustered, and the target face image samples are divided into multiple groups. Each group of target face image samples is the face image of a factory staff member. The electronic device can determine that the pseudo labels of each group of target face image samples are A, B, C, or 11, 12, 13, etc., which are not limited herein.
[0159] In this embodiment, the electronic device can cluster the target face image samples and determine the pseudo-labels corresponding to each target face image sample, so that even when the true identities of the persons to whom the target face image samples belong are unknown, accurate pseudo-labels can be determined to identify the identities of the persons to whom the corresponding target face image samples belong.
[0160] As an implementation manner of an embodiment of the present invention, the step of adjusting the model parameters of the part to be trained based on the first classification loss, the second classification loss, and the constraint loss may include:
[0161] Based on the first classification loss, the second classification loss, and the constraint loss, the loss function value L is calculated according to the following formula: L = L c1 +L c2 +λL kd ; based on the loss function value, the model parameters of the part to be trained are adjusted.
[0162] Wherein, L c1 is the above-mentioned first classification loss, L c2 is the above-mentioned second classification loss, L kd is the above-mentioned constraint loss, and λ is a preset parameter.
[0163] By summing the first classification loss, the second classification loss, and the constraint loss of the preset parameter, the obtained loss function value can accurately represent the difference between the recognition result and the true result of the face image in the target domain, the difference between the recognition result and the true result of the face image in the source domain, and the difference between the face features extracted by the part to be trained and the face features extracted by the corresponding part with fixed model parameters in the initialization model. Therefore, the electronic device can use the above formula to calculate the loss function value L, and based on the loss function value L, adjust the model parameters of the part to be trained to obtain a face recognition model with strong recognition ability.
[0164] Wherein, the value of the preset parameter λ can be set according to the change of the loss function value during the training process combined with actual experience, and no specific limitation is made here.
[0165] In this embodiment, the electronic device can calculate the loss function value based on the first classification loss, the second classification loss, and the constraint loss according to the formula. And based on the loss function value, the model parameters of the part to be trained are adjusted to make the initialization recognition model converge. Through the above formula, the electronic device can accurately calculate the loss function value, enhance the training effect of the initialization recognition model, improve the recognition effect of the face recognition model on the source domain face images, and also make the face recognition model have better performance in recognizing the target domain face images.
[0166] As an implementation manner of an embodiment of the present invention, the step of determining the constraint loss based on the predicted feature and the initial feature may include:
[0167] Calculate the constraint loss L based on the predicted feature and the initial feature according to the following formula kd :
[0168]
[0169] where n is the number of source domain face features, F i is the initial feature corresponding to the i-th source domain face feature, is the predicted feature corresponding to the i-th source domain face feature.
[0170] By calculating the difference between the initial feature corresponding to each source domain face feature and the predicted feature corresponding to the source domain face feature, the difference between each face feature extracted from the part to be trained and the face feature extracted from the corresponding part with fixed model parameters in the initialization model can be obtained. Furthermore, by calculating the variance of the difference values corresponding to each source domain face feature, the obtained constraint loss can accurately represent the difference degree between the predicted feature and the initial feature. Therefore, the electronic device can calculate the constraint loss using the above formula to ensure that an accurate loss function value can be calculated.
[0171] In this embodiment, the electronic device can calculate the constraint loss of each source domain face feature after model training, and thus calculate the constraint loss of the initialization recognition module. Through the above formula, the electronic device can accurately calculate the constraint loss, and use this constraint loss as a supervision to adjust the model parameters of the part to be trained, and a face recognition model with higher accuracy can be obtained.
[0172] As an implementation manner of an embodiment of the present invention, the step of determining the source domain face feature based on the face feature output by the intermediate layer may include:
[0173] Perform dimensionality reduction processing on the face feature output by the intermediate layer to obtain the face feature after dimensionality reduction as the source domain face feature.
[0174] Since the source domain face feature is the output of the intermediate layer of the initialization recognition model and has a high feature dimension, a large storage space is required to store the source domain face feature. To save storage space, the electronic device can perform dimensionality reduction processing on the face feature output by this intermediate layer, thereby significantly reducing the storage space required to store the source domain face feature while basically maintaining the feature information volume.
[0175] For example, methods such as PCA (Principal Component Analysis) dimensionality reduction can be used to reduce the dimensionality of the face features output by the middle layer, obtaining the face features after dimensionality reduction as the source domain face features. The core operation of PCA dimensionality reduction is SVD (Singular Value Decomposition), and SVD can be represented by the formula: AA T = U∑ 2 U T , where A is the matrix to be decomposed, A T is the transpose matrix of A, U is the left singular matrix of A, U T is the transpose matrix of U, and ∑ is a diagonal matrix containing the corresponding eigenvalues. The PCA dimensionality reduction method can delete certain dimensions with correlations in the original data, maximizing the retention of the information carried by the data while reducing the dimensionality of the data.
[0176] In one implementation, the process of obtaining source domain face features from the full amount of face image samples in the source domain and performing dimensionality reduction processing can be as Figure 5 shown. Among them, the initialized recognition model can be a convolutional neural network. conv is the convolutional layer of the convolutional neural network, bn represents batch normalization of the data, relu and tanh are the activation functions used in the activation function layer, and the residual block is the residual block of the convolutional neural network. The parameter fixed part includes the first three residual blocks, and the module parameters of the fourth residual block can be changed for initializing model training. After inputting the full amount of face image samples in the source domain into the convolutional neural network, the electronic device can extract the face features output by the middle layer of the parameter fixed part, and then through dimensionality reduction processing, obtain the face features after dimensionality reduction.
[0177] Correspondingly, before the step of adjusting some model parameters of the initialized recognition model based on the target face image sample and the source domain face features, the method may further include:
[0178] Performing dimensionality restoration processing on the source domain face features to obtain the restored source domain face features.
[0179] In order to try to maintain the recognition ability of the face recognition model for face images in the target domain, since the source domain face features are obtained through dimensionality reduction processing, before the step of adjusting some model parameters of the initialized recognition model based on the target face image sample and the source domain face features, the electronic device can perform dimensionality restoration processing on the source domain face features to obtain the restored source domain face features, and the dimension of the restored source domain face features is the same as the dimension of the face features output by the middle layer of the parameter fixed part.
[0180] In this embodiment, the electronic device can perform dimensionality reduction processing on the face features output by the intermediate layer, and perform dimensionality restoration processing on the source domain face features after dimensionality reduction processing before the step of adjusting some model parameters of the initialized recognition model based on the target face image sample and the source domain face features. Thus, on the basis of basically not affecting the recognition ability of the face recognition model for the target domain face image, the storage space required for storing the source domain face features can be significantly reduced.
[0181] As an implementation manner of the embodiment of the present invention, as Figure 6 shown, the above method may further include:
[0182] S601, obtain the face image to be recognized in the target domain.
[0183] After training the initialized recognition model to obtain a face recognition model for the target domain, since the face recognition model can recognize the face images in the target domain and still maintains the ability to recognize the face images in the source domain, the electronic device can deploy the face recognition model in the actual application scenario, that is, the target domain scenario, and then obtain the face image to be recognized in the target domain.
[0184] For example, when the face recognition model is used for face recognition of the entrance and exit gates in the park, the original model in the gate can be replaced with this face recognition model. This face recognition model has better performance in face recognition of park personnel and basically maintains the ability to recognize the face images of non-park personnel. Then, when a person wants to enter or exit the gate, the electronic device can collect the face image of the person as the face image to be recognized in the target domain.
[0185] S602, recognize the face image to be recognized based on the face recognition model, and determine the identity corresponding to the face image to be recognized.
[0186] Furthermore, the electronic device can recognize the face image to be recognized based on the face recognition model and determine the identity corresponding to the face image to be recognized. After the electronic device recognizes the face image to be recognized, it can determine the identity identifier corresponding to the face image to be recognized and perform different operations according to different identity identifiers. For example, the electronic device can control the gate to open or remain closed and other actions according to the identity identifier corresponding to the face image to be recognized.
[0187] For example, during training, the pseudo-labels corresponding to the target face image samples are "Person A", "Person B", etc. After the electronic device identifies that the identity corresponding to the face image to be recognized is "Person B", it can determine that this person is a person within the industrial park and has the access permission. Then, the electronic device can control the turnstile to open. If it is identified that the identity corresponding to the face image to be recognized is Liu XX, it can be determined that this person is not a person within the industrial park but a certain person outside the park and does not have the access permission. Then, the electronic device can control the turnstile to remain closed.
[0188] In this embodiment, the electronic device can not only identify the target face images in the target domain but also retain the ability to identify the face images in the source domain. Furthermore, for the personnel corresponding to either the source domain or the target domain, the face image to be recognized can be identified based on the face recognition model, and the identity corresponding to the face image to be recognized can be accurately determined.
[0189] Corresponding to the above-mentioned training method of the face recognition model, an embodiment of the present invention further provides a training device for a face recognition model. Next, a training device for a face recognition model provided by an embodiment of the present invention will be introduced.
[0190] As Figure 7 shown, a training device for a face recognition model, the device includes:
[0191] An initialization training module 701, configured to obtain source domain face features and initialize a recognition model;
[0192] Among them, the initialized recognition model is trained based on the full amount of face image samples in the source domain, and the source domain face features are the face features of the full amount of face image samples obtained through the initialized recognition model.
[0193] A target domain sample acquisition module 702, configured to obtain target face image samples in the target domain;
[0194] Among them, the identity labels corresponding to the target face image samples are unknown.
[0195] An incremental training module 703, configured to adjust partial model parameters of the initialized recognition model based on the target face image samples and the source domain face features until the initialized recognition model converges, so as to obtain a face recognition model for the source domain and the target domain.
[0196] It can be seen that in the solution provided by the embodiment of the present invention, the electronic device can obtain the source domain face features and initialize the recognition model, where the initialized recognition model is trained based on the full amount of face image samples in the source domain, and the source domain face features are the face features of the full amount of face image samples obtained through the initialized recognition model; obtain the target face image samples in the target domain, where the identity labels corresponding to the target face image samples are unknown; based on the target face image samples and the source domain face features, adjust some model parameters of the initialized recognition model until the initialized recognition model converges to obtain a face recognition model for the source domain and the target domain. After the initialized model is trained using the full amount of face image samples in the source domain, it saves some source domain face features and fixes some parameters of the initialized model. Furthermore, after using the target face image samples in the target domain and the source domain face features to further train the initialized model, a face recognition model for the source domain and the target domain is obtained. This face recognition model not only maintains the recognition ability for the full amount of face images in the source domain, but also can accurately recognize the target face images in the target domain. In the case where the face data of the source domain and the target domain cannot be used simultaneously to train the face recognition model, it can accurately recognize face images, improving the recognition ability and accuracy of the face recognition model.
[0197] As an implementation manner of the embodiment of the present invention, the above-mentioned initialization training module 701 may include:
[0198] A sample screening unit, configured to screen the full amount of face image samples according to a preset screening strategy to obtain the screened full amount of face image samples;
[0199] Wherein, the preset screening strategy enables the number of identity information corresponding to the screened full amount of face image samples to be not less than a preset number while the number of the screened full amount of face image samples remains unchanged.
[0200] A feature acquisition unit, configured to input the screened full amount of face image samples into the initialized recognition model to obtain the face features output by the middle layer of the initialized recognition model.
[0201] A feature determination unit, configured to determine the source domain face features based on the face features output by the middle layer.
[0202] As an implementation manner of the embodiment of the present invention, the above-mentioned initialized recognition model includes a parameter fixed part and a part to be trained;
[0203] As Figure 8 shown, the above-mentioned incremental training module 703 may include:
[0204] The first input unit 801 is configured to input the target face image sample into the parameter-fixed part and the part to be trained, obtain a first predicted label, and determine a first classification loss based on the first predicted label and the pseudo-label corresponding to the target face image sample.
[0205] The second input unit 802 is configured to input the source domain face features into the part to be trained, obtain a second predicted label, and determine a second classification loss based on the second predicted label and the identity label corresponding to the source domain face features.
[0206] The third input unit 803 is configured to input the source domain face features into the part to be trained and the initial part corresponding to the part to be trained respectively, obtain an estimated feature and an initial feature, and determine a constraint loss based on the estimated feature and the initial feature;
[0207] Wherein, the initial part is the model part corresponding to when the model parameters of the part to be trained are fixed to the model parameters after being trained based on the full amount of face image samples.
[0208] The parameter adjustment unit 804 is configured to adjust the model parameters of the part to be trained based on the first classification loss, the second classification loss, and the constraint loss.
[0209] As an implementation manner of an embodiment of the present invention, the above device may further include:
[0210] The target domain sample clustering module is configured to cluster the target face image samples before the step of inputting the target face image samples into the parameter-fixed part and the part to be trained to obtain a first predicted label, and determine the pseudo-label corresponding to each target face image sample;
[0211] Wherein, the pseudo-label is used to identify the personnel identity to which the corresponding target face image sample belongs.
[0212] As an implementation manner of an embodiment of the present invention, the above parameter adjustment unit 804 may include:
[0213] The loss function value calculation sub-unit is configured to calculate the loss function value L based on the first classification loss, the second classification loss, and the constraint loss according to the following formula:
[0214] L = L c1 + L c2 + λL kd
[0215] Wherein, L c1 is the first classification loss, L c2 is the second classification loss, L kdis the constraint loss, and λ is a preset parameter.
[0216] A parameter adjustment subunit, configured to adjust the model parameters of the part to be trained based on the loss function value.
[0217] As an implementation manner of an embodiment of the present invention, the above-mentioned third input unit 803 may include:
[0218] A constraint loss calculation subunit, configured to calculate the constraint loss L based on the predicted feature and the initial feature according to the following formula kd :
[0219]
[0220] where n is the number of source domain face features, F i is the initial feature corresponding to the i-th source domain face feature, is the predicted feature corresponding to the i-th source domain face feature.
[0221] As an implementation manner of an embodiment of the present invention, the above-mentioned feature determination unit may include:
[0222] A feature dimensionality reduction subunit, configured to perform dimensionality reduction processing on the face features output by the intermediate layer to obtain the dimensionality-reduced face features as the source domain face features.
[0223] The above-mentioned device may further include:
[0224] A feature restoration module, configured to perform dimensionality restoration processing on the source domain face features before the step of adjusting part of the model parameters of the initialization recognition model based on the target face image sample and the source domain face features to obtain the restored source domain face features.
[0225] As an implementation manner of an embodiment of the present invention, the above-mentioned device may further include:
[0226] An image of a face to be recognized acquisition module, configured to acquire an image of a face to be recognized in the target domain;
[0227] An identity determination module, configured to recognize the image of the face to be recognized based on the face recognition model to determine the identity corresponding to the image of the face to be recognized.
[0228] An embodiment of the present invention further provides an electronic device, as Figure 9 shown, including a processor 901, a communication interface 902, a memory 903, and a communication bus 904. Among them, the processor 901, the communication interface 902, and the memory 903 complete communication with each other through the communication bus 904.
[0229] A memory 903 for storing computer programs;
[0230] A processor 901, when executing the programs stored on the memory 903, implements the method steps described in any of the above embodiments.
[0231] The communication bus mentioned in the above electronic device may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, only a thick line is used in the figure, but it does not mean that there is only one bus or one type of bus.
[0232] The communication interface is used for communication between the above electronic device and other devices.
[0233] The memory may include a Random Access Memory (RAM), and may also include a Non-Volatile Memory (NVM), such as at least one disk memory. Optionally, the memory may also be at least one storage device located far from the aforementioned processor.
[0234] The above-mentioned processor may be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it may also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0235] In another embodiment provided by the present invention, a computer-readable storage medium is also provided. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the steps of the method described in any of the above embodiments.
[0236] In another embodiment provided by the present invention, a computer program product containing instructions is also provided. When it runs on a computer, it causes the computer to execute the method steps described in any of the above embodiments.
[0237] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present invention are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or a data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, a hard disk, a magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)).
[0238] It should be noted that, in this document, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variation thereof is intended to cover a non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not expressly listed, or elements inherent to such process, method, article, or device. Without further limitation, an element defined by the phrase "comprising a..." does not exclude the presence of additional identical elements in the process, method, article, or device comprising the element.
[0239] Each embodiment in this specification is described in a related manner. The same or similar parts among the embodiments can be referred to each other, and the differences between each embodiment and other embodiments are emphasized. In particular, for the embodiments of the device, electronic device, computer-readable storage medium, and computer program product, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can refer to the partial description of the method embodiments.
[0240] The above are only the preferred embodiments of the present invention and are not intended to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention are all included within the protection scope of the present invention.
Claims
1. A training method for a face recognition model, characterized in that, The method includes: Obtaining source domain face features and initializing an identification model, where the initialized identification model is trained based on the full set of face image samples in the source domain, and the source domain face features are the face features of the full set of face image samples obtained through the initialized identification model; the initialized identification model includes a parameter-fixed part and a part to be trained; the parameter-fixed part is the model part in the initialized identification model for processing the full set of face image samples to obtain source domain face features; Obtaining target face image samples in the target domain, where the identity labels corresponding to the target face image samples are unknown; Based on the target face image samples and the source domain face features, adjusting some model parameters of the initialized identification model until the initialized identification model converges, to obtain a face recognition model for the source domain and the target domain; where the some model parameters are the part to be trained in the initialized identification model; during the process of adjusting the some model parameters of the initialized identification model, the target face image samples are processed by the parameter-fixed part and the part to be trained, and the source domain face features are processed by the part to be trained.
2. The method according to claim 1, wherein The step of obtaining the source domain face features includes: According to a preset screening strategy, screening the full set of face image samples to obtain the screened full set of face image samples, where the preset screening strategy makes the number of identity information corresponding to the screened full set of face image samples not less than a preset number while the number of the screened full set of face image samples remains unchanged; Inputting the screened full set of face image samples into the initialized identification model to obtain the face features output by the middle layer of the initialized identification model; Based on the face features output by the middle layer, determining the source domain face features.
3. The method according to claim 1, characterized in that, The initialized identification model includes a parameter-fixed part and a part to be trained; The step of adjusting some model parameters of the initialized identification model based on the target face image samples and the source domain face features includes: Inputting the target face image samples into the parameter-fixed part and the part to be trained to obtain a first predicted label, and determining a first classification loss based on the first predicted label and the pseudo label corresponding to the target face image samples; Inputting the source domain face features into the part to be trained to obtain a second predicted label, and determining a second classification loss based on the second predicted label and the identity label corresponding to the source domain face features; Inputting the source domain face features into the part to be trained and the initial part corresponding to the part to be trained respectively to obtain a predicted feature and an initial feature, and determining a constraint loss based on the predicted feature and the initial feature, where the initial part is the model part corresponding to when the model parameters of the part to be trained are fixed to the model parameters after training based on the full set of face image samples; Based on the first classification loss, the second classification loss and the constraint loss, adjusting the model parameters of the part to be trained.
4. The method according to claim 3, wherein Before the step of inputting the target face image sample into the parameter-fixed part and the part to be trained to obtain the first predicted label, the method further includes: Clustering the target face image samples to determine the pseudo-labels corresponding to each target face image sample, where the pseudo-labels are used to identify the personnel identities to which the corresponding target face image samples belong.
5. The method according to claim 3, wherein The step of adjusting the model parameters of the part to be trained based on the first classification loss, the second classification loss, and the constraint loss includes: Calculating the loss function value L according to the following formula based on the first classification loss, the second classification loss, and the constraint loss: L = L c1 + L c2 + λL kd Among them, L c1 is the first classification loss, L c2 is the second classification loss, L kd is the constraint loss, and λ is a preset parameter; Adjusting the model parameters of the part to be trained based on the loss function value.
6. The method according to claim 3, wherein The step of determining the constraint loss based on the estimated feature and the initial feature includes: Based on the predicted features and the initial features, the constraint loss L is calculated according to the following formula kd :[[]]END]] where n is the number of source domain face features, and F i is the initial feature corresponding to the i-th source domain face feature, and is the estimated feature corresponding to the i-th source domain face feature.
7. The method according to claim 2, wherein The step of determining the source domain face feature based on the face feature output by the intermediate layer includes: Performing dimensionality reduction processing on the face feature output by the intermediate layer to obtain the face feature after dimensionality reduction as the source domain face feature; Before the step of adjusting part of the model parameters of the initialized recognition model based on the target face image sample and the source domain face feature, the method further includes: Performing dimensionality restoration processing on the source domain face feature to obtain the restored source domain face feature.
8. The method according to any one of claims 1-6, characterized in that, The method further includes: Obtaining the face image to be recognized in the target domain; Identifying the face image to be recognized based on the face recognition model to determine the identity corresponding to the face image to be recognized.
9. A training device for a face recognition model, characterized in that, The device includes: An initialization training module, configured to obtain source domain face features and an initialized recognition model, where the initialized recognition model is trained based on all face image samples in the source domain, and the source domain face features are the face features of all face image samples obtained through the initialized recognition model; the initialized recognition model includes a parameter-fixed part and a part to be trained; the parameter-fixed part is the model part of the initialized recognition model for processing all face image samples in the source domain to obtain source domain face features; A target domain sample acquisition module, configured to obtain target face image samples in the target domain, where the identity labels corresponding to the target face image samples are unknown; An incremental training module, configured to adjust part of the model parameters of the initialized recognition model based on the target face image samples and the source domain face features until the initialized recognition model converges to obtain a face recognition model for the source domain and the target domain; where the part of the model parameters is the part to be trained in the initialized recognition model; during the process of adjusting part of the model parameters of the initialized recognition model, the target face image samples are processed by the parameter-fixed part and the part to be trained, and the source domain face features are processed by the part to be trained.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method steps described in any one of claims 1-8 are implemented.
Citation Information
Patent Citations
Image recognition model migration method and device, equipment and storage medium
CN112801236A