Heterogeneous face recognition model training method, face recognition method and device

By pre-training the face recognition model and fine-tuning the parameters of the intermediate network layer on the heterogeneous face dataset, the accuracy problem caused by the mid-span modal differences in heterogeneous face recognition is solved, and the recognition accuracy is improved.

CN120236159APending Publication Date: 2025-07-01JIAXING UPHOTON OPTOELECTRONICS TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311867815.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-29
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

There is a problem of cross-modal difference in heterogeneous face recognition, resulting in low model accuracy and lack of enough paired cross-modal face image training data, making it difficult to collect large quantities of near-infrared images.

Method used

By pre-training the face recognition model based on the visible face image dataset, fixing the parameters of the backbone network, only updating the parameters of the intermediate network layer, and fine-tuning the model using the heterogeneous face dataset to improve the recognition accuracy.

Benefits of technology

Without changing the network structure of the face recognition model, the fine-tuning model is used for the heterogeneous face data set to improve the accuracy of heterogeneous face recognition and solve the problem of low recognition accuracy caused by cross-modal differences.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120236159A_ABST
    Figure CN120236159A_ABST
Patent Text Reader

Abstract

A heterogeneous face recognition model training method, and a face recognition method and apparatus, the method comprising: training a first face recognition network based on a first face image data set to obtain a pre-trained first face recognition model, the first face image data set being a visible light face image data set; the structure of the pre-trained first face recognition model is finely adjusted based on a second face image data set, a second face recognition model is obtained, the number of images in the second face image data set is at least two orders of magnitude smaller than the number of images in the first face image data set, and the second face image data set comprises heterogeneous face images; and training a second face recognition model based on the second face image data set to obtain a trained heterogeneous face recognition model, and in the training process of the second face recognition model, fixing network parameters of a backbone network of the second face recognition model, and updating network parameters of a middle network layer of the second face recognition model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of face recognition technology, and more particularly to a method for training a heterogeneous face recognition model, a face recognition method and device. Background Art

[0002] Face recognition is a technology for identifying a person's identity based on a facial image. First, feature vectors are extracted from the face image, and then the similarity between the vectors is calculated through a similarity metric function. Currently, the mainstream solution is to perform feature extraction based on a convolutional neural network and use a cosine function for similarity calculation.

[0003] Traditional face recognition is generally applied to two-dimensional visible light images. However, in an environment with poor lighting conditions, the quality of visible light images is poor, while near-infrared images are more robust to lighting. Therefore, in related scenarios, based on the easy availability of visible light images and the lighting robustness of near-infrared images, visible light images are used as the base library for registration, and near-infrared imaging technology is used to obtain clear near-infrared images for comparison and recognition. This will involve the problem of heterogeneous recognition between visible light images and near-infrared images.

[0004] One problem with heterogeneous face recognition is that visible light face images and near-infrared light face images come from different modalities, so there are large cross-modal differences in their data sources, which will lead to low accuracy of the trained heterogeneous face recognition model. On the other hand, due to the lack of sufficient paired cross-modal face image training data and the difficulty in collecting a large number of near-infrared images, there are great difficulties. Summary of the Invention

[0005] The present application is proposed to solve the above problems. According to one aspect of the present application, there is provided a method for training a heterogeneous face recognition model, the method comprising: training a first face recognition network based on a first face image dataset to obtain a pre-trained first face recognition model, wherein the first face image dataset is a visible light face image dataset; fine-tuning the network structure of the pre-trained first face recognition model based on a second face image dataset to obtain a second face recognition model, wherein the number of images in the second face image dataset is at least two orders of magnitude smaller than the number of images in the first face image dataset, and the second face image dataset at least includes heterogeneous face images different from visible light face images; training the second face recognition model based on the second face image dataset to obtain a trained heterogeneous face recognition model, wherein during the training of the second face recognition model, the network parameters of the backbone network of the second face recognition model are fixed, and the network parameters of the intermediate network layer of the second face recognition model are updated.

[0006] In one embodiment of the present application, updating the network parameters of the intermediate network layer of the second face recognition model includes: updating the network parameters of the neck network of the second face recognition model.

[0007] In one embodiment of the present application, the method further includes: during the training of the second face recognition model, fixing the network parameters of the batch sample normalization layer of the second face recognition model.

[0008] In one embodiment of the present application, the method further includes: during the training of the second face recognition model, setting the learning rate of the second face recognition model to be at least two orders of magnitude smaller than the learning rate of the first face recognition model.

[0009] In one embodiment of the present application, the learning rate of the second face recognition model is not greater than 0.001.

[0010] In one embodiment of the present application, the initial learning rate of the second face recognition model is from 0.00005 to 0.0002.

[0011] In one embodiment of the present application, fine-tuning the network structure of the pre-trained first face recognition model based on the second face image dataset to obtain a second face recognition model includes: removing the classification layer of the first face recognition model; reconstructing a classification layer based on the number of images in the second face image dataset and adding it to the first face recognition model to obtain the second face recognition model.

[0012] In one embodiment of the present application, the second face image dataset includes the visible light face images and the heterogeneous face images, wherein there are some images in the heterogeneous face images that correspond to the face images with the same identity information as the visible light face images.

[0013] In one embodiment of the present application, the heterogeneous face images include near-infrared face images.

[0014] According to another aspect of the present application, there is provided a face recognition method, the method including obtaining a face image to be recognized; performing recognition on the face image to be recognized based on a trained heterogeneous face recognition model to obtain a face recognition result; wherein, the heterogeneous face recognition model is obtained based on the training method of the above-mentioned heterogeneous face recognition model.

[0015] In one embodiment of the present application, the face image to be recognized includes a heterogeneous face image different from the visible light face image.

[0016] In one embodiment of the present application, the heterogeneous face images include near-infrared face images.

[0017] According to another aspect of the present application, a training device for a heterogeneous face recognition model includes a processor and a memory. A computer program is stored on the memory and run by the processor. When the computer program is run by the processor, the processor is caused to execute the above-mentioned training method for the heterogeneous face recognition model.

[0018] According to another aspect of the present application, a face recognition device is provided. The device includes a processor and a memory. A computer program is stored on the memory and run by the processor. When the computer program is run by the processor, the processor is caused to execute the above-mentioned face recognition method.

[0019] According to another aspect of the present application, a storage medium is provided. A computer program is stored on the storage medium and run by a processor. When the computer program is run by the processor, the processor is caused to execute the above-mentioned training method for the heterogeneous face recognition model or the above-mentioned face recognition method.

[0020] The training method, face recognition method and device for the heterogeneous face recognition model of the present application can fine-tune some network parameters of the face recognition model using a heterogeneous face dataset without changing the network structure of the face recognition model, thereby improving the accuracy of heterogeneous face recognition. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] By describing the embodiments of the present application in more detail in conjunction with the accompanying drawings, the above and other objects, features and advantages of the present application will become more obvious. The drawings are used to provide a further understanding of the embodiments of the present application, and constitute a part of the specification. Together with the embodiments of the present application, they are used to explain the present application and do not constitute a limitation to the present application. In the drawings, the same reference numerals generally represent the same components or steps.

[0022] Figure 1 A schematic block diagram showing an example of an electronic device for implementing the training method and device for a heterogeneous face recognition model according to an embodiment of the present application.

[0023] Figure 2 A schematic flowchart showing the training method for a heterogeneous face recognition model according to an embodiment of the present application.

[0024] Figure 3 A schematic flowchart showing the face recognition method according to an embodiment of the present application.

[0025] Figure 4 A schematic structural block diagram showing the training device for a heterogeneous face recognition model according to an embodiment of the present application.

[0026] Figure 5A schematic structural block diagram of a face recognition device according to an embodiment of the present application is shown.

[0027] Figure 6 A schematic structural block diagram of a storage medium according to an embodiment of the present application is shown. Detailed implementation manners

[0028] In order to make the objectives, technical solutions, and advantages of the present application more apparent, exemplary embodiments according to the present application will be described in detail below with reference to the accompanying drawings. Apparently, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments of the present application. It should be understood that the present application is not limited by the exemplary embodiments described herein. Based on the embodiments of the present application described herein, all other embodiments obtained by those skilled in the art without creative efforts shall fall within the protection scope of the present application.

[0029] First, reference is made to Figure 1 to describe an exemplary electronic device 100 for implementing the training method and device of the heterogeneous face recognition model according to the embodiments of the present invention.

[0030] As Figure 1 shown, the electronic device 100 includes one or more processors 102, one or more storage devices 104, an input device 106, and an output device 108, and these components are interconnected through a bus system 110 and / or other forms of connection mechanisms (not shown). It should be noted that Figure 1 the components and structure of the electronic device 100 shown are only exemplary and non-limiting. According to requirements, the electronic device may also have other components and structures.

[0031] The processor 102 may be a central processing unit (CPU) or other forms of processing units with data processing capabilities and / or instruction execution capabilities, and may control other components in the electronic device 100 to perform desired functions.

[0032] The storage device 104 may include one or more computer program products, and the computer program products may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may include, for example, random access memory (RAM) and / or cache memory, etc. The non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc. One or more computer program instructions may be stored on the computer-readable storage media, and the processor 102 may run the program instructions to implement the client functions (implemented by the processor) and / or other desired functions in the embodiments of the present invention described below. Various application programs and various data may also be stored in the computer-readable storage media, such as various data used and / or generated by the application programs, etc.

[0033] The input device 106 may be a device used by a user to input instructions, and may include one or more of a keyboard, a mouse, a microphone, a touch screen, etc. In addition, the input device 106 may also be any interface for receiving information.

[0034] The output device 108 may output various information (such as images or sounds) to the outside (such as a user), and may include one or more of a display, a speaker, etc. In addition, the output device 108 may also be any other device with an output function.

[0035] Exemplarily, an example electronic device for implementing the training method and apparatus of the heterogeneous face recognition model according to the embodiments of the present invention may be implemented in, such as, an access control system, clock-in attendance, video surveillance, etc.

[0036] Next, reference will be made to Figure 2 to describe the training method 200 of the heterogeneous face recognition model according to the embodiments of the present application. Figure 2 FIG. shows a schematic flowchart of the training method 200 of the heterogeneous face recognition model according to the embodiments of the present application. As Figure 2 shown, the training method 200 of the heterogeneous face recognition model according to the embodiments of the present application may include the following steps:

[0037] In step S210, a first face recognition network is trained based on a first face image dataset to obtain a pre-trained first face recognition model, where the first face image dataset is a visible light face image dataset.

[0038] In step S220, the network structure of the pre-trained first face recognition model is fine-tuned based on the second face image dataset to obtain a second face recognition model. The number of images in the second face image dataset is at least two orders of magnitude smaller than that in the first face image dataset, and the second face image dataset includes at least heterogeneous face images different from visible light face images.

[0039] In step S230, the second face recognition model is trained based on the second face image dataset to obtain a trained heterogeneous face recognition model. During the training of the second face recognition model, the network parameters of the backbone network of the second face recognition model are fixed, and the network parameters of the intermediate network layer of the second face recognition model are updated.

[0040] In an embodiment of the present application, when training a heterogeneous face recognition model, first, a visible light face image dataset (referred to as the first face image dataset) is collected to pre-train a face recognition network to obtain a pre-trained face recognition model (referred to as the first face recognition model); thereafter, another face image dataset with a much smaller number of images than that in the visible light image dataset (referred to as the second face image dataset, which includes heterogeneous face images different from visible light images, such as near-infrared face images) is used to train the face recognition model again. Since there are significant differences between the image data included in the second face image dataset and the first face image dataset (because visible light face images are easily obtained, while heterogeneous face images such as near-infrared face images are difficult to obtain), it is usually necessary to first fine-tune the network structure of the face recognition model to adapt to the change in the number of images. At this time, a second face recognition model is obtained, and the second face image dataset can be used to train it.

[0041] Among them, during the training process, the network parameters of the backbone network of the second face recognition model are fixed, and the network parameters of the intermediate network layer of the second face recognition model are updated. This is because:

[0042] The network structure of a face recognition model usually includes a backbone network, an intermediate network (such as a neck network), and a head network. Among them, the backbone network is the main component of the model, usually a convolutional neural network (CNN) or a residual neural network (Resnet), etc. The backbone network is generally responsible for extracting features of the input image for subsequent processing and analysis. The backbone network usually has many layers and many parameters and can extract high-level feature representations of the image. The intermediate network is the intermediate layer connecting the backbone network and the head network. The main role of the intermediate network is to reduce the dimension or adjust the features from the backbone network to better meet the task requirements. The intermediate network can use convolutional layers, pooling layers, or fully connected layers, etc. The head network is the last layer of the model, usually a classifier or a regressor. The structure of the head network varies according to different tasks. For image classification tasks, a softmax classifier can be used; for object detection tasks, a bounding box regressor and a classifier can be used, etc. In this application, for image classification features, softmax loss classification can be used, or other loss classifications can also be used.

[0043] For a face recognition model, the first layer of the network usually extracts the contours and edges of the face from the original picture, and different convolutional kernels extract different edge features. The second layer of the network combines the edge information learned in the first layer to form local features of the face, such as eyes and mouth. The subsequent layers gradually combine the features of the previous layers to form the appearance of the face. From local to global, from simple to complex, the more layers there are, the more accurate the model effect is. That is the above-mentioned backbone network structure. Through the backbone network structure, highly discriminative face features can be extracted. According to the structural principles of the above intermediate network and head network, by adding or modifying the structures of the intermediate network and the head network, the model can be easily applied to different tasks and datasets, thereby improving the generalization ability and performance of the model. In heterogeneous face recognition, the visible light image and the near-infrared image of the same face are extremely similar but there are differences. The inventor found that the feature differences between the visible light image and the near-infrared image can remove their respective specific parts when reducing the dimension in the intermediate network layer, so as to be mapped to a common latent space and obtain unified features.

[0044] For this reason, in the embodiments of the present application, the face recognition model is pre-trained with a visible light image data set containing a large number of images, so that the backbone network structure of the face recognition model is capable of extracting highly discriminative face features and has strong generalization ability. When the model is fine-tuned on a heterogeneous data set with a small number of images (i.e., a small data set), the model will gradually simulate the distribution of the small data set during fine-tuning training, thus losing the generalization ability trained previously. Therefore, in order to ensure that overfitting does not occur during fine-tuning on a small data set, in the present application, the parameters of the backbone network are fixed, and only the parameters of the intermediate network layer are updated, ensuring that the highly discriminative features obtained by pre-training remain unchanged. At the same time, the characteristics of the heterogeneous images on the small data set are removed in the intermediate network layer, thereby solving the problem of low recognition accuracy caused by cross-modal differences in data sources in heterogeneous face recognition and improving the accuracy of heterogeneous face recognition.

[0045] Generally, the training method of the heterogeneous face recognition model according to the embodiments of the present application can fine-tune some network parameters of the face recognition model using a heterogeneous face data set without changing the network structure of the face recognition model, improving the accuracy of heterogeneous face recognition.

[0046] In the embodiments of the present application, in step S210, the first face recognition network is trained based on the first face image data set to obtain a pre-trained first face recognition model. Among them, the first face image data set is a visible light face image data set. First, before pre-training the first face recognition model, the first face image data set needs to be obtained. The first face image data set can be an open-source face image data set of visible light or other suitable face image data sets, and no specific limitation is made thereto. As a possible implementation manner, an open-source face image data set of visible light is selected because the open-source face image data set does not require self-construction of a face image data set, greatly reducing the development cost. At the same time, the open-source face image data set has been widely used in the face recognition technology in the field of computer vision. The annotation accuracy of the face image data in the data set is high, the type richness of the face image data is relatively high, and in addition, the quantity scale of various face image data in the open-source face image data set is huge. Therefore, selecting an open-source face image data set for face recognition model training is beneficial to improving the training efficiency of the trained model and improving the recognition accuracy of the face recognition model.

[0047] In this embodiment, the first face image dataset may be an open-source large-scale visible-light face image dataset, such as the Glint360K dataset or other suitable face image datasets, which are not specifically limited herein. The Glint360K dataset includes multiple target objects, and each target object corresponds to multiple face images in a single modality, that is, visible-light face images. The Glint360K dataset is a relatively large face recognition dataset in the current face recognition field. Of course, according to actual needs, a face image dataset with an increased or decreased sampling range can be selected. The Glint360K dataset contains more than 17 million face images of 360,000 people, which has a significant improvement in the types of faces and the number of images compared to other face image datasets. Therefore, it is more conducive to the training of face recognition models.

[0048] In this embodiment, before training the first face recognition network model based on the first face image dataset, image preprocessing operations on the face images are also included. Image preprocessing is often affected by different image acquisition environments, such as the brightness of the light source and the performance of the device, resulting in problems such as noise and insufficient contrast. In addition, factors such as the distance and focal length make the size and position of the face in the entire image uncertain. To ensure the consistency of the face size, position, and face image quality in the face image, image preprocessing must be performed. At the same time, the images in conventional face recognition datasets often include the entire human head and part of the environmental background, and the face is often tilted (tilted head, side face). Before inputting such images into the model, it is necessary to crop the "real face" part of the image, remove irrelevant background information, and align the cropped face images before they can be used for training. The face image preprocessing operations include: face detection, face key point detection, face alignment, face cropping, and image normalization.

[0049] Exemplarily, the image preprocessing operation first performs face detection. Face detection refers to detecting the position and bounding box of a face in an image. The goal of this task is to determine whether there is a face in the image and identify its approximate position. Face detection is usually used as a preprocessing step to locate faces in higher-level tasks such as face recognition and expression recognition. Secondly, face key point detection is carried out. After face detection is completed, it is also necessary to use a model to predict the positions of face key points based on the results of face detection. Face key points include the specific coordinates of the center of the left eye, the center of the right eye, the tip of the nose, the left corner of the mouth, and the right corner of the mouth. Among them, the face key point detection model can be a Multi-task Cascaded Convolutional Networks (MTCNN), a Tweaked Convolutional Neural Networks (TCNN), a Tasks-Constrained Deep Convolutional Network (TCDCN), etc. After that, face alignment is required. After obtaining the face key points in the face image, the transformation matrix between these two points can be calculated through the relative position relationship between the face key points in the image and the key points of the standard portrait, and then the scaling factor, rotation angle, and translation parameters between the face in the image and the reference face, as well as the coordinates of the rotated face box, can be obtained. Then, face cropping is carried out. The aligned face image is cropped according to a unified size regulation to obtain a face image of a suitable size for face recognition model training. Among them, the size range of image cropping can be 112*112 pixel size, or it can be an image of other suitable sizes, and no specific limitation is made in this regard. Finally, normalization processing of the image is carried out. Normalization of the face image is to adjust the face image to the same size and resolution for the convenience of subsequent processing of the face recognition model. The normalization of the face image can be in the interval [-1.0, 1.0], or it can be in other suitable intervals, and no specific limitation is made in this regard.

[0050] In this embodiment, the first face recognition model is pre-trained based on the first face image dataset, that is, the heterogeneous face recognition model is pre-trained based on the large-scale visible light image dataset Glint360K. For the convenience of describing the face recognition method, the network model of this applicant's face recognition selects one of the Resnet50 network models in the Deep Residual Network (Resnet), and it can also be other suitable models, which are not specifically limited herein. For example, it can be a Light-weight Convolutional Neural Network (Light CNN), a Resnet101 network model, a Resnet18 network model, etc. Resnet has a deeper network structure compared to traditional convolutional network models. By introducing residual connections, the problem of gradient disappearance during the training of deep networks is solved, effectively improving the performance of the model. This convolutional neural network Resnet50 is composed of 50 convolutional layers connected in sequence, so it is called Resnet50. The structure of the Resnet50 network is simple, and its algorithm idea is also very concise, which is the calculation of 50 convolutional layers to extract different features of the image, and the image is classified through the last convolutional layer, that is, the fully connected layer. The core of Resnet50 is convolution and residual, and the core of convolution is feature extraction. In the network model, the information flow is passed from the previous layer to the next layer in sequence, and the output of each layer needs to be processed by an activation function. This way of passing information in sequence is prone to the problem of gradient disappearance, especially in deep networks. Resnet50 solves the problem of gradient disappearance by introducing residual connections in the network, allowing information to directly skip and transfer between network layers. Using Resnet50 in the deep learning model as the basic network and pre-training it with a large-scale visible light face image dataset, a better basic feature extractor is obtained, and the need to train a new model from scratch is avoided, improving the training efficiency and accuracy of the model.

[0051] In an embodiment of the present application, fine-tuning the network structure of the pre-trained first face recognition model based on the second face image dataset to obtain a second face recognition model includes: removing the classification layer of the first face recognition model; reconstructing a classification layer based on the number of image categories in the second face image dataset and adding it to the first face recognition model to obtain the second face recognition model. The second face image dataset can be selected from the NIR-VIS heterogeneous face image dataset or other suitable heterogeneous face image datasets, and no specific limitation is made thereto. The second face image dataset includes multiple target objects, each target object includes multiple modalities, and each modality corresponds to multiple face images. For example, the CASIA NIR-VIS 2.0 heterogeneous face recognition dataset includes two modalities of visible light face images and near-infrared face images of the target object.

[0052] In this embodiment, fine-tuning the network structure of the pre-trained first face recognition model based on the second face image dataset means removing the classification layer of the first face recognition model; reconstructing a classification layer based on the number of images in the second face image dataset and adding it to the first face recognition model to obtain the second face recognition model. That is, removing the classification layer in the pre-trained heterogeneous face recognition network and only retaining the feature extraction layer except the classification layer, and at the same time reconstructing the classification layer based on the number of the CASIA NIR-VIS 2.0 heterogeneous face recognition dataset. Since the number of categories in different datasets is different, when training a face recognition model using different datasets, the classification layer in the face recognition model needs to be changed to be applicable to the new dataset. For example, if the first face image dataset for training has 10 people and the output dimension of feature extraction is 128 dimensions, then the weight size of the constructed classification layer is 128 * 10. If the second face image dataset for training has 100 people and the output dimension of feature extraction is 128 dimensions, then the weight size of the constructed classification layer is 128 * 100. Reconstructing the classification layer of the pre-trained heterogeneous face recognition model based on the number of categories in the second face image dataset, that is, the number of categories in the heterogeneous face image dataset, to obtain the second face recognition model.

[0053] In the embodiments of the present application, the second face image dataset includes visible light face images and heterogeneous face images, where some of the heterogeneous face images correspond to the face images with the same identity information as the visible light face images. The second face image dataset is the NIR-VIS heterogeneous face image dataset described above. This heterogeneous face image dataset can be the CASIA NIR-VIS2.0 heterogeneous face recognition dataset or other suitable heterogeneous face image datasets, and no specific limitation is made thereto. The CASIA NIR-VIS2.0 heterogeneous face recognition dataset includes visible light face images and near-infrared face images of the target object, and is a relatively large publicly available face database across the near-infrared and visible light spectra, which is widely used in the performance evaluation of near-infrared-visible light heterogeneous faces. The images in this database come from more than 750 individuals. Each individual has 1-22 visible light images and 5-50 near-infrared images, with approximately more than 17,000 images in total. The images between the two domains do not have a completely one-to-one relationship but are randomly captured. However, there are some heterogeneous face images that correspond to the face images with the same identity information as the visible light face images, that is, the face images of the same person in different modalities. This database also includes variations in illumination, expression, pose, distance, and whether wearing glasses, making it a very challenging database.

[0054] In the embodiments of the present application, the heterogeneous face images include near-infrared face images. When fine-tuning the second face recognition model, the second face image dataset is required. The second face image dataset is the NIR-VIS heterogeneous face image dataset described above. This heterogeneous face image dataset includes visible light face images and heterogeneous face images, where the heterogeneous face images can be near-infrared face images or other suitable heterogeneous face images, and no specific limitation is made thereto. For example, the heterogeneous face images can be sketch face images, comic face images, and near-infrared face images, etc. The second face image dataset can be the CASIA NIR-VIS2.0 heterogeneous face recognition dataset or other suitable heterogeneous face image datasets, and no specific limitation is made thereto. The CASIA NIR-VIS2.0 heterogeneous face recognition dataset includes near-infrared face images (i.e., NIR images) and visible light face images (i.e., VIS images) of multiple individuals. Face feature decoupling training and face recognition tests are performed on the CASIA NIR-VIS2.0 heterogeneous face recognition dataset. This database is an open-source, relatively large and challenging near-infrared and visible light heterogeneous face recognition dataset, and is a very important benchmark dataset in the evaluation of near-infrared and visible light heterogeneous face recognition.

[0055] In an embodiment of the present application, in step S230, the second face recognition model is trained based on the second face image dataset to obtain a trained heterogeneous face recognition model. During the training of the second face recognition model, the network parameters of the backbone network of the second face recognition model are fixed. Training the second face recognition model using the second face image dataset means training the second face recognition model using the CASIA NIR-VIS 2.0 heterogeneous face recognition dataset. During the training process, the network parameters of the backbone network of the second face recognition model are fixed. In deep learning, a model is usually divided into three parts: a backbone network, a neck network, and a head network. The backbone network is the main part of the model and is usually a convolutional neural network (CNN) or a residual neural network (Resnet), etc. The backbone network is responsible for extracting the features of the input image for subsequent processing and analysis. The backbone network usually has many layers and many parameters and can extract high-level feature representations of the image. The neck network is the intermediate layer connecting the backbone network and the head network. The main role of the neck network is to reduce the dimension or adjust the features from the backbone network to better meet the task requirements. The neck network can use convolutional layers, pooling layers, or fully connected layers, etc. The head network is the last layer of the model and is usually a classifier or a regressor. The head network generates the final output by inputting the features processed by the head network. The structure of the head network varies according to different tasks. For example, for image classification tasks, a softmax classifier can be used; for object detection tasks, a bounding box regressor and a classifier, etc., can be used. In the embodiment of the present application, for image classification features, arcface loss classification is used, and other loss classifications, such as softmax loss classification, etc., can also be used, and no specific limitation is made in this regard.

[0056] In this embodiment, the network parameters of the backbone network of the second face recognition model are fixed. The backbone network can extract highly discriminative face features. The first face recognition network is trained using the first face image dataset as described above, that is, the heterogeneous face recognition model is pre-trained based on the large-scale visible light image dataset Glint360K. Due to this pre-training, the backbone network already has highly discriminative face features and strong generalization ability. When the second face recognition model is trained using the second face image dataset, that is, the CASIA NIR-VIS2.0 heterogeneous face recognition dataset is used to train the pre-trained heterogeneous face dataset. Since the number of face image datasets used in the first pre-training and the second training differs greatly, the visible light image dataset used in the first pre-training belongs to a large dataset, and the heterogeneous face image dataset used in the second training belongs to a small dataset. When training for the second time, the model will gradually simulate the distribution of the current small dataset during training, thus losing the generalization ability of the previous large dataset training. Therefore, in order to ensure that overfitting does not occur during training on the small dataset, the network parameters of the backbone network are selected to be fixed.

[0057] In the embodiment of the present application, updating the network parameters of the intermediate network layer of the second face recognition model in step S230 includes: updating the network parameters of the neck network of the second face recognition model. The neck network is the intermediate layer connecting the backbone network and the head network. The main function of the neck network is to reduce the dimension or adjust the features from the backbone network to better meet the task requirements. In heterogeneous face recognition, the visible light image and the near-infrared image of the same target object are extremely similar but there are differences. Therefore, when reducing the dimension of the neck network layer, the respective specific parts of the two can be removed, so as to be mapped to the common latent space and obtain unified features. The common latent space refers to projecting face images taken under different domains, that is, different spectra, into a common representation space, and the expression of the face image identity features is not affected by the spectrum. By fixing the network parameters of the backbone network and updating the network parameters of the neck network, it is ensured that the highly discriminative features obtained by the model during the first pre-training remain unchanged, the generalization ability of the model is improved, and at the same time, the phenomenon of overfitting can be prevented.

[0058] In an embodiment of the present application, during the training of the second face recognition model, the network parameters of the batch sample normalization layer of the second face recognition model are fixed. The batch sample normalization layer first calculates the pixel values of all points in each channel, obtains the mean and variance, then divides the pixel value of each point minus the mean by the variance on each channel to obtain the pixel value of that point, and finally connects it to the activation function, which can accelerate the convergence of the network. To a certain extent, the parameters of the batch sample normalization layer reflect the distribution of the dataset. Similarly, to ensure that the face recognition model is not affected by the distribution of the small dataset as much as possible, once the distributions of the test data and the training data are different, the performance of the network will be greatly reduced. Therefore, the network parameters of the batch sample normalization layer are set to be fixed.

[0059] In an embodiment of the present application, during the training of the second face recognition model, the learning rate of the second face recognition model is set to be at least two orders of magnitude smaller than the learning rate of the first face recognition model. The learning rate controls the step size of the gradient update after each backpropagation. During the first training, that is, the pre-training, the model network has already obtained relatively high performance. If a relatively large learning rate is still used during the second model training, it is very easy for the network to jump out of the local optimal point. Therefore, the learning rate during the second model training is at least two orders of magnitude smaller than the learning rate of the first face recognition model. For example, if the learning rate for the pre-training of the first face recognition model is set to 0.01, then the learning rate for the training of the second face recognition model can be set to 0.0001.

[0060] In an embodiment of the present application, the learning rate of the second face recognition model is not greater than 0.001. The learning rate controls the step size of the gradient update after each backpropagation. During the first model pre-training, relatively high performance has been obtained. If a relatively large learning rate is set, it will cause the model network to jump out of the local optimal point. Therefore, setting the learning rate not to be greater than 0.001 is to ensure the performance and generalization ability of the model. As described above, the network parameters of the backbone network have been fixed, and only the network parameters of the neck network are updated. The neck network is in a deeper part of the model network and has fewer parameters, and does not require too large a learning rate for update either, otherwise it will also jump out of the local optimum. In an embodiment of the present application, the initial value of the learning rate of the second face recognition model can be set to 0.00005 to 0.0002, or other suitable values, and no specific limitation is made thereto.

[0061] Therefore, the training method of the heterogeneous face recognition model according to the embodiments of the present application can fine-tune some network parameters of the face recognition model using a heterogeneous face dataset without changing the network structure of the face recognition model, thereby improving the accuracy of heterogeneous face recognition.

[0062] The following is combined with Figure 3Describe the face recognition method provided according to another aspect of the present application. Figure 3 FIG. 300 shows a schematic flowchart of a face recognition method according to an embodiment of the present application. The face recognition method 300 according to an embodiment of the present application may include the following steps:

[0063] In step S310, obtain a face image to be recognized;

[0064] In step S320, recognize the face image to be recognized based on the trained heterogeneous face recognition model to obtain a face recognition result; wherein, the heterogeneous face recognition model is obtained based on the training method of the heterogeneous face recognition model described above.

[0065] Therefore, the face recognition method according to an embodiment of the present application can fine-tune some network parameters of the face recognition model using a heterogeneous face dataset without changing the network structure of the face recognition model, improving the accuracy of heterogeneous face recognition.

[0066] In an embodiment of the present application, the face image to be recognized includes a heterogeneous face image different from a visible light face image. Based on the previously trained heterogeneous face recognition model, the heterogeneous face image is input into the trained heterogeneous face recognition model to obtain a recognition result of the face image. The heterogeneous face image and the visible light face image belong to different modalities. For example, the visible light image is one modality, and the near-infrared light image is one modality. The heterogeneous face image may be a face image of another modality different from the visible light face image, and no specific limitation is made thereto.

[0067] In an embodiment of the present application, the heterogeneous face image includes a near-infrared face image. The heterogeneous face image may also be other suitable heterogeneous face images, and no specific limitation is made thereto. For example, the heterogeneous face image may be a sketch face image, a cartoon face image, a near-infrared face image, etc.

[0068] Next, in conjunction with Figure 4 Describe a training device for a heterogeneous face recognition model provided according to another aspect of the present application. Figure 4 FIG. 400 shows a schematic structural block diagram of a training device 400 for a heterogeneous face recognition model according to an embodiment of the present application. As Figure 4As shown in the figure, the training device 400 for heterogeneous face recognition models includes a memory 410 and a processor 420, where: a computer-executable program run by the processor 420 is stored on the memory 410. When the computer-executable program is run by the processor 420, the processor 420 is caused to execute the foregoing training method 200 for heterogeneous face recognition models. Those skilled in the art can understand the structures and specific operations of the various modules in the training device 400 for heterogeneous face recognition models according to the embodiments of the present application in combination with the foregoing content. For the sake of brevity, it will not be elaborated here. Therefore, the training device for heterogeneous face recognition models according to the embodiments of the present application can fine-tune some network parameters of the face recognition model using a heterogeneous face dataset without changing the network structure of the face recognition model, improving the accuracy of heterogeneous face recognition.

[0069] Next, in combination with Figure 5 a description is given of a face recognition device provided according to another aspect of the present application. Figure 5 FIG. shows a schematic structural block diagram of a face recognition device 500 according to an embodiment of the present application. As Figure 5 shown, the face recognition device 500 includes a memory 510 and a processor 520, where: a computer-executable program run by the processor 520 is stored on the memory 510. When the computer-executable program is run by the processor 520, the processor 520 is caused to execute the foregoing face recognition method 300. Those skilled in the art can understand the structures and specific operations of the various modules in the face recognition device 500 according to the embodiments of the present application in combination with the foregoing content. For the sake of brevity, it will not be elaborated here. Therefore, the face recognition device according to the embodiments of the present application can fine-tune some network parameters of the face recognition model using a heterogeneous face dataset without changing the network structure of the face recognition model, improving the accuracy of heterogeneous face recognition.

[0070] Next, in combination with Figure 6 a description is given of a storage medium provided according to another aspect of the present application. Figure 6 FIG. shows a schematic structural block diagram of a storage medium 600 according to an embodiment of the present application. As Figure 6 shown, for the storage medium 600, a computer program is stored thereon, and when the computer program is run by a processor, the processor is caused to execute the foregoing training method for heterogeneous face recognition models according to the embodiments of the present application. The storage medium may include, for example, a memory card of a smart phone, a storage component of a tablet computer, a hard disk of a personal computer, a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a portable compact disc read-only memory (CD-ROM), a USB memory, or any combination of the above storage media. The computer-readable storage medium may be any combination of one or more computer-readable storage media.

[0071] Based on the above description, the training method, face recognition method and device of the heterogeneous face recognition model according to the embodiments of the present application can fine-tune some network parameters of the face recognition model using a heterogeneous face dataset without changing the network structure of the face recognition model, thereby improving the accuracy of heterogeneous face recognition.

[0072] Although example embodiments have been described herein with reference to the accompanying drawings, it should be understood that the above example embodiments are merely exemplary and are not intended to limit the scope of the present application. Those of ordinary skill in the art can make various changes and modifications therein without departing from the scope and spirit of the present application. All such changes and modifications are intended to be included within the scope of the present application as claimed in the appended claims.

[0073] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in connection with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0074] In several embodiments provided by the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed.

[0075] In the specification provided herein, a large number of specific details are set forth. However, it can be understood that the embodiments of the present application can be practiced without these specific details. In some instances, well-known methods, structures, and technologies have not been shown in detail so as not to obscure the understanding of this specification.

[0076] Similarly, it should be understood that, in order to streamline this application and assist in understanding one or more of the various inventive aspects, in the description of the exemplary embodiments of this application, the various features of this application are sometimes grouped together into a single embodiment, figure, or description thereof. However, the methods of this application should not be construed as reflecting an intention that the claimed application requires more features than are expressly recited in each claim. Rather, as reflected in the corresponding claims, the inventive point lies in that the corresponding technical problems can be solved with features less than all the features of a single disclosed embodiment. Therefore, the claims following the detailed description are hereby expressly incorporated into the detailed description, where each claim itself serves as a separate embodiment of this application.

[0077] Those skilled in the art will understand that, except for features that are mutually exclusive, any combination can be used to combine all the features disclosed in this specification (including the accompanying claims, abstract, and drawings) and all the processes or units of any method or device so disclosed. Unless otherwise expressly stated, each feature disclosed in this specification (including the accompanying claims, abstract, and drawings) can be replaced by an alternative feature that provides the same, equivalent, or similar purpose.

[0078] In addition, those skilled in the art can understand that, although some of the embodiments described herein include certain features included in other embodiments rather than other features, the combination of features of different embodiments means that it is within the scope of this application and forms different embodiments. For example, in the claims, any one of the claimed embodiments can be used in any combination.

[0079] The various component embodiments of this application can be implemented in hardware, or in software modules running on one or more processors, or in a combination thereof. Those skilled in the art should understand that a microprocessor or a digital signal processor (DSP) can be used in practice to implement some or all of the functions of some of the modules according to the embodiments of this application. This application can also be implemented as a program (for example, a computer program and a computer program product) for executing part or all of the methods described herein. Such a program implementing this application can be stored on a computer-readable medium, or can be in the form of one or more signals. Such signals can be downloaded from an Internet website, or provided on a carrier signal, or provided in any other form.

[0080] It should be noted that the above embodiments are illustrative of the present application rather than restrictive of the present application, and those skilled in the art can design alternative embodiments without departing from the scope of the appended claims. In the claims, any reference signs placed between parentheses shall not be construed as limiting the claim. The word "comprising" does not exclude the presence of elements or steps not listed in the claim. The word "a" or "an" preceding an element does not exclude the presence of a plurality of such elements. The present application can be implemented by means of hardware including several different elements and by means of a suitably programmed computer. In the unit claims listing several abnormal detection devices of a train traction system, several of these abnormal detection devices of the train traction system can be embodied by the same item of hardware. The use of the words first, second, and third, etc. does not denote any order. These words can be interpreted as names.

[0081] As described above, the specific embodiments of the present application or the description of the specific embodiments are only for illustration, and the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered within the protection scope of the present application. The protection scope of the present application shall be subject to the protection scope of the claims.

Claims

1. A training method for a heterogeneous face recognition model, characterized in that, The method includes: Training a first face recognition network based on a first face image dataset to obtain a pre-trained first face recognition model, where the first face image dataset is a visible light face image dataset; Fine-tuning the network structure of the pre-trained first face recognition model based on a second face image dataset to obtain a second face recognition model, where the number of images in the second face image dataset is at least two orders of magnitude smaller than the number of images in the first face image dataset, and the second face image dataset includes at least heterogeneous face images different from visible light face images; Training the second face recognition model based on the second face image dataset to obtain a trained heterogeneous face recognition model, where during the training of the second face recognition model, the network parameters of the backbone network of the second face recognition model are fixed, and the network parameters of the intermediate network layer of the second face recognition model are updated.

2. The method according to claim 1, characterized in that, The updating of the network parameters of the intermediate network layer of the second face recognition model includes: Updating the network parameters of the neck network of the second face recognition model.

3. The method according to claim 1, characterized in that, The method further includes: during the training of the second face recognition model, fixing the network parameters of the batch sample normalization layer of the second face recognition model.

4. The method according to any one of claims 1 to 3, characterized in that, The method further includes: during the training of the second face recognition model, setting the learning rate of the second face recognition model to be at least two orders of magnitude smaller than the learning rate of the first face recognition model.

5. The method according to claim 4, wherein The learning rate of the second face recognition model is not greater than 0.

001.

6. The method according to claim 5, wherein The initial learning rate of the second face recognition model is from 0.00005 to 0.0002.

7. The method according to any one of claims 1 to 3, characterized in that, The fine-tuning of the network structure of the pre-trained first face recognition model based on the second face image dataset to obtain a second face recognition model includes: Removing the classification layer of the first face recognition model; Reconstructing a classification layer based on the number of images in the second face image dataset and adding it to the first face recognition model to obtain the second face recognition model.

8. The method according to any one of claims 1 to 3, characterized in that, The second face image dataset includes the visible light face images and the heterogeneous face images, where there are some images in the heterogeneous face images that correspond to the face images with the same identity information as the visible light face images.

9. The method according to any one of claims 1 to 3, characterized in that The heterogeneous face images include near-infrared face images.

10. A face recognition method, characterized in that, The method includes: Obtaining a face image to be recognized; Recognizing the face image to be recognized based on the trained heterogeneous face recognition model to obtain a face recognition result; Wherein, the heterogeneous face recognition model is obtained based on the training method of the heterogeneous face recognition model according to any one of claims 1-9.

11. The method according to claim 10, wherein The face image to be recognized includes heterogeneous face images different from visible light face images.

12. The method according to claim 11, wherein The heterogeneous face images include near-infrared face images.

13. A training device for a heterogeneous face recognition model, characterized in that, The device includes a processor and a memory, and a computer program run by the processor is stored on the memory. When the computer program is run by the processor, the processor is caused to execute the method for training a heterogeneous face recognition model according to any one of claims 1-9.

14. A face recognition device, characterized in that, The device includes a processor and a memory, and a computer program run by the processor is stored on the memory. When the computer program is run by the processor, the processor is caused to execute the face recognition method according to any one of claims 10-12.

15. A storage medium, characterized in that, A computer program run by a processor is stored on the storage medium. When the computer program is run by the processor, the processor is caused to execute the method for training a heterogeneous face recognition model according to any one of claims 1-9 or execute the face recognition method according to any one of claims 10-12.