A method, device and electronic device for constructing a cross-modal face recognition model
By acquiring and processing face image data in visible light and near-infrared modes, using key point detection and neural network training, the problem of complex and low accuracy of cross-modal face recognition model training is solved, and efficient cross-modal face recognition is achieved.
Patent Information
- Application Number
- CN202111418637.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-25
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2041-11-25
AI Technical Summary
The existing cross-modal face recognition model is complex in training and has low recognition accuracy when switching lighting environments, making it difficult to effectively recognize faces in visible light and near-infrared modes.
By obtaining the face image data set containing visible light and near-infrared modes, the face key point detection algorithm is used to extract key points, and different combinations are input into the neural network model. The neural network model is trained using the correlation degree until a model is obtained for cross-modal face recognition.
The training process is simplified, the accuracy of cross-modal face recognition is improved, and the face recognition can be directly performed in visible light and near-infrared modes without mode conversion.
Smart Images

Figure CN114038045B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of face recognition, and in particular, to a method, device and electronic device for constructing a cross-modal face recognition model. Background Art
[0002] When using a visible light - near infrared adaptive camera for face recognition, the different working modes of the camera can be switched according to the lighting environment. When the light intensity meets the threshold set by the ISP module, it can be switched to the visible light mode, and conversely, to the near infrared mode. Since the distribution range of near infrared cameras is not as extensive as that of visible light, and there are privacy issues for users, it is difficult to collect a large amount of near infrared image data. To solve the modality difference, the current cross-modal face recognition methods mainly include: (1) Based on visible light images, design and extract modality-invariant features, and through transfer learning, transform face images from one modality to another. This method has the problem of domain inadaptability, poor model generalization ability, and low recognition accuracy; (2) Extract the image features of the two modalities, combine them into new feature information, and project them onto a common subspace. This method is limited by the number of IDs in the current face recognition dataset and it is difficult to train a model with high accuracy; (3) Use a generative adversarial network to generate an approximate near infrared image dataset to approximate the face feature distribution trend of near infrared images. This method extracts features based on visible light images and generates approximate near infrared images to train the face recognition model in the near infrared mode, making the training process of the face recognition model in the near infrared mode complex and the recognition accuracy low. Summary of the Invention
[0003] In view of this, embodiments of the present invention provide a method, device and electronic device for constructing a cross-modal face recognition model to solve the technical problems of complex training of the face recognition model in a cross-modal environment and low face recognition accuracy in the prior art.
[0004] The technical solutions proposed by the present invention are as follows:
[0005] In the first aspect of the embodiment of the present invention, a method for constructing a cross-modal face recognition model is provided. The method for constructing a cross-modal face recognition model includes: obtaining a face image dataset, where the face image dataset includes face images obtained in the visible light mode and face images obtained in the near-infrared mode; using a face key point detection algorithm to perform key point detection on the face images in the face image dataset; extracting multiple key points included in each face image from the face image dataset, and selecting a sub-image where a target key point is located from the multiple key points; inputting different combinations formed by the face image dataset and the sub-images where the multiple target key points corresponding to each face image in the face image dataset are located into a target neural network model; using a classification module in the target neural network model to determine the association degree between each face image and each corresponding combination; and training the target neural network model according to the association degree between each face image and each corresponding combination until a face recognition model for performing face recognition in the visible light mode and the near-infrared mode is obtained.
[0006] Optionally, after using the face key point detection algorithm to perform key point recognition on the face images in the face image dataset and before extracting multiple key points included in each face image from the face image dataset and selecting a sub-image where a target key point is located from the multiple key points, the method includes: correcting the key points in the recognized face images according to a preset face template to complete the face alignment operation.
[0007] Optionally, before inputting different combinations formed by the face image dataset and the sub-images where the multiple target key points corresponding to each face image in the face image dataset are located into the target neural network model, the method further includes: obtaining a pre-trained face image dataset; using the pre-trained face image dataset to determine a pre-trained neural network model; and using the obtained face image dataset to train the pre-trained neural network model to obtain the target neural network model.
[0008] Optionally, training the target neural network model includes: determining the loss value of the loss function corresponding to the target neural network model during the training process until the loss value of the loss function meets a preset condition to obtain a face recognition model for performing face recognition in the visible light mode and the near-infrared mode.
[0009] In a second aspect of the embodiments of the present invention, a cross-modal face recognition model construction device is provided. The cross-modal face recognition model construction device includes: an acquisition module, configured to acquire a face image data set, where the face image data set includes face images acquired in a visible light mode and face images acquired in a near-infrared mode; a detection module, configured to perform key point detection on the face images in the face image data set by using a face key point detection algorithm; a selection module, configured to extract multiple key points included in each face image from the face image data set, and select a sub-image where a target key point is located from the multiple key points; an input module, configured to input different combinations formed by the face image data set and the sub-images where the multiple target key points corresponding to each face image in the face image data set are located into a target neural network model; a determination module, configured to use a classification module in the target neural network model to determine the association degree between each face image and each corresponding combination; a training module, configured to train the target neural network model according to the association degree between each face image and each corresponding combination until a face recognition model for performing face recognition in the visible light mode and the near-infrared mode is obtained.
[0010] Optionally, the device further includes: an alignment module, configured to correct the key points in the recognized face image according to a preset face template to complete the face alignment operation.
[0011] Optionally, the device further includes: a first acquisition module, configured to acquire a pre-trained face image data set; a first determination module, configured to determine a pre-trained neural network model by using the pre-trained face image data set; a first training module, configured to train the pre-trained neural network model by using the acquired face image data set to obtain the target neural network model.
[0012] Optionally, the device further includes: a second determination module, configured to determine the loss value of a loss function corresponding to the target neural network model during the training process until the loss value of the loss function meets a preset condition to obtain a face recognition model for performing face recognition in the visible light mode and the near-infrared mode.
[0013] In a third aspect of the embodiments of the present invention, a computer-readable storage medium is provided. The computer-readable storage medium stores computer instructions, and the computer instructions are used to cause the computer to execute the cross-modal face recognition model construction method as described in the first aspect and any one of the first aspect of the embodiments of the present invention.
[0014] A fourth aspect of an embodiment of the present invention provides an electronic device, including: a memory and a processor, which are communicatively connected to each other. The memory stores computer instructions, and the processor executes the computer instructions to execute the cross-modal face recognition model construction method as described in the first aspect and any one of the first aspects of the embodiments of the present invention.
[0015] The technical solution provided by the present invention has the following effects:
[0016] The cross-modal face recognition model construction method provided by the embodiment of the present invention obtains a face image data set, where the face image data set includes face images obtained in the visible light mode and face images obtained in the near-infrared mode; uses a face key point detection algorithm to perform key point detection on the face images in the face image data set; extracts multiple key points included in each face image from the face image data set, and selects a sub-image where the target key point is located from the multiple key points; inputs different combinations formed by the face image data set and the sub-images where the multiple target key points corresponding to each face image in the face image data set are located into a target neural network model; uses a classification module in the target neural network model to determine the correlation degree between each face image and each corresponding combination; and trains the target neural network model according to the correlation degree between each face image and each corresponding combination until a face recognition model for performing face recognition in the visible light mode and the near-infrared mode is obtained. This method uses the correlation degree between each face image and each corresponding combination to train the neural network model, and the training process is simple; moreover, the face recognition in two modes can be directly completed simultaneously using this correlation degree without conversion, improving the accuracy of face recognition in the cross-modal mode. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for use in the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0018] Figure 1 is a flowchart of a cross-modal face recognition model construction method according to an embodiment of the present invention;
[0019] Figure 2 is a structural block diagram of a cross-modal face recognition model construction device according to an embodiment of the present invention;
[0020] Figure 3Schematic diagram of the structure of a computer-readable storage medium provided according to an embodiment of the present invention;
[0021] Figure 4 Schematic diagram of the structure of an electronic device provided according to an embodiment of the present invention. Detailed implementation manners
[0022] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0023] An embodiment of the present invention provides a method for constructing a cross-modal face recognition model, as Figure 1 shown, the method includes the following steps:
[0024] Step S101: Obtain a face image dataset, where the face image dataset includes face images obtained in the visible light mode and face images obtained in the near-infrared mode. Specifically, the images in the face image dataset are face images obtained in the visible light mode and face images obtained in the near-infrared mode for the same face ID. Before obtaining the face image dataset in the multi-modal environment, a data acquisition platform is built using a near-infrared and a visible light camera, and then the local dataset is collected using this platform. During the collection process, it is necessary to ensure that each face ID image set includes near-infrared images and visible light images. In one embodiment, the number of collected face IDs is ten thousand.
[0025] Step S102: Use a face key point detection algorithm to perform key point detection on the face images in the face image dataset. Specifically, before performing face detection on the collected image set, obtain the face detection frame of each image, and then use a face detection algorithm to detect the target face in the dataset, and use a face key point detection algorithm to perform key point detection on the near-infrared images and visible light images in the face image dataset respectively, to obtain the face key points of the near-infrared images and the face key points of the visible light images. Among them, the face detection algorithm is to locate the face position in the image set and provide an initial input frame for the subsequent face key point detection algorithm. The face key point detection algorithm is to locate the key point positions of a given face image, including 68 key point positions such as eyebrows, eyes, nose, mouth, and face contour. The embodiments of the present application do not limit this face detection algorithm and face key point detection algorithm, as long as they can meet the requirements of face detection and key point recognition.
[0026] Step S103: Extract multiple key points included in each face image from the face image dataset, and select the sub-image where the target key point is located from the multiple key points. Specifically, after obtaining the face image dataset, extract 68 key points included in each face image, and select 5 target key points from these 68 key points, including the center points of the left and right eyes, the tip of the nose, and the two corners of the mouth. Crop out four sub-images of the left eye, right eye, nose, and mouth according to the positions of these 5 key points. Among them, the size of each regional sub-image is adjusted to a unified size of 256*256.
[0027] Step S104: Input different combinations formed by the face image dataset and the sub-images where multiple target key points corresponding to each face image in the face image dataset are located into the target neural network model. Specifically, after obtaining the face image and the sub-images where multiple target key points are located, input them into the feature extraction network module in the target neural network model to obtain the features of the face image and the sub-images where multiple target key points are located. Combine the features of the face image and the sub-images where multiple target key points are located in a certain order, and then input the different combinations formed into the feature fusion module of the target neural network model.
[0028] In one embodiment, input the four cropped sub-images of the left eye, right eye, nose, and mouth and the corresponding entire face image into the backbone network NetA, that is, the feature extraction module of the target neural network model. Among them, the backbone network refers to a lightweight deep residual learning framework. Then, input the combined features into the feature fusion module of the target neural network model, where this module refers to a network module composed of fully connected layers.
[0029] Step S105: Determine the correlation degree between each face image and each corresponding combination by using the classification module in the target neural network model. Specifically, after combining the features of the entire face image and the sub-images where multiple target key points are located, use the classification module in the target neural network model to classify and identify the combined features, and learn the correlation degree between each face image and each corresponding combination. For example, after performing different forms of combined splicing on the sub-images where the left eye, right eye, nose, and mouth are located, only one combination of splicing results can preferably represent the fine-grained characterization information of the face individual image, which is not affected by the modality environment, denoted as combination M, and other combination forms are sub-optimally representative of the fine-grained characterization information of the face individual image and are susceptible to the modality environment, denoted as combination N. Assign learnable correlation coefficients to each combination for feature concatenation, input them into the classification model, and use circle loss for supervised learning. Circle loss includes the triplet loss function tripletloss, which is suitable for supervised learning of paired different combined features. Set the range of this correlation coefficient to (0, 1), and perform a random initialization assignment operation on this correlation coefficient to be 0.4 - 0.5. Through supervised learning, this correlation coefficient gradually converges and shows the following distribution: the correlation coefficient corresponding to combination M gradually tends to 1, and the correlation coefficient corresponding to combination N gradually tends to 0.
[0030] Step S106: Train the target neural network model according to the correlation degree between each face image and each corresponding combination until a model for face recognition in visible light mode and near-infrared mode is obtained. Specifically, during the process of adaptively learning the correlation degree between each face image and each corresponding combination, this correlation degree can assist in supervising the training of the target neural network model, enabling the network to learn the main features of the face in a multi-modal environment, which are not affected by cross-modal, until the training converges, and finally obtaining a model for face recognition in visible light mode and near-infrared mode. By learning the correlation degree of facial region features, the training of the target neural network is strengthened, so that in cross-modal face recognition, the key region feature combination with the largest correlation degree is used as the final main face feature, improving the accuracy of cross-modal face recognition.
[0031] The method for constructing a cross-modal face recognition model provided by the embodiment of the present invention includes obtaining a face image dataset, where the face image dataset includes face images obtained in the visible light mode and face images obtained in the near-infrared mode; using a face key point detection algorithm to perform key point detection on the face images in the face image dataset; extracting multiple key points included in each face image from the face image dataset, and selecting the sub-image where the target key point is located from the multiple key points; inputting different combinations formed by the face image dataset and the sub-images where the multiple target key points corresponding to each face image in the face image dataset are located into the target neural network model; using the classification module in the target neural network model to determine the correlation degree between each face image and each corresponding combination; and training the target neural network model according to the correlation degree between each face image and each corresponding combination until a face recognition model for performing face recognition in the visible light mode and the near-infrared mode is obtained. This method uses the correlation degree between each face image and each corresponding combination to train the neural network model, and the training process is simple; moreover, using this correlation degree can directly complete face recognition in both modes at the same time without conversion, improving the accuracy of face recognition in the cross-modal mode.
[0032] As an optional implementation manner of the embodiment of the present invention, after step 102 and before step 103, the method further includes: correcting the key points in the recognized face image according to a preset face template to complete the face alignment operation.
[0033] Specifically, after obtaining the face key points, first set a face standard template, and then perform an alignment operation on the face image according to this template, that is, correct the key points in the detected face image. Among them, the face standard template means that five points, such as the center points of two eyes, the tip of the nose, and the corner points of two mouths, have fixed values and do not change with different faces. Among them, on the coordinate axis with the point in the upper left corner of the template image as the coordinate origin, the position information coordinates (x, y) of each key point of the entire face image in the corresponding entire face image can be, for example: the coordinate of the position where the left eye is located is (38.29, 51.69), the coordinate of the position where the right eye is located is (73.53, 51.50), the coordinate of the position where the nose is located is (56.05, 71.73), and the coordinates of the positions corresponding to both sides of the mouth are (41.54, 92.36) and (70.72, 92.20).
[0034] Face alignment can calculate the affine transformation matrix according to the coordinates of five key points at the corresponding positions of the existing face and the coordinates of five points of the standard template, and then multiply it with the existing face image to complete face alignment.
[0035] In one embodiment, a face image A is matched with a standard template B as a face template. Among them, the face image A includes five key points A1 to A5, and the standard template B includes five key points B1 to B5. Through the point-to-point correspondence, an affine transformation is used to solve the transformation matrix M. Multiplying A1 to A5 by M can complete face alignment.
[0036] As an optional implementation manner of an embodiment of the present invention, before inputting different combinations formed by a face image data set and sub-images where multiple target key points corresponding to each face image in the face image data set are located into a target neural network model, the method further includes: obtaining a pre-trained face image data set; determining a pre-trained neural network model by using the pre-trained face image data set; training the pre-trained neural network model by using the obtained face image data set to obtain a target neural network model. Among them, the obtained face image data set refers to a standard face image data set obtained after correcting the key points in the recognized face image.
[0037] Specifically, before obtaining the pre-trained face image data set, the face image is first processed.
[0038] In one embodiment, the face image is scaled to 256*256 and grayscaled to obtain the gray value of each pixel of the image, and the values of the three RGB channels of each pixel are set to this gray value.
[0039] Then, the publicly available face recognition data set glint360k is obtained, and a backbone network NetA pre-training model adapted to the application platform, that is, a pre-trained neural network model, is obtained by modifying based on mobilefacenet. Among them, the face recognition data set glint360k has 360,000 category numbers and 17 million photos; mobilefacenet is a lightweight network running on mobile devices; the accuracy of the backbone network NetA pre-training model on the publicly available benchmark test set is lfw-99.66, cfp-96.04, agedb-95.8.
[0040] After determining the pre-trained neural network model, the face image data set obtained above, which includes face images obtained in the visible light mode and face images obtained in the near-infrared mode, is used, and hyperparameters such as the learning rate and batch_size are set to fine-tune the pre-trained neural network model to obtain a target neural network model. Specifically, network fine-tuning through transfer learning can accelerate convergence and improve the generalization ability of the model. In one embodiment, the learning rate is set to 0.0001, the batch size is set to 128, and other corresponding hyperparameters remain unchanged.
[0041] As an alternative implementation of the embodiment of the present invention, after determining the target neural network model, the model is trained until a model for face recognition in the visible light mode and the near-infrared mode is obtained. Specifically, first, feature vectors are extracted from each key sub-image and the entire face image input into the target neural network model, the feature vectors are combined according to certain rules, and then input into a feature fusion module composed of fully connected layers. Then, the final feature vector is determined based on the learnable association degree between each face image and each combination thereof and the feature vector output by the feature fusion module. Finally, the loss value of the circle loss of the loss function corresponding to the target neural network model during the training process is determined until the model converges, and a face recognition model for face recognition in the visible light mode and the near-infrared mode is obtained. Specifically, when the loss value reaches a certain value and remains stable, the model converges and the training ends, and a model for face recognition in the visible light mode and the near-infrared mode is obtained.
[0042] In one embodiment, 512-dimensional feature vectors are extracted from the four regional sub-images of the left eye, right eye, nose, and mouth and the entire face image cropped out in the pre-trained model of the backbone network NetA, denoted as f1, f2, f3, f4, f5, which represent the feature vectors of the left eye, right eye, nose, mouth, and the entire face image in sequence. Among them, the dimension of the extracted vector does not have a specific value. Here, 512-dimensional feature vectors are selected for extraction, which is only set considering the running speed and model accuracy of the currently deployed model. In this solution, the dimension of the extracted vector is not specifically limited.
[0043] The five extracted feature vectors are combined pairwise in sequence to obtain ten 1024-dimensional feature vectors, and the ten feature vectors are input into a fully connected layer Fc to obtain the feature vector fa, and the association degree between the feature sub-images of the face key points is learned.
[0044] The feature vector fa is input into ten fully connected layers, and ten 1024-dimensional feature vectors are obtained. Then, ten different learnable adaptive weights r are set to characterize the association degree strength between the feature sub-images of the face key points, and the final 512-dimensional feature vector fb is obtained by weighted calculation based on the weight r and the ten 1024-dimensional feature vectors obtained by inputting into the ten fully connected layers. Finally, the circle loss is used as the loss function, and the target neural network model is trained until a model for face recognition in the visible light mode and the near-infrared mode is obtained.
[0045] In one embodiment, n near-infrared images and n visible light images are input into the backbone network NetA, and the corresponding features A and B are obtained. Among them, their respective features are randomly combined by facial features such as the nose and eyes. Then, the feature is input into a fully connected layer, and supervised learning is performed through a loss function to achieve registration.
[0046] An embodiment of the present invention further provides a cross-modal face recognition model construction device, as Figure 2 shown. The device includes:
[0047] An acquisition module 401, configured to acquire a face image data set, where the face image data set includes face images acquired in the visible light mode and face images acquired in the near-infrared mode; for detailed content, refer to the relevant description of step S101 in the above method embodiment.
[0048] A detection module 402, configured to perform key point detection on the face images in the face image data set by using a face key point detection algorithm; for detailed content, refer to the relevant description of step S102 in the above method embodiment.
[0049] A selection module 403, configured to extract multiple key points included in each face image from the face image data set, and select a sub-image where a target key point is located from the multiple key points; for detailed content, refer to the relevant description of step S103 in the above method embodiment.
[0050] An input module 404, configured to input different combinations formed by the face image data set and the sub-images where multiple target key points corresponding to each face image in the face image data set are located into a target neural network model; for detailed content, refer to the relevant description of step S104 in the above method embodiment.
[0051] A determination module 405, configured to use a classification module in the target neural network model to determine the association degree between each face image and each corresponding combination; for detailed content, refer to the relevant description of step S105 in the above method embodiment.
[0052] A training module 406, configured to train the target neural network model according to the association degree between each face image and each corresponding combination until a face recognition model for performing face recognition in the visible light mode and the near-infrared mode is obtained; for detailed content, refer to the relevant description of step S106 in the above method embodiment.
[0053] The cross-modal face recognition model construction device provided by the embodiment of the present invention acquires a face image data set, where the face image data set includes face images acquired in the visible light mode and face images acquired in the near-infrared mode; uses a face key point detection algorithm to perform key point detection on the face images in the face image data set; extracts multiple key points included in each face image from the face image data set, and selects the sub-image where the target key point is located from the multiple key points; inputs different combinations formed by the face image data set and the sub-images where the multiple target key points corresponding to each face image in the face image data set are located into the target neural network model; uses the classification module in the target neural network model to determine the association degree between each face image and each corresponding combination; trains the target neural network model according to the association degree between each face image and each corresponding combination until a face recognition model for performing face recognition in the visible light mode and the near-infrared mode is obtained. This method uses the association degree between each face image and each corresponding combination to train the neural network model, and the training process is simple; moreover, the face recognition in two modes can be directly completed simultaneously using this association degree without conversion, improving the accuracy of face recognition in the cross-modal mode.
[0054] As an optional implementation manner of the embodiment of the present invention, the device further includes: an alignment module, configured to correct the key points in the recognized face image according to a preset face template to complete the face alignment operation.
[0055] As an optional implementation manner of the embodiment of the present invention, the device further includes: a first acquisition module, configured to acquire a pre-trained face image data set; a first determination module, configured to determine a pre-trained neural network model by using the pre-trained face image data set; a first training module, configured to train the pre-trained neural network model by using the acquired face image data set to obtain a target neural network model.
[0056] As an optional implementation manner of the embodiment of the present invention, the device further includes: a second determination module, configured to determine the loss value of the loss function corresponding to the target neural network model during the training process until the loss value of the loss function meets a preset condition to obtain a face recognition model for performing face recognition in the visible light mode and the near-infrared mode.
[0057] For the detailed function description of the cross-modal face recognition model construction device provided by the embodiment of the present invention, please refer to the description of the cross-modal face recognition model construction method in the above embodiment.
[0058] The embodiment of the present invention further provides a storage medium, such as Figure 3As shown, a computer program 601 is stored thereon. When the instructions are executed by a processor, the steps of the cross-modal face recognition model construction method in the above embodiments are implemented. The storage medium also stores audio-visual stream data, feature frame data, interaction request signaling, encrypted data, and a preset data size, etc. Among them, the storage medium can be a magnetic disk, an optical disc, a read-only memory (ROM), a random access memory (RAM), a flash memory, a hard disk drive (HDD), or a solid-state drive (SSD), etc.; the storage medium can also include a combination of the above types of memories.
[0059] Those skilled in the art can understand that to implement all or part of the processes in the above method embodiments, it can be completed by instructing relevant hardware through a computer program. The program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the above method embodiments. Among them, the storage medium can be a magnetic disk, an optical disc, a read-only memory (ROM), a random access memory (RAM), a flash memory, a hard disk drive (HDD), or a solid-state drive (SSD), etc.; the storage medium can also include a combination of the above types of memories.
[0060] The embodiment of the present invention also provides an electronic device, such as Figure 4 As shown, the electronic device may include a processor 51 and a memory 52. The processor 51 and the memory 52 can be connected through a bus or other means. Figure 4 Taking the connection through the bus as an example.
[0061] The processor 51 can be a central processing unit (CPU). The processor 51 can also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. chips, or a combination of the above types of chips.
[0062] The memory 52, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs, non-transitory computer-executable programs, and modules, such as the corresponding program instructions / modules in the embodiments of the present invention. The processor 51 executes various functional applications and data processing of the processor by running the non-transitory software programs, instructions, and modules stored in the memory 52, that is, implements the cross-modal face recognition model construction method in the above method embodiments.
[0063] The memory 52 may include a program storage area and a data storage area. Among them, the program storage area can store an operating device and application programs required for at least one function; the data storage area can store data created by the processor 51 and the like. In addition, the memory 52 may include a high-speed random access memory, and may also include a non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory 52 may optionally include a memory remotely set relative to the processor 51, and these remote memories can be connected to the processor 51 through a network. Examples of the above networks include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0064] The one or more modules are stored in the memory 52 and, when executed by the processor 51, execute the cross-modal face recognition model construction method in the Figure 1 embodiments shown.
[0065] Specific details of the above electronic device can be understood by referring to the corresponding relevant descriptions and effects in the Figure 1 embodiments shown, and will not be elaborated here.
[0066] Although the embodiments of the present invention have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the present invention, and such modifications and variations all fall within the scope defined by the appended claims.
Claims
1. A method for constructing a cross-modal face recognition model, characterized in that, It includes the following steps: Obtain a face image dataset, where the face image dataset includes face images obtained in the visible light mode and face images obtained in the near-infrared mode; Use a face key point detection algorithm to perform key point detection on the face images in the face image dataset; Extract multiple key points included in each face image from the face image dataset, and select a sub-image where the target key point is located from the multiple key points; Input different combinations formed by the face image dataset and the sub-images where the multiple target key points corresponding to each face image in the face image dataset are located into a target neural network model; Use the classification module in the target neural network model to determine the association degree between each face image and each corresponding combination; Train the target neural network model according to the association degree between each face image and each corresponding combination until a face recognition model for face recognition in the visible light mode and the near-infrared mode is obtained.
2. The method according to claim 1, characterized in that, After using the face key point detection algorithm to perform key point detection on the face images in the face image dataset, and before extracting multiple key points included in each face image from the face image dataset and selecting a sub-image where the target key point is located from the multiple key points, the method includes: Perform face alignment operation by correcting the key points in the recognized face image according to a preset face template.
3. The method according to claim 1, characterized in that Before inputting different combinations formed by the face image dataset and the sub-images where the multiple target key points corresponding to each face image in the face image dataset are located into the target neural network model, the method further includes: Obtain a pre-trained face image dataset; Determine a pre-trained neural network model using the pre-trained face image dataset; Train the pre-trained neural network model using the obtained face image dataset to obtain the target neural network model.
4. The method according to claim 1, characterized in that Training the target neural network model includes: Determine the loss value of the loss function corresponding to the target neural network model during the training process until the loss value of the loss function meets a preset condition to obtain a face recognition model for face recognition in the visible light mode and the near-infrared mode.
5. An apparatus for constructing a cross-modal face recognition model, characterized in that, It includes: An acquisition module, configured to obtain a face image dataset, where the face image dataset includes face images obtained in the visible light mode and face images obtained in the near-infrared mode; A detection module, configured to use a face key point detection algorithm to perform key point detection on the face images in the face image dataset; A selection module, configured to extract multiple key points included in each face image from the face image dataset, and select a sub-image where the target key point is located from the multiple key points; An input module, configured to input different combinations formed by the face image dataset and the sub-images where the multiple target key points corresponding to each face image in the face image dataset are located into a target neural network model; A determination module, configured to determine the association degree between each face image and each corresponding combination by using the classification module in the target neural network model; A training module, configured to train the target neural network model according to the association degree between each face image and each corresponding combination until a face recognition model for face recognition in visible light mode and near-infrared mode is obtained.
6. The device according to claim 5, characterized in that, The device further includes: An alignment module, configured to correct key points in the recognized face image according to a preset face template to complete face alignment operation.
7. The device according to claim 5, characterized in that The device further includes: A first acquisition module, configured to acquire a pre-trained face image dataset; A first determination module, configured to determine a pre-trained neural network model by using the pre-trained face image dataset; A first training module, configured to train the pre-trained neural network model by using the acquired face image dataset to obtain the target neural network model.
8. The device according to claim 5, characterized in that The device further includes: A second determination module, configured to determine the loss value of the loss function corresponding to the target neural network model during training until the loss value of the loss function meets a preset condition to obtain a face recognition model for face recognition in visible light mode and near-infrared mode.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions, and the computer instructions are used to cause the computer to execute the cross-modal face recognition model construction method according to any one of claims 1-4.
10. An electronic device, characterized in that, Including: A memory and a processor, the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the computer instructions to execute the cross-modal face recognition model construction method according to any one of claims 1-4.
Citation Information
Patent Citations
A multi-modal feature fusion method and device based on a convolutional neural network
CN109583569A
Face recognition method, device and electronic equipment, and computer non-volatile readable storage medium
US20210312163A1