Two-dimensional image assisted three-dimensional face recognition method, system and device and medium

By reconstructing 3D face data from 2D images and converting it into normal component maps, the problems of scarcity and high cost of 3D face recognition data are solved, and an efficient and accurate 3D face recognition system is realized.

CN120635968APending Publication Date: 2025-09-12SECOND AFFILIATED HOSPITAL OF COLLEGE OF MEDICINEOF XIAN JIAOTONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510923378.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-04
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

Existing 3D face recognition technology suffers from data scarcity, insufficient generation quality, and high cost of relying on real 3D scanning, making it difficult to achieve efficient deployment and recognition in data-scarce scenarios.

Method used

By reconstructing high-quality 3D face data from large-scale 2D face images and converting them into normal component maps, a 2D convolutional network is used for training to generate structured normal component maps for feature extraction and recognition.

Benefits of technology

It achieves high-precision, low-cost 3D face recognition without real 3D training samples, improves recognition accuracy and ease of deployment, and reduces the difficulty and cost of data collection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120635968A_ABST
    Figure CN120635968A_ABST
Patent Text Reader

Abstract

The invention discloses a two-dimensional image-assisted three-dimensional face recognition method, system and device and a medium, and the recognition method comprises the steps: obtaining 3D faces of a registration set and a to-be-recognized set; generating normal component diagrams in corresponding directions for the 3D faces of the registration set and the to-be-recognized set; utilizing a 3D face recognition model to extract 3D face depth features of the registration set and the to-be-recognized set according to the normal component graph; calculating the distance between the 3D face depth features of the to-be-recognized set and the registration set; fusing the distances between the 3D face depth features to obtain a fused distance; and regarding the category of the 3D face data in the registration set with the minimum fusion distance as a recognition result of the 3D face data in the to-be-recognized set. According to the method, a visualized 3D representation mode (normal component graph) is adopted as model input, so that the 2D convolutional neural network can be directly used for a 3D recognition task, reconstructed 3D face data and a 2D face recognition pre-training model are linked, and a complex point cloud processing or graph network structure does not need to be introduced. According to the design, an existing pre-training model can be reused, the development and training cost is reduced, good platform compatibility and engineering deployment efficiency are achieved, and the 3D recognition technology can be popularized to practical application scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of face recognition technology, and in particular relates to a two-dimensional image-assisted three-dimensional face recognition method, system, device and medium. Background Art

[0002] Face recognition, a key research area in computer vision and biometrics, has been widely applied in fields such as security, identity authentication, and smart devices. With the development of deep convolutional neural networks (DCNNs), face recognition based on 2D (two-dimensional) images has achieved significant breakthroughs. This is primarily due to the availability of massive publicly available 2D face image datasets. For example, FaceNet, trained on over 200 million images, significantly improved recognition accuracy and generalization. In contrast, 3D (three-dimensional) face recognition offers advantages such as robustness to changes in illumination and pose, making it suitable for identification in complex scenarios. However, the development of deep learning-based 3D face recognition technology is severely constrained by the scarcity of high-quality 3D face datasets, the high cost of acquiring them, the complex acquisition process, and the reliance on specialized equipment and controlled environments. For example, the commonly used FRGCv2 (Face Recognition Grand Challenge version 2) and Bosphorus datasets contain only a few thousand 3D samples, far from meeting the large-scale training data requirements of deep models.

[0003] To alleviate the scarcity of 3D data, researchers have proposed a variety of enhancement strategies, including transformation-based enhancements (such as rotation, scaling, and adding noise) and synthesis-based enhancements (such as 3D morphable models and generative adversarial models). However, most of these methods are still limited to the 3D domain and struggle to overcome the limitations of the original 3D data itself. In recent years, with the development of single-image 3D reconstruction technology, a new direction has emerged: reconstructing 3D faces from a large number of 2D face images, thereby overcoming the shortcomings and deficiencies of existing technologies in a low-cost, large-scale manner.

[0004] Current 3D face recognition enhancement methods generally have the following problems and shortcomings:

[0005] (1) Data scarcity remains prominent: Existing enhancement methods, such as geometric transformation or deformable modeling (3DMM (3DMorphable Model), GPMM (geometric process maintenance model)), although they can expand sample diversity to a certain extent, are unable to generate new samples with significant identity differences, resulting in limited deep model training and insufficient generalization capabilities. Faces generated by linear models such as 3DMM often lack realism and have limited expression richness. The application of methods such as generative adversarial methods in the 3D field is challenged by input alignment, label quality, and expression authenticity, and it is difficult to guarantee recognition performance by generating samples.

[0006] (2) Reliance on supervised information from real 3D scans: Most existing methods rely on real 3D scan data as a training or verification benchmark, which limits the deployability of the model in data-scarce scenarios and makes it difficult to achieve cost control.

[0007] (3) The potential of 3D reconstruction data has not been systematically verified and explored: Although existing methods can reconstruct 3D facial shapes from 2D images, most methods have not conducted in-depth research on the actual contribution of these reconstructed 3D data in 3D face recognition training, and lack systematic performance verification and analysis.

[0008] In summary, existing technologies find it difficult to strike a balance between data scale, generation quality, recognition performance and training cost, which becomes a key bottleneck for the further development of 3D face recognition. Summary of the Invention

[0009] In order to overcome the scarcity of 3D face recognition training data and limited data enhancement methods in the existing technology, the purpose of the present invention is to propose a two-dimensional image-assisted three-dimensional face recognition method, system, device and medium. This method can utilize large-scale existing 2D face images to generate high-quality 3D face data through efficient reconstruction, and convert the 3D face data into a structured normal component map to input into a standard 2D convolutional network for training. In this way, even without real 3D training samples, a high-precision and easy-to-deploy 3D face recognition system can be achieved.

[0010] To achieve the above object, the technical solution adopted by the present invention is as follows:

[0011] A two-dimensional image-assisted three-dimensional face recognition method comprises the following steps:

[0012] Get the 3D faces of the registration set and the set to be recognized;

[0013] Generate normal component maps of the 3D faces of the registration set and the set to be identified;

[0014] Using the 3D face recognition model, the 3D face depth features of the registration set and the set to be recognized are extracted according to the normal component map;

[0015] Calculate the distance between the 3D face depth features of the set to be identified and the registered set;

[0016] Fuse the distances between 3D face depth features to obtain the fusion distance;

[0017] The 3D face data category in the registration set with the smallest fusion distance is regarded as the result of recognition of the 3D face data in the recognition set.

[0018] Furthermore, the 3D face recognition model is determined through the following process:

[0019] Based on the 2D face image, pre-train the deep face recognition model to obtain the pre-trained 2D image face recognition model;

[0020] Reconstruct a 3D face model using a 2D image to obtain 3D face data;

[0021] The 3D face data is converted into a normal component map, and the pre-trained 2D image face recognition model is fine-tuned to obtain a 3D face recognition model.

[0022] Furthermore, a normal component map is generated for the 3D faces of the registration set and the set to be identified, including:

[0023] The 3D face data is mapped into a structured image of size m×n×3. Each pixel in the structured image is a 3D point coordinate (x, y, z). Within each pixel neighborhood, a local plane is fitted and the unit normal vector of each point is calculated. m is the number of sampling points along the vertical direction, and n is the number of sampling points along the horizontal direction.

[0024] Split the unit normal vector of each point into normal components in the x, y, and z directions to form a normal component graph in the x, y, and z directions;

[0025] Using the loss function and gradient backpropagation, the normal component maps in the x, y, and z directions are respectively fine-tuned to the pre-trained 2D image face recognition model to obtain the 3D face recognition model in the x, y, and z directions.

[0026] Furthermore, based on the 2D face image, a deep face recognition model is pre-trained to obtain a pre-trained 2D image face recognition model, including: using a 2D face image dataset, pre-training the SphereFace network architecture using a loss function, and obtaining a pre-trained 2D image face recognition model.

[0027] Furthermore, the 2D image is used to reconstruct a 3D face model to obtain 3D face data, including:

[0028] The neural network ExpNet is used to regress 3D facial shape parameters and expression parameters from a single 2D face image, complete 3D face reconstruction, and obtain 3D face data.

[0029] Furthermore, using the 3D face recognition model, the 3D face depth features of the registration set and the set to be recognized are extracted according to the normal component map, using the following formula:

[0030]

[0031] in, is the x-direction depth feature of the 3D face data to be identified, is the x-direction depth feature of the 3D face data in the registration set, f x is the 3D face recognition model in the x direction, is the x-direction normal component map of the 3D face data to be identified, is the x-direction normal component map of the 3D face data in the registration set;

[0032]

[0033] in, is the y-direction depth feature of the 3D face data to be identified, is the y-direction depth feature of the 3D face data in the registration set, f y It is the 3D face recognition model in the y direction. is the y-direction normal component map of the 3D face data to be identified, is the y-direction normal component map of the 3D face data in the registration set;

[0034]

[0035] in, is the z-direction depth feature of the 3D face data to be identified, is the z-direction depth feature of the 3D face data in the registration set, f z It is the 3D face recognition model in the z direction. is the z-direction normal component map of the 3D face data to be identified, It is the z-direction normal component map of the 3D face data in the registration set.

[0036] Furthermore, the distance between the 3D face depth features of the to-be-recognized set and the registered set is calculated by the following process:

[0037] Calculate the distance d between the depth features in the x direction x :

[0038]

[0039] Among them, d x is the distance between depth features in the x direction;

[0040]

[0041] Among them, d y is the distance between depth features in the y direction;

[0042]

[0043] Among them, d z is the distance between depth features in the z direction;

[0044] The fusion distance is calculated as follows:

[0045] d=1 / 3(d x +d y +d z )

[0046] Where d is the fusion distance.

[0047] A two-dimensional image-assisted three-dimensional face recognition system, comprising:

[0048] 3D face acquisition module, used to obtain 3D faces in the registration set and the set to be recognized;

[0049] A normal component map generation module is used to generate a normal component map of the 3D faces in the registration set and the set to be identified;

[0050] The deep feature extraction module is used to extract the 3D face deep features of the registration set and the set to be recognized based on the normal component map using the 3D face recognition model;

[0051] The distance calculation module between depth features is used to calculate the distance between the depth features of the 3D faces in the set to be identified and the registered set;

[0052] The fusion module is used to fuse the distances between 3D face depth features to obtain the fusion distance;

[0053] The face recognition result confirmation module is used to regard the 3D face data category in the registration set with the smallest fusion distance as the result of the recognition of the 3D face data in the set to be recognized.

[0054] An electronic device includes a memory, a processor, and a computer program stored in the memory and runnable on the processor. When the processor executes the computer program, the two-dimensional image-assisted three-dimensional face recognition method is implemented.

[0055] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the two-dimensional image-assisted three-dimensional face recognition method.

[0056] Compared with the prior art, the present invention has the following beneficial effects:

[0057] The present invention extracts fine-grained geometric features in three-dimensional space by converting the reconstructed 3D face into a normal component map. These normal component maps can accurately describe the changes in local facial curvature, enhance the recognition ability and geometric discriminability of the model, and effectively make up for the shortcomings of 2D image features in spatial structure perception. Experiments show that the Rank-1 recognition accuracy of the model trained only with reconstructed data significantly exceeds that of the baseline model, verifying the significant improvement of the identity discriminability of the present invention. The present invention adopts an image-based 3D representation (normal component map) as model input, so that a 2D convolutional neural network (such as SphereFace) can be directly used for 3D recognition tasks, linking the reconstructed 3D face data with the 2D face recognition pre-training model without introducing complex point cloud processing or graph network structure. This design can not only reuse existing pre-trained models and reduce development and training costs, but also has good platform compatibility and engineering deployment efficiency, which helps to promote 3D recognition technology to practical application scenarios. The present invention is compatible with existing 2D recognition network structures and has the advantages of simple deployment and strong portability.

[0058] Furthermore, this invention reconstructs 3D facial data from 2D facial images, replacing the traditional reliance on high-precision 3D scanning. This significantly reduces the difficulty and cost of data acquisition and enables large-scale training, providing a solution to the critical bottleneck of "scarce training data" in the field of 3D face recognition. The reconstruction model used (such as ExpNet) can efficiently generate large-scale 3D data from public 2D face libraries (such as VGGFace2), demonstrating good scalability and engineering feasibility. BRIEF DESCRIPTION OF THE DRAWINGS

[0059] Figure 1 This is a flow chart of the 2D image-assisted 3D face recognition method proposed in the present invention;

[0060] Figure 2 are the normal component diagrams in the x, y, and z directions, where (a) is the normal component diagram in the x direction, (b) is the normal component diagram in the y direction, and (c) is the normal component diagram in the z direction;

[0061] Figure 3 This is a flow chart of the two-dimensional image-assisted three-dimensional face recognition method of the present invention;

[0062] Figure 4 Schematic diagram of a two-dimensional image-assisted three-dimensional face recognition system. DETAILED DESCRIPTION

[0063] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.

[0064] It should be noted that the terms "first", "second", etc. in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0065] See also Figure 1 and Figure 3 The present invention provides a two-dimensional image-assisted three-dimensional face recognition method, which includes four key stages:

[0066] Get the 3D faces of the registration set and the set to be recognized;

[0067] Generate normal component maps of the 3D faces of the registration set and the set to be identified;

[0068] Using the 3D face recognition model, the 3D face depth features of the registration set and the set to be recognized are extracted according to the normal component map;

[0069] Calculate the distance between the 3D face depth features of the set to be identified and the registered set;

[0070] Fuse the distances between 3D face depth features to obtain the fusion distance;

[0071] The 3D face data category in the registration set with the smallest fusion distance is regarded as the result of recognition of the 3D face data in the recognition set.

[0072] The specific steps are as follows: 1) pre-training a deep face recognition model based on large-scale 2D face images to obtain a pre-trained model;

[0073] This paper first selects a large-scale 2D face image dataset, such as VGGFace2, and uses the SphereFace network architecture as a face recognition model for pre-training. This yields a pre-trained face recognition model for 2D images. Specifically, a 20-layer SphereFace convolutional neural network, denoted as Sphere20, is used, which has strong representation learning capabilities. The input image is an RGB image of size 112×112. The loss function is the cross-entropy loss:

[0074]

[0075] Among them, y ij is the true label of the i-th sample in the j-th category, is the output probability of the softmax layer of the convolutional neural network.

[0076] The purpose of pre-training is to learn robust facial features from rich 2D images, laying the foundation for subsequent migration to the 3D field.

[0077] 2) Use 2D images to reconstruct large-scale, high-quality 3D face models to obtain 3D face data;

[0078] This paper uses the neural network ExpNet model to directly regress 3D facial shape and expression parameters from a single 2D face image, completing 3D face reconstruction and obtaining 3D face data. This process avoids the computational complexity and error accumulation problems of traditional 3D deformable model fitting. The reconstruction process is based on the 3D Morphable Model (3DMM) expression:

[0079]

[0080] in, Indicates the shape of the average face of a 3D face, S represents the shape principal component, E represents the shape principal component, and They represent the 3D face shape parameters and expression parameters output by the neural network ExpNet model respectively.

[0081] The purpose of 3D face reconstruction is to obtain 3D face data from massive 2D images, thereby providing data for the training of 3D face recognition models.

[0082] 3) Extracting a structured normal component map based on the 3D face data and using it to fine-tune the pre-trained model to obtain a 3D face recognition model;

[0083] In order to enable the above-mentioned pre-trained convolutional neural network Sphere20 to effectively process the reconstructed 3D face data, the present invention converts the reconstructed 3D face data into a normal component map for fine-tuning the pre-trained network.

[0084] First, the reconstructed 3D face data is mapped into a structured image P of size m×n×3, where m is the number of sampling points along the vertical direction, n is the number of sampling points along the horizontal direction, and each pixel is a 3D point coordinate (x, y, z). In each pixel neighborhood (5×5 window), a local plane S is fitted. ij , calculate the normal vector Split the unit normal vector of each point into normal components in the x, y, and z directions to form three two-dimensional images N x 、N y 、N z , respectively, represent the normal component diagrams in the x, y, and z directions, see Figure 2 (a), (b), (c) and (d).

[0085] Using the loss function and gradient back propagation, the normal component map in the x direction is fine-tuned to the pre-trained 2D image face recognition model to obtain the 3D face recognition model f in the x direction. x ;

[0086] Using the loss function, after gradient back propagation, the normal component map in the y direction is fine-tuned to the pre-trained 2D image face recognition model to obtain the 3D face recognition model f in the y direction. y ;

[0087] Using the loss function, after gradient back propagation, the normal component map in the z direction is fine-tuned to the pre-trained 2D image face recognition model to obtain the 3D face recognition model f in the z direction. z .

[0088] The loss function adopts cross entropy loss.

[0089] 4) Perform 3D face recognition based on multi-directional feature fusion through a 3D face recognition model;

[0090] In the test phase, the present invention uses the 3D face recognition model in the x, y, and z directions to generate the corresponding x, y, and z direction normal component maps N for the 3D faces of the registration set (a data set with known identities, which the system uses as a comparison benchmark) and the 3D faces of the to-be-recognized set (samples with unknown identities, used for testing or recognition). x 、N y 、N z , extract the depth features in the x direction:

[0091]

[0092] in, is the x-direction depth feature of the 3D face data to be identified, is the x-direction depth feature of the 3D face data in the registration set, fx is the 3D face recognition model in the x direction, is the x-direction normal component map of the 3D face data to be identified, It is the normal component map of the x-direction of the 3D face data in the registration set.

[0093] Extract the depth feature in the y direction:

[0094]

[0095] in, is the y-direction depth feature of the 3D face data to be identified, is the y-direction depth feature of the 3D face data in the registration set, f y It is the 3D face recognition model in the y direction. is the y-direction normal component map of the 3D face data to be identified, It is the y-direction normal component map of the 3D face data in the registration set.

[0096] Extract depth features in the z direction:

[0097]

[0098] in, is the z-direction depth feature of the 3D face data to be identified, is the z-direction depth feature of the 3D face data in the registration set, f z It is the 3D face recognition model in the z direction. is the z-direction normal component map of the 3D face data to be identified, It is the z-direction normal component map of the 3D face data in the registration set.

[0099] Calculate the distance d between the depth features in the x direction x :

[0100]

[0101] Calculate the distance d between the depth features in the y direction y :

[0102]

[0103] Calculate the distance d between the depth features in the z direction x :

[0104]

[0105] Calculate the characteristic distance d in the x direction accordingly x , characteristic distance d in the y direction y, characteristic distance d in the z direction z , the feature distances in the three directions are graded and fused to obtain the fusion distance:

[0106] d=1 / 3(d x +d y +d z )

[0107] Where d is the fusion distance.

[0108] The 3D face data category in the registration set with the smallest fusion distance is regarded as the result of the recognition of the 3D face data in the recognition set, and the 3D face recognition is completed.

[0109] 5) Experimental verification

[0110] This paper presents a comprehensive evaluation of the proposed method. First, we introduce the training and testing datasets and the evaluation protocol. Then, we verify the effectiveness of reconstructing 3D faces in recognition tasks. Finally, we compare the proposed method with a baseline model.

[0111] 5.1) Dataset

[0112] The VGGFace2 dataset contains 3.31 million face images from 9,131 identities. The images are sourced from Google and cover a wide range of age, pose, lighting, ethnicity, and occupation. Also provided are face bounding boxes, five key points, and estimated age and pose labels. Standard data augmentation (such as horizontal flipping and random crops) is applied during training.

[0113] The BU-3DFE dataset consists of 2,500 3D face scans from 100 subjects. Each subject displays six typical facial expressions (happiness, disgust, fear, anger, surprise, and sadness), each with four intensity levels, as well as a neutral expression. The neutral scans are used to construct the image library, while the remaining 2,400 scans are used for testing. This dataset is challenging due to its large expression variation.

[0114] The FRGC v2 dataset contains 4007 textured 3D face scans from 466 people, including 1642 samples with different expressions, all collected under controlled lighting conditions. The first scan of each subject is used as the gallery, and the rest are used as test samples.

[0115] The Bosphorus dataset contains 4,666 3D scans of 105 subjects (60 men and 45 women), covering a wide range of poses, expressions, occlusions, and other variations. This study used a subset of 2,902 scans containing only expression variations. The first neutral scan of each subject served as the library, while the remaining 2,797 3D faces served as the test set.

[0116] In the experiments, the VGGFace2 dataset was used for training. The Sphere20 network was first pre-trained using raw 2D images. 3D faces were then reconstructed from these images and converted into normal component maps for model fine-tuning. No real 3D data was used during training; the model was trained entirely on 2D images and the 3D information generated from them.

[0117] 5.2) Evaluation of the effectiveness of the proposed method

[0118] To evaluate the effectiveness of the proposed 2D-assisted 3D face recognition method, we verified the contribution of the reconstructed 3D face to recognition performance. During pre-training, 112×112 RGB images from VGGFace2 were fed into the Sphere20 network, which output a 9131-dimensional class probability vector. During fine-tuning, the fully connected and softmax layers were reinitialized and trained using a cross-entropy loss.

[0119] Tables 1, 2, and 3 compare the Rank-1 recognition rates of the baseline model (Sphere20 trained only with 2D images) and the proposed method on the BU-3DFE dataset. The experimental results show that although the model is trained only on 3D data generated from 2D images and does not use any real 3D face training data, its recognition accuracy on the three public datasets significantly exceeds that of the baseline model trained with 2D images, demonstrating the effectiveness of the proposed 2D-assisted 3D face recognition method.

[0120] Table 1. Face recognition evaluation results of BU-3DFE dataset

[0121]

[0122] Table 2 Face recognition evaluation results of FRGC v2 dataset

[0123]

[0124] Table 3. Bosphorus dataset face recognition evaluation results

[0125]

[0126] The present invention generates 3D face data based on 2D images for training 3D face recognition. It uses 2D images to reconstruct large-scale 3D face models instead of real 3D scans, achieving a low-cost and high-efficiency data generation solution.

[0127] The 3D face recognition method of the present invention combines a multi-branch network structure with directional feature fusion, processes the normal component maps of the three directions respectively through a three-branch network, and fuses them at the decision layer to improve the recognition accuracy and system robustness.

[0128] See also Figure 4 The two-dimensional image-assisted three-dimensional face recognition system of the present invention comprises:

[0129] 3D face acquisition module, used to obtain 3D faces in the registration set and the set to be recognized;

[0130] A normal component map generation module is used to generate a normal component map of the 3D faces in the registration set and the set to be identified;

[0131] The deep feature extraction module is used to extract the 3D face depth features of the registration set and the set to be recognized based on the normal component map using the 3D face recognition model;

[0132] The distance calculation module between depth features is used to calculate the distance between the depth features of the 3D faces in the to-be-recognized set and the registered set;

[0133] The fusion module is used to fuse the distances between 3D face depth features to obtain the fusion distance;

[0134] The face recognition result confirmation module is used to regard the 3D face data category in the registration set with the smallest fusion distance as the result of the recognition of the 3D face data in the set to be recognized.

[0135] All relevant contents of each step involved in the embodiment of the aforementioned two-dimensional image-assisted three-dimensional face recognition method can be referred to the functional description of the functional modules corresponding to the two-dimensional image-assisted three-dimensional face recognition system in the embodiment of the present invention, and will not be repeated here.

[0136] In one embodiment of the present invention, an electronic device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the 2D-assisted deep 3D face recognition method when executing the computer program.

[0137] In one embodiment of the present invention, a computer-readable storage medium is provided, specifically a computer-readable storage medium (Memory), which is a memory device in a computer device for storing programs and data. It is understandable that the computer-readable storage medium here can include both built-in storage media in the computer device and, of course, extended storage media supported by the computer device. The computer-readable storage medium provides a storage space, which stores the operating system of the terminal. In addition, one or more instructions suitable for being loaded and executed by the processor are also stored in the storage space. These instructions can be one or more computer programs (including program codes). It should be noted that the computer-readable storage medium here can be a high-speed RAM memory or a non-volatile memory, such as at least one disk memory. The processor can load and execute one or more instructions stored in the computer-readable storage medium to implement the 2D-assisted deep 3D face recognition method in the above embodiment.

[0138] It will be understood by those skilled in the art that embodiments of the present invention may be provided as methods, systems, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0139] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0140] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1The function specified in one or more boxes.

[0141] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0142] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to the above embodiments, ordinary technicians in the field should understand that the specific implementation methods of the present invention can still be modified or replaced by equivalents. Any modification or equivalent replacement that does not depart from the spirit and scope of the present invention should be covered by the scope of protection of the claims of the present invention.

Claims

1. A two-dimensional image-assisted three-dimensional face recognition method, characterized in that: The following steps are involved: Get the 3D faces of the registration set and the set to be recognized; Generate normal component maps of the 3D faces of the registration set and the set to be identified; Using the 3D face recognition model, the 3D face depth features of the registration set and the set to be recognized are extracted according to the normal component map; Calculate the distance between the 3D face depth features of the set to be identified and the registered set; Fuse the distances between 3D face depth features to obtain the fusion distance; The 3D face data category in the registration set with the smallest fusion distance is regarded as the result of recognition of the 3D face data in the recognition set.

2. The two-dimensional image-assisted three-dimensional face recognition method according to claim 1, characterized in that: The 3D face recognition model is determined through the following process: Based on the 2D face image, pre-train the deep face recognition model to obtain the pre-trained 2D image face recognition model; Reconstruct a 3D face model using a 2D image to obtain 3D face data; The 3D face data is converted into a normal component map, and the pre-trained 2D image face recognition model is fine-tuned to obtain a 3D face recognition model.

3. The two-dimensional image-assisted three-dimensional face recognition method according to claim 2, characterized in that: Generate a normal component map of the 3D faces of the registration set and the set to be identified, including: The 3D face data is mapped into a structured image of size m×n×3. Each pixel in the structured image is a 3D point coordinate (x, y, z). Within each pixel neighborhood, a local plane is fitted and the unit normal vector of each point is calculated. m is the number of sampling points along the vertical direction, and n is the number of sampling points along the horizontal direction. Split the unit normal vector of each point into normal components in the x, y, and z directions to form a normal component graph in the x, y, and z directions; Using the loss function and gradient backpropagation, the normal component maps in the x, y, and z directions are respectively fine-tuned to the pre-trained 2D image face recognition model to obtain the 3D face recognition model in the x, y, and z directions.

4. The two-dimensional image-assisted three-dimensional face recognition method according to claim 1, characterized in that: Based on 2D face images, a deep face recognition model is pre-trained to obtain a pre-trained 2D image face recognition model, including: using a 2D face image dataset, pre-training the SphereFace network architecture using a loss function, and obtaining a pre-trained 2D image face recognition model.

5. The two-dimensional image-assisted three-dimensional face recognition method according to claim 2, characterized in that: Use 2D images to reconstruct a 3D face model and obtain 3D face data, including: The neural network ExpNet is used to regress 3D facial shape parameters and expression parameters from a single 2D face image, complete 3D face reconstruction, and obtain 3D face data.

6. The two-dimensional image-assisted three-dimensional face recognition method according to claim 1, characterized in that: Using the 3D face recognition model, the 3D face depth features of the registration set and the set to be recognized are extracted according to the normal component map, using the following formula: in, is the x-direction depth feature of the 3D face data to be identified, is the x-direction depth feature of the 3D face data in the registration set, f x is the 3D face recognition model in the x direction, is the x-direction normal component map of the 3D face data to be identified, is the x-direction normal component map of the 3D face data in the registration set; in, is the y-direction depth feature of the 3D face data to be identified, is the y-direction depth feature of the 3D face data in the registration set, f y It is the 3D face recognition model in the y direction. is the y-direction normal component map of the 3D face data to be identified, is the y-direction normal component map of the 3D face data in the registration set; in, is the z-direction depth feature of the 3D face data to be identified, is the z-direction depth feature of the 3D face data in the registration set, f z It is the 3D face recognition model in the z direction. is the z-direction normal component map of the 3D face data to be identified, It is the z-direction normal component map of the 3D face data in the registration set.

7. The two-dimensional image-assisted three-dimensional face recognition method according to claim 1, characterized in that: Calculate the distance between the 3D face depth features of the set to be identified and the registered set through the following process: Calculate the distance d between the depth features in the x direction x : Among them, d x is the distance between depth features in the x direction; Among them, d y is the distance between depth features in the y direction; Among them, d z is the distance between depth features in the z direction; The fusion distance is calculated as follows: d=1 / 3(d x +d y +d z ) Where d is the fusion distance.

8. A two-dimensional image-assisted three-dimensional face recognition system, characterized in that: include: 3D face acquisition module, used to obtain 3D faces in the registration set and the set to be recognized; A normal component map generation module is used to generate a normal component map of the 3D faces in the registration set and the set to be identified; The deep feature extraction module is used to extract the 3D face depth features of the registration set and the set to be recognized based on the normal component map using the 3D face recognition model; The distance calculation module between depth features is used to calculate the distance between the depth features of the 3D faces in the to-be-recognized set and the registered set; The fusion module is used to fuse the distances between 3D face depth features to obtain the fusion distance; The face recognition result confirmation module is used to regard the 3D face data category in the registration set with the smallest fusion distance as the result of the recognition of the 3D face data in the set to be recognized.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the two-dimensional image-assisted three-dimensional face recognition method as described in any one of claims 1 to 7 is implemented.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the two-dimensional image-assisted three-dimensional face recognition method as described in any one of claims 1 to 7 is implemented.