High-precision face matching method and system based on deep learning driving
Through the high-precision face matching method based on deep learning, the data dependence, computing resource requirements, real-time and accuracy of the face recognition system in the prior art are solved, and an efficient, accurate and secure face matching system is realized.
Patent Information
- Application Number
- CN202510595683.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-09
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-05-09
AI Technical Summary
In the prior art, face recognition systems have problems such as strong data dependence, high computing resource requirements, difficult to balance real-time and accuracy, and insufficient security.
High-precision face matching method based on deep learning is adopted to collect sample data, establish sample expansion models, perform sample expansion and standardization, and train matching models for face matching. The system includes image acquisition, preprocessing, face matching and feedback modules, and uses the VGGNet model to perform feature extraction and matching recognition.
It effectively reduces the cost of data labeling and computing resource requirements, improves the accuracy and robustness of facial recognition, enhances the system's real-time response capabilities and security, and is suitable for security, access control and finance fields.
Smart Images

Figure CN120108025A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image or video recognition or understanding, and in particular to a high-precision face matching method and system driven by deep learning. Background Art
[0002] Face recognition is a technology that uses facial features to identify people. It usually collects images or video streams containing faces to detect and track faces and achieve face recognition. With the rapid development of information technology and the increasing demand for security in society, face recognition technology, as a convenient and efficient biometric identification method, has gradually been widely used in security monitoring, access control, financial transactions and other fields.
[0003] Early face recognition systems mainly relied on hand-crafted feature extraction methods, such as local binary pattern (LBP), principal component analysis (PCA) and linear discriminant analysis (LDA). Although these methods can achieve certain results under certain conditions, their performance significantly decreases when faced with complex environmental changes, such as lighting changes, posture differences, occlusion, etc. At the same time, the means of adversarial face recognition are also changing with each passing day. The original face recognition methods and systems are difficult to meet the needs of practical applications. In recent years, the development of deep learning technology, especially convolutional neural networks (CNN), has provided a new way to solve these problems. Methods based on deep learning can automatically learn richer and more discriminative feature representations from a large amount of data, thereby greatly improving the accuracy and robustness of face recognition. As one of the classic CNN architectures, VGGNet has performed well in multiple computer vision tasks such as image classification and object detection due to its simple and effective structural design, and is widely used in the field of face recognition.
[0004] Although deep learning-based face recognition systems have made significant progress, they still face some challenges, including: (1) High data dependence: The training of high-quality models requires the support of large-scale labeled data sets. The cost of data acquisition and labeling is high, especially when privacy protection is involved, making data collection more difficult. (2) High computing resource requirements: Deep learning models usually require a large amount of computing resources for training and inference, and high-precision models often contain millions or even hundreds of millions of parameters, which places high demands on hardware devices and limits their application in resource-constrained environments. (3) It is difficult to strike a balance between real-time performance and accuracy. To achieve real-time performance, many application scenarios require that the face matching process be completed within a limited time, which may result in reduced recognition accuracy. How to improve the system's real-time response capability while ensuring high accuracy is an urgent problem to be solved. (4) Impact on security: Traditional face recognition systems are vulnerable to spoofing attacks, such as using photos or videos to impersonate legitimate users. In addition to integrating liveness detection technology to enhance the security of the system, it becomes particularly important to better identify the corresponding synthetic photos and videos. Summary of the invention
[0005] The present invention solves the problems existing in the prior art and provides a high-precision face matching method and system driven by deep learning.
[0006] The technical solution adopted by the present invention is a high-precision face matching method driven by deep learning, which comprises the following steps: S1, collect sample data and preprocess them to obtain sample data pairs; S2. Establish a sample expansion model to expand the sample data pairs until the preset conditions are met; S3, normalize the sample data pairs processed by S2 using the constructed rules; S4, training the input matching model with the sample data processed by S3; S5. Use the trained matching model for face matching.
[0007] Preferably, in S1, sample data is collected and invalid data is cleaned; a unique identification code is assigned to the cleaned sample data based on user information; Create mixed positive sample data pairs and mixed negative sample data pairs.
[0008] Preferably, the cleaned sample data is randomly matched, and matching thresholds α and β are set, α>β>0; The sample data pairs whose matching similarity results are greater than the threshold α and whose unique identification codes are different are regarded as fixed negative sample data pairs; The sample data pairs whose matching similarity results are less than the threshold β and whose unique identification codes are the same are fixed positive sample data pairs; Randomly select a number of sample data pairs whose matching similarity results are less than a threshold α and greater than a threshold β, which are random negative sample data pairs and random positive sample data pairs; A fixed positive sample data pair and a random positive sample data pair are considered as a mixed positive sample data pair; a fixed negative sample data pair and a random negative sample data pair are considered as a mixed negative sample data pair.
[0009] Preferably, the mixed negative sample data pairs also include filtered negative sample data pairs established based on associated users after extracting user information.
[0010] Preferably, the sample expansion model comprises a generator and a discriminator connected in sequence, and the generator comprises a feature extraction module and a synthetic feature generation module connected in sequence.
[0011] Preferably, the feature extraction module includes a plurality of convolution blocks and a plurality of fully connected layers connected in sequence, and any of the convolution blocks outputs corresponding features and results; The synthetic feature generation module includes a number of synthetic blocks connected in sequence, any synthetic block inputs the result of the previous level, the corresponding convolution block outputs the corresponding features and random noise, and up-sampling is performed.
[0012] Preferably, the sample data pairs processed in S2 include the same number of mixed positive sample data pairs and mixed negative sample data pairs.
[0013] Preferably, in S3, normalizing the sample data includes sequentially performing a unified sample size, center cropping, a unified sample format, and a normalization process on the sample data.
[0014] Preferably, in S4, the matching model includes VGG16, a similarity calculation module and a binary classification module connected in sequence.
[0015] A high-precision face matching system based on deep learning drive, the system comprising: One or more image acquisition modules, used for acquiring images to be matched; A preprocessing module, used to normalize the matching images according to the constructed rules; A face matching module, which uses the high-precision face matching method based on deep learning to extract features and perform matching and recognition on faces; The feedback module corresponds to the image acquisition module setting and is used to feed back the matching results of the face matching module.
[0016] The present invention relates to a high-precision face matching method and system driven by deep learning. Sample data are collected and preprocessed. After obtaining sample data pairs, a sample expansion model is established to expand the sample data pairs until preset conditions are met; the processed sample data pairs are normalized according to constructed rules, the processed sample data pairs are input into a matching model for training, and the trained matching model is used for face matching; the system uses one or more image acquisition modules to acquire images to be matched, uses a preprocessing module to normalize the images to be matched according to the constructed rules, uses a face matching module of the method to extract features and perform matching recognition on faces, and finally uses a feedback module to feed back the matching results of the face matching module.
[0017] The beneficial effects of the present invention are: (1) Expand the sample acquisition method and establish a sample expansion model to solve the problem of high data demand and high dependence on annotation for training high-quality models. This can expand the data set in an orderly and efficient manner, thereby reducing the difficulty of training the matching model and effectively training the recognition of synthetic photos and videos. (2) Through multi-level feature learning and large-scale data training, the method and its system can effectively cope with illumination changes, posture differences and occlusion problems, improve recognition accuracy and robustness, and better prevent fraud attacks on the basis of integrated liveness detection, making it more suitable for security, access control and finance. (3) The sample expansion model and the matching model are separated, and the classic VGGNet model is used for actual matching training, testing and application to solve the problems of low recognition accuracy, poor robustness and insufficient security in the existing technology. The training difficulty is low, the computing resource requirements are low, the real-time performance is strong, and the accuracy is guaranteed. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 is a flow chart of the method of the present invention; Figure 2 It is a schematic structural diagram of the sample expansion model of the present invention; Figure 3 It is a schematic block diagram of the structure of the matching model of the present invention; Figure 4 It is a schematic diagram of the system structure of the present invention. DETAILED DESCRIPTION
[0019] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0020] like Figure 1 The present invention relates to a high-precision face matching method based on deep learning, and the method comprises the following steps: S1, collect sample data and preprocess them to obtain sample data pairs; S2. Establish a sample expansion model to expand the sample data pairs until the preset conditions are met; S3, normalize the sample data pairs processed by S2 using the constructed rules; S4, training the input matching model with the sample data processed by S3; S5. Use the trained matching model for face matching.
[0021] The method content is described below in conjunction with specific implementation methods.
[0022] S1, collect sample data and preprocess them to obtain sample data pairs; In this embodiment, the datasets used include the FIW (Families in the Wild) dataset, the KinFaceW-I dataset, and the KinFaceW-II dataset; in particular, the FIW dataset contains a large amount of data in which photos of human faces are grouped by person and by family. Considering the problem of misidentification of close relatives in face recognition, the existence of these sample data can help the method better train faces that are easily misidentified; During the implementation process, the collected sample data need to be cleaned of invalid data. The sample data here are mainly image data, so some sample data with obviously unrecognizable features, missing images or basic information are screened out.
[0023] Based on user information, a unique identification code is assigned to the cleaned sample data. Generally, the user ID is used as a label for comparison. In practical applications, the encoding of the unique identification code follows preset rules, such as assigning associated unique identification codes to members of the same family group. Create mixed positive sample data pairs and mixed negative sample data pairs.
[0024] In the present invention, mixed positive sample data pairs and mixed negative sample data pairs can effectively expand the amount of basic sample data. The positive sample data pair refers to two image samples of the same face (the person's image), and the negative sample data pair refers to two image samples of different faces (not the person's image). In the implementation process of this method, the mixed negative sample data pairs include fixed negative sample data pairs, random negative sample data pairs and screened negative sample data pairs; The mixed positive sample data pairs include fixed positive sample data pairs, random positive sample data pairs and positive sample data pairs expanded by a sample expansion model; Finally, the number of mixed positive sample data pairs and mixed negative sample data pairs is equal.
[0025] S1.1. Fixing of fixed positive sample data pairs and fixed negative sample data pairs; Random matching is performed on the cleaned sample data. This matching refers to machine matching of the cleaned sample data, such as extracting features with VGGNet and calculating the angle of feature vectors. In the specific implementation of the present invention, a double matching threshold, α and β, is set, α>β>0; If there is a sample data pair whose matching similarity result is greater than the threshold α and whose unique identification codes are different, it means that the two faces are very similar and need to be used as negative samples for training and adjusting the parameters of the matching model. Therefore, it is used as a fixed negative sample data pair. If there is a sample data pair whose matching similarity result is less than the threshold β and whose unique identification code is the same, it means that this may be a face recognition error caused by the same individual under different environmental parameters or objective factors. It is necessary to always use it as a positive sample for training and adjusting the parameters of the matching model, so it is used as a fixed positive sample data pair.
[0026] S1.2, obtaining random negative sample data pairs and random positive sample data pairs; The cleaned sample data is randomly matched and after removing the fixed positive sample data pairs and the fixed negative sample data pairs, a number of sample data pairs whose matching similarity results are less than the threshold α and greater than the threshold β are randomly selected from the sample pool as random negative sample data pairs and random positive sample data pairs; S1.3, obtaining the screening negative sample data pairs; The mixed negative sample data pair also includes a screening negative sample data pair established based on associated users after extracting user information, that is, the aforementioned "assigning associated unique identification codes to members of the same family group", and randomly extracting facial data of different members of the same family group to combine them into a screening negative sample data pair.
[0027] After completing the above processing, the number of negative sample data pairs may be too large, and the training data needs to be further expanded, so it is necessary to actively expand the sample data.
[0028] S2. Establish a sample expansion model to expand the sample data pairs until the preset conditions are met; In the implementation process of the present invention, the adversarial generative network is used as the backbone network for sample expansion; Figure 2 As shown, specifically, the sample expansion model includes a generator G and a discriminator (discriminator) D connected in sequence, and the generator includes a feature extraction module and a synthetic feature generation module connected in sequence; The feature extraction module includes several convolution blocks and several fully connected layers connected in sequence, such as 3 convolution blocks and 2 fully connected layers. Each convolution block includes a convolution layer, a normalization layer and an activation function. After the sample data passes through the feature extraction module, each convolution block outputs the corresponding features g1, g2, g3 and the result g4; The synthetic feature generation module includes several synthetic blocks connected in sequence, 3 in this case, each synthetic block includes a concatenation layer and an upsampling layer, each synthetic block inputs the result of the previous level, and the corresponding convolution block outputs the corresponding features and random noise. The first synthetic block inputs g4, feature g3 and random noise z, and outputs P1, the second synthetic block inputs g3', feature g2 and random noise z, and outputs P2, the third synthetic block inputs g2', feature g1 and random noise z, and outputs P3, and finally passes the corresponding activation function to obtain the generated sample data; The sample data of the input feature extraction module and the sample data generated by the synthetic feature generation module are paired and input into the discriminator for identification, and the parameter adjustment of the sample expansion model is achieved by setting the loss function. The loss function includes the loss g_loss of the generator G and the loss d_loss of the discriminator D. d_loss includes real_loss that makes the output of the discriminator for the probability of true samples close to the true value and fake_loss that makes the output of the discriminator for the probability of false samples close to the false value. This is content that is easy for technical personnel in this field to understand, and technical personnel in this field can set it by themselves according to actual needs.
[0029] Due to the particularity of the adversarial generative network, the number of random positive sample data pairs that can be selected in the expanded sample pool increases, so that the processed sample data pairs include the same number of mixed positive sample data pairs and mixed negative sample data pairs.
[0030] S3, normalize the sample data pairs processed by S2 using the constructed rules; In order to ensure that the model obtains standardized input, this method further normalizes the sample data during the specific implementation process, including performing unified sample size (Resize), center cropping (CenterCrop), unified sample format (ToTensor), and normalization (Normalize) on the sample data in sequence; Specifically, in the Resize step, all images are resized to a short side length of 255 pixels, while the aspect ratio of the original image remains unchanged; in CenterCrop, the resized image is cropped to 224 × 224 pixels; in ToTensor, the image is converted into a tensor to facilitate passing it into the model for training; Normalize normalizes the pixel values of the image, that is, the original image's 0~255 pixel value size will be mapped to the [0,1] interval.
[0031] S4, training the input matching model with the sample data processed by S3; like Figure 3 As shown, the matching model includes VGG16, a similarity calculation module and a binary classification module connected in sequence.
[0032] S4.1. When selecting the specific backbone network of VGGNet, the present invention uses three VGG neural networks of different depths to verify the experimental results for comparison, including VGG11 (8 convolutional layers + 3 fully connected layers), VGG13 (10 convolutional layers + 3 fully connected layers) and VGG16 (13 convolutional layers + 3 fully connected layers), where the convolution kernel size is fixed to 3 × 3, the convolution step size is fixed to 1, and the padding operation is fixed to 1. The image passes through the convolutional layer and the fully connected layer in the VGG model, and finally outputs a 2622-dimensional feature vector, which is slightly different from the standard VGG model. The main reason is that the pre-trained VGG model originally performed a classification task of a total of 2.6 million images with 2622 categories, and the output is a 2622-dimensional probability, so it can be regarded as having extracted the features of the face image; S4.2. After the input sample data passes through the backbone network of VGGNet, two 2622-dimensional image features are obtained, which are recorded as feature A and feature B respectively. The similarity calculation module uses the cosine similarity method to calculate the difference between feature A and feature B as the input of the classification model; in order to explore the impact of different similarity calculation methods on the effect of identity authentication tasks, the two measurement methods of Euclidean distance and Manhattan distance are also used to calculate the difference of images; S4.3. Prediction and classification are performed using a binary classification module. In the specific implementation process, the present invention uses a 5-layer linear fully connected network with an input dimension of 2622. The input vector and output vector dimensions of each layer are [2622, 1024], [1024, 512], [512, 256], [256, 128], and [128, 2], respectively. The former is the input dimension, and the latter is the mapping dimension. The output obtained by each fully connected layer is then input into the Relu activation function. The final two-dimensional vector represents the probability that the input image is the target image. The maximum value of one dimension of the two-dimensional vector is taken as the prediction result. If the maximum value is the first dimension, it means that the prediction result is 0. If the maximum value is the second dimension, it means that the prediction result is 1. S4.4. Finally, the prediction result of the model is obtained based on the two-dimensional vector output by the binary classification model. A result of 0 indicates that the input image pair is an illegal image, and a result of 1 indicates that the input image pair is a legal image. The prediction result and the actual label are used to perform back-propagation training on the model using the cross entropy loss function, and the optimization function is used to obtain the optimal model parameters.
[0033] The batch_size of all models in the experiment was set to 64, the epoch was set to 150, and the learning rate was set to 0.05. After the actual experiment, for the specific selection of the VGGNet backbone network, two loss functions, absolute value loss and square loss, were used to compare with the cross entropy function. The evaluation indicators used accuracy, precision, recall, and F1 score. The comparison results are shown in Table 1. Table 1. Comparison of loss results of different VGGNet backbone networks
[0034] As can be seen from Table 1, the performance of cross entropy loss is better than absolute value loss and square loss on all three models. When using square loss and absolute value loss, as the number of network layers of the VGG model increases, the model effect is limited, only increasing by about 5%. Finally, this method uses VGG16 as the backbone network.
[0035] S5. Use the trained matching model for face matching.
[0036] The present invention also relates to a high-precision face matching system based on deep learning drive, such as Figure 4 As shown, the system comprises: One or more image acquisition modules, such as cameras, video cameras, etc., are used to acquire the images to be matched. In the actual image acquisition process, the interference of the environment should be minimized; The preprocessing module is used to normalize the matching image according to the constructed rules, including but not limited to normalizing the face obtained from the collected image, removing edge interference information, obtaining the ROI area, etc.; A face matching module, which uses the high-precision face matching method based on deep learning to extract features and perform matching and recognition on faces; In the specific implementation process, a computer-readable storage medium is involved, on which a high-precision face matching program driven by deep learning is stored, and when the program is executed by a processor, the high-precision face matching method driven by deep learning is implemented; In the specific implementation process, a computer device is also involved, including a memory, a processor, and a computer program stored in the memory and run on the processor. When the processor executes the program, the above-mentioned high-precision face matching method driven by deep learning is implemented.
[0037] The feedback module corresponds to the image acquisition module setting and is used to feedback the matching results of the face matching module. It is generally a playback device or integrated in the acquisition device.
[0038] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0039] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0040] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0041] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process in the computer or other programmable device. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0042] Although the preferred embodiments of the present invention have been described, those skilled in the art may make other changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present invention.
[0043] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalents, the present invention is also intended to include these modifications and variations.
Claims
1. A high-precision face matching method based on deep learning, characterized in that: The method comprises the following steps: S1, collect sample data and preprocess them to obtain sample data pairs; S2. Establish a sample expansion model to expand the sample data pairs until the preset conditions are met; S3, normalize the sample data pairs processed by S2 using the constructed rules; S4, training the input matching model with the sample data processed by S3; S5. Use the trained matching model for face matching.
2. According to claim 1, a high-precision face matching method based on deep learning drive is characterized in that: In S1, sample data is collected and invalid data is cleaned; a unique identification code is assigned to the cleaned sample data based on user information; Create mixed positive sample data pairs and mixed negative sample data pairs.
3. The high-precision face matching method based on deep learning drive according to claim 2, characterized in that: Randomly match the cleaned sample data, set the matching thresholds α and β, α>β>0; The sample data pairs whose matching similarity results are greater than the threshold α and whose unique identification codes are different are regarded as fixed negative sample data pairs; The sample data pairs whose matching similarity results are less than the threshold β and whose unique identification codes are the same are fixed positive sample data pairs; Randomly select a number of sample data pairs whose matching similarity results are less than a threshold α and greater than a threshold β, which are random negative sample data pairs and random positive sample data pairs; A fixed positive sample data pair and a random positive sample data pair are considered as a mixed positive sample data pair; a fixed negative sample data pair and a random negative sample data pair are considered as a mixed negative sample data pair.
4. The high-precision face matching method based on deep learning drive according to claim 3, characterized in that: The mixed negative sample data pairs also include filtered negative sample data pairs established based on associated users after extracting user information.
5. The high-precision face matching method based on deep learning drive according to claim 3, characterized in that: The sample expansion model includes a generator and a discriminator connected in sequence, and the generator includes a feature extraction module and a synthetic feature generation module connected in sequence.
6. The high-precision face matching method based on deep learning drive according to claim 5, characterized in that: The feature extraction module includes a plurality of convolution blocks and a plurality of fully connected layers connected in sequence, and any of the convolution blocks outputs corresponding features and results; The synthetic feature generation module includes a number of synthetic blocks connected in sequence, any synthetic block inputs the result of the previous level, the corresponding convolution block outputs the corresponding features and random noise, and up-sampling is performed.
7. The high-precision face matching method based on deep learning drive according to claim 1, characterized in that: The sample data pairs processed by S2 include the same number of mixed positive sample data pairs and mixed negative sample data pairs.
8. The high-precision face matching method based on deep learning drive according to claim 1, characterized in that: In S3, normalizing the sample data includes sequentially performing unification of sample size, center cropping, unification of sample format, and normalization processing on the sample data.
9. The high-precision face matching method based on deep learning drive according to claim 1, characterized in that: In S4, the matching model includes VGG16, a similarity calculation module and a binary classification module connected sequentially.
10. A high-precision face matching system based on deep learning, characterized in that: The system comprises: One or more image acquisition modules, used for acquiring images to be matched; A preprocessing module, used to normalize the matching images according to the constructed rules; A face matching module, using the high-precision face matching method based on deep learning drive as described in any one of claims 1 to 9, for feature extraction, matching and recognition of faces; The feedback module corresponds to the image acquisition module setting and is used to feed back the matching results of the face matching module.
Citation Information
Patent Citations
Power use business handling assisting system based on human face recognition
CN108319941A
Audio similarity matching method and device and storage medium
CN111143604A
Blood relationship automatic recognition method and device based on face image analysis
CN113496219A
Recognition training method and device for face wearing mask, electronic equipment and storage medium
CN114596618A
Breast ultrasound image tumor segmentation method based on cGAN
CN116993649A