Three-segment face recognition method based on dual-spectrum fusion
By adopting a three-stage method of dual spectral fusion in face recognition technology, combining fractional differential convolutional neural network and weighted compact local graph structure algorithm, the problem that the existing technology cannot fight against multiple deception methods at the same time in infrared thermal imaging environment is solved, and efficient and accurate face recognition is achieved.
Patent Information
- Application Number
- CN202210878274.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-25
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2042-07-25
AI Technical Summary
Existing facial recognition technology is difficult to effectively fight at the same time when facing multiple deception methods, especially in infrared thermal imaging environments, monochromatic spectral recognition technology cannot resist advanced deception methods such as 3D models or mask heads at the same time.
The three-stage face recognition method based on dual spectral fusion is adopted. By acquiring visible light and thermal infrared images, a fractional differential convolutional neural network and a weighted compact local graph structure algorithm are used for preliminary detection, and then the image fusion is fusion-based and the distance difference function is used for identification.
It realizes effective recognition of multiple deception methods in infrared thermal imaging environment, improves the accuracy and operation speed of face recognition, and can fight against attacks from visible light, thermal infrared and multiple deception methods at the same time.
Smart Images

Figure CN115188054B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of face recognition technology, and more specifically to the field of face recognition technology based on the fusion of visible light and thermal infrared images. Background Art
[0002] With the development of artificial intelligence, face recognition has been widely used in identity confirmation. Face recognition technology is a biometric technology that recognizes people based on their facial features. However, even though face recognition technology is very advanced today, fake and disguised face technology still exists and has become a major obstacle to achieving higher breakthroughs in face recognition technology. At present, face recognition technology mainly faces three types of fraud methods. For example, Figure 1 As shown in the figure, the first method is to use the user's face picture (a), which is low-cost and easy to implement; the second is to use the user's face video (b), among which the video containing facial expressions and facial movements is the most deceptive; the last is to use the user's 3D model or mask headgear ((c)(d)(e)), which uses 3D technology to synthesize the face and can imitate facial movements, which is more deceptive than photos and videos. Faced with the proliferation of counterfeiting methods, liveness detection came into being. Liveness detection is a key technology to determine the user's true physiological characteristics and can effectively resist the attacks of the above-mentioned deceptive methods. At present, the following liveness detection technologies are widely used. The first is a method based on micro-texture, which collects fake faces multiple times to show the difference from real faces. This method can compare certain differences, but is easily disturbed by light, especially for video attacks. The second is a method based on motion information, which extracts specific motion information of the face to make judgments. This method is the most widely used in the existing technology, but it has high requirements for user interaction, and the effect is not ideal when the attacker hollows out the photo to make corresponding actions. The third is a method based on dual spectrum, which determines the true and false faces through the differences in spectral reflectivity of various materials, and uses multiple bands to combine to make the true and false faces show a large difference, so as to distinguish them. This method is very effective in liveness detection, but the collection conditions are relatively strict; the fourth is a multiple feature fusion algorithm that combines the above methods.
[0003] With the continuous improvement of recognition technology, a large number of existing face recognition technologies are based on visible light environments. However, there is a lack of research on face recognition technology that integrates dual-spectrum visible light and infrared thermal imaging. In recent years, representative infrared thermal imaging face recognition technologies mainly include: Yangyang Lian et al. proposed in 2020 to use component analysis PCA algorithm to judge the facial occlusion state, introduce infrared thermal imaging technology, and use BP network to locate facial acupoints to recognize faces. Vicente Pavez et al. proposed in 2022 to create a facial thermal imaging database, use StyleCLIP supervised operation to input the latent space of visible images, and add some required attributes to the visible face. Then use FaceNet architecture to create robust thermal imaging face recognition. The above methods collect monochromatic spectra and use neural networks to identify images. With the advancement of fraud methods, monochromatic spectrum recognition technology can only solve some problems. In the infrared thermal imaging environment, there will be temperature differences when using the user's 3D model or mask headgear, so as to achieve the purpose of preventing deception. When different monochromatic spectra are used to complete face recognition, it is impossible to fight against multiple deception methods at the same time. Summary of the invention
[0004] The purpose of the present invention is to solve the technical problem that when different monochromatic spectra are used to complete face recognition, it is impossible to simultaneously fight against multiple deception methods. The present invention provides a three-stage face recognition method based on dual-spectrum fusion.
[0005] In order to achieve the above-mentioned purpose, the present invention specifically adopts the following technical solutions:
[0006] The three-stage face recognition method based on dual-spectrum fusion includes the following steps:
[0007] The steps include:
[0008] Step 1, obtain the visible light image and thermal infrared image of the target to be identified (use two sensors to obtain visible light image and thermal infrared imaging respectively. Since the thermal imager outputs a single-channel signal in AV format, the signal needs to be converted into a single-channel digital image format through a data acquisition board. At the same time, the visible light image also needs to be converted into a single-channel format so that the image size is consistent during subsequent image fusion);
[0009] Step 2, first stage detection: input the visible light image into the convolutional neural network face recognition algorithm based on fractional differential to obtain the first recognition result; input the thermal infrared image into the face recognition algorithm with weighted compact local graph structure to obtain the second recognition result;
[0010] Step 3, second stage detection: If both the first recognition result and the second recognition result are true, the visible light image and the thermal infrared image after the first stage detection are fused and registered to obtain a fused image, and the distance difference L between the visible light image after the first stage detection and the fused image is calculated. A , the distance difference L between the thermal infrared image and the fused image after the first stage of detection B , and sum L A +L B That is, the total distance difference Ds is obtained, and the nonlinear function is introduced to obtain the distance difference function. The value range of the distance difference function is (0,1). When the total distance difference Ds is less than the threshold q=0.4, the recognition result is true (the distance difference function is used to detect the difference between the images before and after fusion, and the correctly fused image is obtained as the recognition data; the distance difference function solves the problem that the mask cannot reflect the accurate temperature difference during thermal imaging, and the imaging is different from the naked face, resulting in the inability to display some feature points, and thereby prevents mask attacks);
[0011] Step 4, third stage detection: When the recognition result of the second stage detection is true, the fused image is input as recognition data into the trained (already trained and verified) dual-spectrum fusion face recognition network model to obtain the third recognition result.
[0012] In the technical solution of the present application: currently available data sets are all based on monochromatic spectra, so it is necessary to collect data sets that fuse visible light and infrared thermal imaging. In the first stage, the purpose is to solve the problem of single visible light image and thermal infrared image recognition. For visible light face recognition, the present invention uses a convolutional neural network face recognition algorithm based on fractional differentials. This method extracts more facial features by adding an attention mechanism and using fractional differentials to process node functions. While enhancing the robustness of the network, the ArcFace loss function is finally used to optimize the model, and iterative training is performed in the network to complete face recognition; a face recognition algorithm based on a weighted compact local graph structure is selected. This algorithm enriches the feature information of the image by extracting closer pixel information, using a reasonable weighting strategy, and paying close attention to the difference between the central pixel and the neighboring pixels. The algorithm has shown excellent performance on the HD thermal infrared face data set; the second stage aims to improve the security of the system against mask deception. Considering that the attacker wearing a mask can cause the system to fail to fuse correctly, because the mask cannot reflect the accurate temperature difference under thermal imaging, and the imaging is different from that of a naked face, which causes some feature points to fail to display. Therefore, the present invention designs a distance difference function to detect the difference between the images before and after fusion, and uses it to prevent mask attacks. SThe smaller the value, the better the image fusion effect. In particular, when the value is greater than the threshold q=0.4, the fusion effect becomes blurred or even garbled. Therefore, when an attacker uses a mask to attack, this value can be used to eliminate it; the third stage solves the recognition problem of dual-spectrum face image fusion. The image features are seriously lost during the dual-spectrum fusion process. The present invention designs a network that can fully extract the features of the fused image. This network MF-FRNet nests the innovative MCC-ANet network and MF-FRMCNet network. The dual-spectrum fusion face recognition network performs some simple channel and size transformation processing on the input image, and then passes through a multi-convolutional layer cascade mean network to further process the fused image, and performs a series of feature detail enhancements to obtain a multi-dimensional feature vector. Finally, the multi-dimensional vector is converted into the corresponding eigenvalue through the MF-FRMCNet network for processing to obtain the final recognition result, including the position of the target and the target center. The recognition method of the present application shows excellent anti-spoofing performance. Compared with the previous face recognition, the present invention aims to prevent attacks by currently popular deception means. It utilizes the complementary advantages of thermal infrared images and visible light images to realize three-stage liveness detection with dual-spectrum image fusion as the main line, which fully reflects the idea of dual-spectrum fusion for liveness detection. At the same time, it can complete the recognition of different attack means at different stages, greatly improving the recognition accuracy and running speed, and can fight against multiple deception means at the same time.
[0013] Furthermore, in step 3, the visible light image and the thermal infrared image are registered using the surf algorithm, and then the visible light image and the thermal infrared image are fused using a method based on PCNN and IFS (this method can retain image detail information to a greater extent and make the fused image clearer).
[0014] Furthermore, the dual-spectrum fusion face recognition network includes 1 convolution layer, 1 maximum pooling layer, 3 multi-convolution layer cascade mean networks, 2 filter cascades, 1 average pooling layer, 1 convolution layer and 1 classification network, and the feature map is generated and then input into the face recognition classification network.
[0015] Furthermore, the multi-convolutional layer cascade mean network includes 2 maximum pooling layers, 7 convolutional layers and 2 filter cascades.
[0016] Furthermore, the face recognition classification network includes 1 CBAM layer, 1 convolutional layer and 3 fully connected layers. The recognition result is output from the 3 fully connected layers and after softmax (for auxiliary classification).
[0017] Furthermore, in step 3, the steps of designing the distance difference function are as follows:
[0018] Step I: define the distance between the pixels of the visible light image A and the fused image C as d(A, C); the distance between the pixels of the thermal infrared image B and the fused image C as d(B, C).
[0019]
[0020]
[0021] Where, N = m × n, A i , C i , B i are the pixel values of the visible light image A, thermal infrared image B and fused image C;
[0022] (The fused image is composed of two images, so when designing the function, the relationship between the three should be considered. Since the facial features in the image need to be reflected) Use a sliding window to divide the image into blocks, and then calculate the distance difference for each sub-image. The calculation form is:
[0023] D 0 (A,B,C)=s A d(A,C)+s B d(B,C) (3)
[0024] In formula (3), s A 、s B is the proportion of the standard deviation of the image in the sliding window, which further reflects the similarity between the source image and the fused image. The calculation form is:
[0025]
[0026]
[0027] Among them, std(A,w),std(B,w) are the standard deviations in the image sliding window, and their calculation form is:
[0028]
[0029]
[0030] where μ A , μ B is the grayscale mean, and then calculate the total distance difference, that is:
[0031]
[0032]
[0033]
[0034] Step II: (In order to better judge the fusion effect and prevent mask deception) introduce a nonlinear function and substitute equation (10) into it, which is expressed as:
[0035]
[0036] Formula (11) is the final distance difference function and its value range is (0,1) (a large number of experiments show that it is concluded that: D S The smaller the value, the better the image fusion effect. In particular, when the value is greater than the threshold q=0.4, the fusion effect becomes blurred or even garbled. Therefore, when an attacker uses a mask to attack, this value can be used to exclude him.)
[0037] Furthermore, the training method of the dual-spectrum fusion face recognition network model includes:
[0038] Step A, obtaining training data (fused correctly fused image), wherein the training data includes corresponding true labels;
[0039] Step B, inputting the training data into an untrained dual-spectrum fusion face recognition network model to obtain an output result;
[0040] Step C: determining a loss function based on the output result and the true label;
[0041] Step D: iteratively train the dual-spectrum fusion face recognition network based on the loss function to obtain a trained dual-spectrum fusion face recognition network.
[0042] Furthermore, (to further improve the accuracy) the loss function uses category prediction and center point prediction for association measurement, including face category loss and center point positioning loss, defined as:
[0043]
[0044] Among them, n is the size of the data set loaded at one time during training, that is, batchsize, is the predicted probability value of each category, θ i is the true value of the center point and face category (the three fully connected layers of the face recognition classification network output an N+2-dimensional low-dimensional vector, 2 is the predicted center point coordinate After N dimensions pass through softmax, we get the probability of each category The category with the highest probability is the final recognition result, θ i is the true value of the center point and face category);
[0045] Center point positioning loss L location It also uses the mean square error form, which is defined as follows:
[0046]
[0047] In the formula is the predicted value of the center point coordinate, x i ,y i True value;
[0048] The global loss function includes the above-mentioned face classification loss and center point positioning loss, so the function is defined as the sum of face classification loss and center point positioning loss. Substituting into equations (12) and (13) we can get Loss, which is calculated as follows:
[0049] Loss = log(L class +L location ) (14)
[0050]
[0051] (The nonlinear log function is introduced in the formula, which can effectively solve the problem of gradient disappearance and enable the network to converge quickly).
[0052] The beneficial effects of the present invention are as follows:
[0053] 1. The present invention aims to prevent attacks by currently popular deceptive means. It uses the complementary advantages of thermal infrared images and visible light images to achieve three-stage liveness detection with dual-spectrum image fusion as the main line, which fully embodies the idea of dual-spectrum fusion for liveness detection. At the same time, it can complete the identification of different attack means at different stages, greatly improving the accuracy of identification and the speed of operation, and can simultaneously fight against multiple deceptive means;
[0054] 2. The present invention collects a data set of visible light and infrared thermal imaging fusion, designs a distance difference function to register the fused image, constructs a facial temperature difference extraction network for the fused image, obtains facial features, and constructs a discrimination network for attacks by different deception methods, thereby determining the stage of recognition failure and being able to fight against multiple deception methods at the same time;
[0055] 3. According to the complementary characteristics of thermal infrared imaging and visible light images, the present invention designs a dual-spectrum fusion three-stage face recognition network that can fully extract image features and accurately identify based on registration fusion and face recognition network. The three-stage method can complete the identification of different attack methods at different stages, greatly improving the accuracy of identification and the speed of operation. The distance difference function and neural network model designed by the invention, while ensuring accurate image fusion, the neural network also has the functions of strong image feature extraction ability, strong fitting data samples, and weight sharing to reduce model size, which can solve more advanced fraud methods. Attack methods that cannot be identified by general face recognition networks are completed. For the features of the images at different stages, each stage of the three-stage method performs corresponding method identification based on different features, minimizing the loss of image features in each stage as much as possible, thereby maximizing the recognition efficiency;
[0056] 4. In view of the differences in image data from different sensors, and the complementarity of visible light and thermal infrared, the dual-light face image registration and fusion algorithm of the present application effectively combines visible light and thermal infrared images, fully extracts image texture information, and integrates them to generate high-quality images. Image registration plays a vital role as the preprocessing part of later image fusion. Taking into account the high real-time requirements of the system, after comparing the advantages and disadvantages of several traditional registration methods, the present invention finally uses the upgraded version of the SIFT algorithm, the SURF algorithm; traditional fusion methods are prone to shortcomings such as blurred edges and unclear features. The present invention uses a thermal infrared and visible light image fusion method based on multi-scale transformation and norm optimization. This method can well retain image detail information and has very high image quality;
[0057] 5. The distance difference function of the present application is intended to improve the security of the system against mask deception. Considering that the attacker wearing a mask causes the system to fail to fuse correctly, because the mask cannot reflect the accurate temperature difference under thermal imaging, and the imaging is different from the naked face, resulting in some feature points not being displayed. Therefore, the present invention designs a distance difference function to detect the difference between the images before and after fusion, and thereby prevent mask attacks;
[0058] 6. The present invention adopts a three-stage method based on dual-spectrum fusion for face recognition, which is suitable for dealing with advanced deception methods such as 3D models or masks and headgear, and has great application value in technical fields such as payment security and security security;
[0059] 7. Since there are relatively few studies on face recognition using dual-spectrum fusion images and a lack of a test model that matches the fusion image, the present invention proposes a distance difference function that is suitable for fusion images and determines whether the fusion is accurate, which reduces the feature loss caused by the fusion process and is conducive to improving the efficiency of subsequent network feature extraction;
[0060] 8. The present invention proposes a dual-spectrum fusion face recognition network MF-FRNet suitable for visible light and thermal infrared fusion images. The network has the following advantages: a. The MCC-ANet network is embedded in the MF-FRNet network to further strengthen feature extraction and reduce the features lost in the fusion process; and multi-level cascade output is adopted to obtain different feature maps at different stages, and different weights are assigned to optimize the calculation of the subsequent recognition network; b. The MCC-ANet network combines the characteristics of the dual-spectrum fusion image, and processes the feature maps of different dimensions differently to obtain more feature details; c. The attention mechanism is added to the face recognition network MF-FRMCNet to highlight the more important edge information in the feature map;
[0061] 9. The present invention proposes an innovative loss function for the training process, which includes both category loss and center point positioning loss. The loss function is defined by a joint measurement method and a nonlinear function is added, which can effectively solve the problem of gradient disappearance and can converge quickly. BRIEF DESCRIPTION OF THE DRAWINGS
[0062] Figure 1 It is a schematic diagram of existing attack methods;
[0063] Figure 2 It is the overall framework diagram of the present invention;
[0064] Figure 3 is a flow chart of the distance difference function of the present invention;
[0065] Figure 4 It is a comparative test data diagram of the present invention;
[0066] Figure 5 It is the dual-spectrum fusion face recognition network structure of the present invention;
[0067] Figure 6 It is a multi-convolutional layer cascade mean network structure of the present invention;
[0068] Figure 7 It is the face recognition classification network of the present invention;
[0069] Figure 8 It is the overall process framework of the present invention;
[0070] Fig. 9 is a drawing of the photo verification result of the legal user in Table 1 of the present invention;
[0071] Fig.10 is a diagram of the video verification result of the legal user in Table 1 of the present invention;
[0072] Fig.11 is a diagram of the mask verification result of the legal user in Table 1 of the present invention;
[0073] Fig.12 is a drawing of the 3D headgear verification result of the legal user in Table 1 of the present invention;
[0074] Fig.13 It is a drawing of the verification result of the 3D simulation model of the legitimate user in Table 1 of the present invention. DETAILED DESCRIPTION
[0075] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in combination with the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments.
[0076] Therefore, based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making any creative work shall fall within the scope of protection of the present invention.
[0077] Example 1
[0078] like Figure 2-8 As shown, the three-stage face recognition method based on dual-spectrum fusion includes the following steps:
[0079] Step 1, obtain the visible light image and thermal infrared image of the target to be identified (acquisition, since the thermal imager outputs a single-channel signal in AV format, the signal needs to be converted into a single-channel digital image format through a data acquisition board, and the visible light image also needs to be converted into a single-channel format so that the image size is consistent during subsequent image fusion);
[0080] Step 2, first stage detection: input the visible light image into the convolutional neural network face recognition algorithm based on fractional differential to obtain the first recognition result; input the thermal infrared image into the face recognition algorithm with weighted compact local graph structure to obtain the second recognition result;
[0081] Step 3, second stage detection: If both the first recognition result and the second recognition result are true, the visible light image and the thermal infrared image after the first stage detection are fused and registered to obtain a fused image, and the distance difference L between the visible light image after the first stage detection and the fused image is calculated. A , the distance difference L between the thermal infrared image and the fused image after the first stage of detection B , and sum L A +L BThat is, the total distance difference Ds is obtained, and the nonlinear function is introduced to obtain the distance difference function. The value range of the distance difference function is (0,1). When the total distance difference Ds is less than the threshold value q=0.4, the recognition result is true (the distance difference function is used to detect the difference between the images before and after fusion, and the fused image with correct fusion is obtained as the recognition data; the distance difference function solves the problem that the mask cannot reflect the accurate temperature difference during thermal imaging, and the imaging is different from the naked face, which leads to the problem that some feature points cannot be displayed, and this is used to prevent mask attacks); Among them, the visible light image and the thermal infrared image are registered by the surf algorithm, and then the method based on PCNN and IFS is used to fuse the visible light image and the thermal infrared image (this method can retain the image detail information to a greater extent and make the fused image clearer);
[0082] The steps for designing the distance difference function are as follows:
[0083] Step I: define the distance between the pixels of the visible light image A and the fused image C as d(A, C); the distance between the pixels of the thermal infrared image B and the fused image C as d(B, C).
[0084]
[0085]
[0086] Where, N = m × n, A i , C i , B i are the pixel values of visible light image A, thermal infrared image B and fused image C;
[0087] (The fused image is composed of two images, so when designing the function, the relationship between the three should be considered. Since the facial features in the image need to be reflected) Use a sliding window to divide the image into blocks, and then calculate the distance difference for each sub-image. The calculation form is:
[0088] D 0 (A,B,C)=s A d(A,C)+s B d(B,C) (3)
[0089] In formula (3), s A 、s B is the proportion of the standard deviation of the image in the sliding window, which further reflects the similarity between the source image and the fused image. The calculation form is:
[0090]
[0091]
[0092] Among them, std(A,w),std(B,w) are the standard deviations in the image sliding window, and their calculation form is:
[0093]
[0094]
[0095] where μ A , μ B is the grayscale mean, and then calculate the total distance difference, that is:
[0096]
[0097]
[0098]
[0099] Step II: (In order to better judge the fusion effect and prevent mask deception) introduce a nonlinear function and substitute equation (10) into it, which is expressed as:
[0100]
[0101] Formula (11) is the final distance difference function and its value range is (0,1) (a large number of experiments show that it is concluded that: D S The smaller the value, the better the image fusion effect. In particular, when the value is greater than the threshold q=0.4, the fusion effect becomes blurred or even garbled. Therefore, when an attacker uses a mask to attack, this value can be used to exclude him);
[0102] Step 4, third stage detection: when the recognition result of the second stage detection is true, the fused image is input as recognition data into the trained (trained and verified) dual-spectrum fusion face recognition network model to obtain the third recognition result; the dual-spectrum fusion face recognition network includes 1 convolution layer, 1 maximum pooling layer, 3 multi-convolution layer cascade mean networks, 2 filter cascades, 1 average pooling layer, 1 convolution layer and 1 classification network, and the feature map is generated and input into the face recognition classification network; the multi-convolution layer cascade mean network includes 2 maximum pooling layers, 7 convolution layers and 2 filter cascades; the face recognition classification network includes 1 CBAM layer, 1 convolution layer and 3 fully connected layers, and the recognition result is output from the 3 fully connected layers and after softmax (for auxiliary classification);
[0103] The training method of the dual-spectrum fusion face recognition network model includes:
[0104] Step A, obtaining training data (fused correctly fused image), wherein the training data includes corresponding true labels;
[0105] Step B, inputting the training data into an untrained dual-spectrum fusion face recognition network model to obtain an output result;
[0106] Step C: determining a loss function based on the output result and the true label;
[0107] Step D, iteratively training the dual-spectrum fusion face recognition network based on the loss function to obtain a trained dual-spectrum fusion face recognition network;
[0108] (To further improve the accuracy) The loss function uses category prediction and center point prediction for association measurement, including face category loss and center point positioning loss, defined as:
[0109]
[0110] Among them, n is the size of the data set loaded at one time during training, that is, batchsize, is the predicted probability value of each category, θ i is the true value of the center point and face category (the three fully connected layers of the face recognition classification network output an N+2-dimensional low-dimensional vector, 2 is the predicted center point coordinate The N-dimensional vector is passed through softmax to obtain the probability of each category The category with the highest probability is the final recognition result, θ i is the true value of the center point and face category);
[0111] Center point positioning loss L location It also uses the mean square error form, which is defined as follows:
[0112]
[0113] In the formula is the predicted value of the center point coordinate, x i ,y i True value;
[0114] The global loss function includes the above-mentioned face classification loss and center point positioning loss, so the function is defined as the sum of face classification loss and center point positioning loss. Substituting into equations (12) and (13) we can get Loss, which is calculated as follows:
[0115] Loss = log(L class +L location ) (14)
[0116]
[0117] (The nonlinear log function is introduced in the formula, which can effectively solve the problem of gradient disappearance and enable the network to converge quickly).
[0118] Example 2
[0119] Based on Example 1, Figure 4 As shown in the figure, 50 fake samples with masks and 50 fake samples with hoods and 50 real samples without masks and 50 real samples without hoods were selected for testing, with a total of 200 samples. The following scatter plot was drawn according to their Ds values, as shown in the figure. Figure 4 As shown, the green dots represent those that failed, and the red dots represent those that succeeded. Because of the mask, there is a temperature difference, which will lead to poor fusion effect, and even blur or even garbled code. From the table below, we can see that the Ds value of the sample with a mask is larger than the Ds value of the sample without a mask, which also verifies the idea that wearing a mask will affect the fusion effect. It can also be seen that there is a clear dividing line Ds=0.4 between the red and green points, so 0.4 can be used as the detection threshold, that is, only when the required distance difference function value is less than 0.4, can the verification be successful. Therefore, when an attacker uses a mask to attack, this value can be used to exclude it.
[0120] Example 3
[0121] Based on Example 1, Figure 5-7 As shown in Figure 1, a dual-spectral fusion face recognition network (Multispectral Fusion Face Recognition Network, MF-FRNet) that can highlight the characteristics of fused images is shown in Figure 1. Figure 5As shown in the figure, this network aims to enhance the feature details lost in the fusion process. It includes 1 convolution layer, 1 maximum pooling layer, 3 multi-convolution layer cascade mean networks, 2 filter cascades, 1 average pooling layer, 1 convolution layer and 1 classification network. The input image size of this network is 416×416×1. After adjusting the channel through a 5×5 convolution layer, a 416×416×32 feature map is generated; then, after a 7×7 maximum pooling layer with a step size of 2, features are further extracted to generate a 208×208×64 feature map; then, after a multi-convolution layer cascade mean network, a feature map of size 104×104×128 is output at each layer; the present invention adopts multi-level cascade output to extract features to the greatest extent, so the last convolution mean layer passes through a filter cascade and is consistent with the previous one. The three 104×104×128 images are cascaded together through a filter. This output process makes up for the loss of features in the fusion process to a large extent. According to the importance of different channel features, each channel is adaptively processed to increase the extraction of important features. Then it is downsampled through a 5×5 average pooling layer, and finally the number of channels is adjusted through a 1×1 convolution to generate a 104×104×256 feature map as the input of the face recognition classification network MF-FRMCNet. After processing, the recognition result is finally obtained.
[0122] The Multi Convolution Cascade Average Network (MCC-ANet) combines the features of the dual-spectral fusion image and processes the feature maps of different dimensions differently to obtain more feature details. Figure 6As shown in the figure, it includes 2 maximum pooling layers, 7 convolutional layers and two filter cascades. First, the maximum pooling layer in MF-FRNet outputs a 208×208×64 feature map as the input of the first convolutional mean layer. Then it passes through two convolutional layers and two maximum pooling layers respectively. It includes 1 5×5 maximum pooling layer with a step size of 2, 1 3×3 convolutional layer with a step size of 2, 1 3×3 maximum pooling layer with a step size of 2, and 1 5×5 convolutional layer with a step size of 2. Finally, four feature maps with different feature details but the size of 104×104×64 are generated. The results after convolution are divided into two steps for different processing. In the first step, the results after different convolutional layers are passed through the filter cascade together, and then adjusted through a 1×1 convolutional layer to adjust the channel, and then a 104×104×128 feature map is output. In the second step, the results after different convolutional layers correspond to a 3×3 or 5×5 convolutional layer respectively. The purpose here is to adjust the channels and generate four 104×104×128 feature maps. Then, after filtering, upsampling is performed to generate 208×208×64 feature maps, which are used as the input of the next convolutional mean layer. The purpose of the entire MCC-ANet network is to reduce the interference caused by the fusion process after training. It greatly enhances the extraction of feature details.
[0123] The recognition network (MF-FRMCNet) of the present invention is as follows Figure 7 As shown in the figure, the Softmax function is used as the classification basis. This network consists of 1 CBAM layer, 1 convolution layer, and 3 fully connected layers. First, the attention mechanism CBAM is used to highlight the more important edge information in the feature map. This mechanism can deal with more and more advanced fraud methods. Then, the dimension of the feature map is changed through a 1×1 convolution layer, and then the obtained feature dimension is subjected to logistic regression through 3 fully connected layers to obtain N+2 dimensions (2 is the predicted center point coordinates). ), and then the remaining N dimensions are processed by softmax to obtain the probability of each category. The category with the highest probability is the final recognition result.
[0124] The face recognition technology based on the fusion of visible light and thermal infrared uses the fused image as input and outputs the probability of each category from 0 to 1.
[0125] Example 4
[0126] Based on Example 1, the training process includes: in the training process of the dual-spectrum fusion face recognition network (MF-FRNet), the fused image in the data set is converted into an image of size 416×416, the epoch is set to 100 times during training, the batchsize is set to 32, and the network weight is updated after one training. The loss function is shown in Formula 15, the back propagation process uses SGD (stochastic gradient descent), the initial learning rate is set to 0.005, and the learning rate uses a fixed step decay strategy. And the model is saved every 10 iterations, and the model with the lowest loss is finally selected as the optimal model, in order to better calculate the most accurate probability distribution through the weight parameters in the overall recognition process.
[0127] Example 5
[0128] like Figure 2-13 As shown, the three-stage face recognition method based on dual-spectrum fusion includes the following steps:
[0129] [1] Collecting data sets. The present invention uses a binocular depth camera and an infrared thermal imager to collect visible light images and thermal infrared images. 500 visible light images of legal users are collected by the binocular depth camera, and 500 infrared thermal images of legal users are collected by the infrared thermal imager. A total of 1,000 images are collected, including 10 categories of people. Since the thermal imager outputs a single-channel signal in AV format, the signal needs to be converted into a single-channel digital image format through a data acquisition board. At the same time, the visible light image also needs to be converted into a single-channel format so that the image size is consistent during subsequent image fusion. Next, the collected images are fused and registered to facilitate the production of a special fusion image dataset.
[0130] [2] Create the first-stage dataset. Use the 1,000 images in [1] to annotate the real labels and create a visible light face recognition dataset and a thermal infrared face recognition dataset respectively, in preparation for the network training in the first stage.
[0131] [3] The first stage of three-stage detection. The face recognition algorithm with weighted compact local graph structure and the convolutional neural network face recognition algorithm with fractional differential are used to perform 100 rounds of iterative training using the data set in [2] with a learning rate of 0.05. After the training, the optimal model is selected for face recognition. This stage is used to identify some relatively simple deception methods, such as legitimate user photos, legitimate user videos, etc.
[0132] [4] Image registration and fusion. The surf algorithm is used to register the images, and then the PCNN and IFS methods are used to perform image registration and fusion. This method can retain image detail information to a greater extent and make the fused image clearer. It prepares for the distance difference calculation proposed in the second stage of detection and the data set in the third stage of detection.
[0133] [5] The second stage of three-stage detection. If the first stage recognition is passed, the second stage detection is entered. In the registration and fusion stage, the visible light and thermal infrared images that have passed the first stage recognition are fused and registered to obtain a fused image. According to the distance difference function proposed in the present invention, its calculation form is shown in formula (11), that is, the two source images and the fused image are divided into blocks using the sliding window method, and then the distance difference of each sub-block of the source image and the fused image is calculated and summed to obtain the distance difference, and finally a nonlinear function is introduced to make its value range (0,1). The threshold q=0.4 obtained from a large number of comparative experiments is used as the judgment standard. Only when the distance difference Ds is less than the threshold, can it pass the verification. Since the mask cannot reflect the accurate temperature difference during thermal imaging, the imaging is different from the naked face, resulting in the problem that some feature points cannot be displayed, which will lead to a poor fusion effect. Therefore, the use of this function can effectively prevent mask deception.
[0134] [6] Create the third-stage DFFD dataset. Use the 500 visible light images and 500 thermal infrared images in [1], obtain 500 fused images through the registration and fusion algorithm in [4], annotate the true labels of the face fused images, and create a fused image dataset. Use the 500 fused images as the DFFD (Dualoptical fusion Face Database) dataset for the third-stage network training.
[0135] [7] Divide the dataset. The DFFD dataset produced in [6] is divided into training set and test set in the ratio of 8:2.
[0136] [8] Build the network. Use the deep learning framework Pytorch and Python programming language to build the dual-spectrum fusion face recognition network MF-FRNet. In the MF-FRNet network, the innovative MCC-ANet network and MF-FRMCNet network are nested. The detailed overall framework is as follows: Figure 5 As shown in the figure. The face recognition network of dual-spectrum fusion performs some simple channel and size transformation on the input image at the beginning, and then passes through a multi-convolutional layer cascade mean network to further process the fused image and perform a series of feature detail enhancement to obtain a multi-dimensional feature vector. Finally, the multi-dimensional vector is converted into the corresponding feature value through the MF-FRMCNet network for processing to obtain the final recognition result.
[0137] [9] Set network parameters. During the training of the dual-spectrum fusion face recognition network, the fused images in the DFFD training data set are converted to images of size 416×416. The batch size is set to 32 during training, and the loss function is shown in Equation 15. SGD (stochastic gradient descent) is used to update the parameters through back propagation of the loss function. The initial learning rate is set to 0.005, and the learning rate adopts a fixed step decay strategy. The dual-spectrum fusion face recognition network is iterated 100 times, and the model is saved every 10 iterations. Finally, the model with the lowest loss is selected as the optimal model. The purpose is to better calculate the most accurate probability distribution through the weight parameters in the overall recognition process.
[0138]
[10] The third stage of the three-stage detection. Load the optimal model obtained in the above steps for recognition, and select the category corresponding to the maximum probability as the recognition result. If there is a recognition result, the verification passes, otherwise it fails.
[0139]
[11] Experimental test phase. As shown in Table 1, different types of deception methods were put into the network for testing. A total of 200 legitimate user photos, 200 legitimate user videos, 100 real-time tests with legitimate user masks, 100 real-time tests with legitimate user 3D headgear, and 100 tests with legitimate user 3D simulation models (containing thermal infrared information and able to simulate live facial movements). When using legitimate user photos and videos, although visible light detection can pass, the first stage verification fails because there is no thermal infrared information; when using legitimate user masks or headgear, the first stage is not enough to exclude them, so the second stage is used for fusion and distance difference calculation. Because the distance difference is greater than the threshold, it is excluded, and the verification fails; when using the legitimate user's 3D simulation model, because it is very different from the real person, it cannot be excluded by the first and second stages, and can only be recognized at a deeper level by the third stage MF-FRNet network. Because the network does not have any output results, the verification fails.
[0140] Table 1 Analysis of the stages of failure of deception
[0141] Test method Is it a deceptive method? Failure at which stage Passed Verify the results Legal user photo yes Phase 1 no like Fig. 9 Shown Legitimate user video yes Phase 1 no like Fig.10 Shown Legitimate User Mask yes Phase II no like Fig.11 Shown Legal user 3D headgear yes Phase II no like Fig.12 Shown 3D simulation model of legitimate users yes Phase 3 no As Fig.13 shown Legal users no ---- yes
[0142] From the above evaluation results, it can be seen that the system has shown excellent anti-spoofing performance. Compared with previous face recognition, the present invention aims to prevent attacks by currently popular deception methods, and uses the complementary advantages of thermal infrared images and visible light images to achieve three-stage liveness detection with dual-spectrum image fusion as the main line, which fully reflects the idea of dual-spectrum fusion for liveness detection. From the results, it can be seen that the distance difference function and the face recognition network also reflect their role, that is, while ensuring the correct fusion, they also show excellent recognition accuracy.
Claims
1. Three-segment face recognition method based on dual-spectrum fusion, It is characterized in that The steps include: Step 1: Obtain a visible light image and a thermal infrared image of the target to be identified; Step 2, first stage detection: input the visible light image into the convolutional neural network face recognition algorithm based on fractional differential to obtain the first recognition result; input the thermal infrared image into the face recognition algorithm with weighted compact local graph structure to obtain the second recognition result; Step 3, second stage detection: If both the first recognition result and the second recognition result are true, the visible light image and the thermal infrared image after the first stage detection are fused and registered to obtain a fused image, and the distance difference between the visible light image after the first stage detection and the fused image is calculated. , the distance difference between the thermal infrared image and the fused image after the first stage of detection , and sum + That is, the total distance difference Ds is obtained, and a nonlinear function is introduced to obtain a distance difference function, the value range of which is (0,1). When the total distance difference Ds is less than the threshold q=0.4, the recognition result is true; Step 4, third stage detection: when the recognition result of the second stage detection is true, the fused image is input as recognition data into the trained and generated dual-spectrum fusion face recognition network model to obtain a third recognition result; The dual-spectrum fusion face recognition network includes 1 convolution layer, 1 maximum pooling layer, 3 multi-convolution layer cascade mean networks, 2 filter cascades, 1 average pooling layer, 1 convolution layer and 1 classification network. After generating the feature map, it is input into the face recognition classification network; The multi-convolutional layer cascade mean network includes 2 maximum pooling layers, 7 convolutional layers and 2 filter cascades; The face recognition classification network includes 1 CBAM layer, 1 convolutional layer and 3 fully connected layers. The recognition results are output from the 3 fully connected layers and after softmax.
2. According to claim 1, the three-stage face recognition method based on dual-spectrum fusion, It is characterized in that In step 3, the visible light image and the thermal infrared image are registered using the surf algorithm, and then the visible light image and the thermal infrared image are fused using a method based on PCNN and IFS.
3. The three-stage face recognition method based on dual-spectrum fusion according to claim 1, It is characterized in that In step 3, the steps of designing the distance difference function are as follows: Step I: Define the distance between the pixels of the visible light image A and the fused image C as ; The distance between the pixels of thermal infrared image B and fused image C is , (1) (2) in, , , , are the pixel values of the visible light image A, thermal infrared image B and fused image C; Use a sliding window to divide the image into blocks, and then calculate the distance difference for each sub-image. The calculation form is: (3) In formula (3) , is the proportion of the standard deviation of the image in the sliding window, which further reflects the similarity between the source image and the fused image. The calculation form is: (4) (5) in, , is the standard deviation in the image sliding window, which is calculated as: (6) (7) in , is the grayscale mean, and then calculate the total distance difference, that is: (8) (9) (10) Step II: Introduce nonlinear function and substitute equation (10) into it, which is expressed as: (11) Formula (11) is the final distance difference function and its value range is (0,1).
4. The three-stage face recognition method based on dual-spectrum fusion according to claim 1 or 3, It is characterized in that The training method of the dual-spectrum fusion face recognition network model includes: Step A: obtaining training data, wherein the training data includes corresponding real labels; Step B, inputting the training data into an untrained dual-spectrum fusion face recognition network model to obtain an output result; Step C: determining a loss function based on the output result and the true label; Step D: iteratively train the dual-spectrum fusion face recognition network based on the loss function to obtain a trained dual-spectrum fusion face recognition network.
5. The three-stage face recognition method based on dual-spectrum fusion according to claim 4, It is characterized in that The loss function uses category prediction and center point prediction for association measurement, including face category loss and center point positioning loss, defined as: (12) Among them, n is the size of the data set loaded at one time during training, that is, batchsize, is the predicted probability value of each category, is the true value of the center point and face category; Center point positioning loss It also uses the mean square error form, which is defined as follows: (13) In the formula is the predicted value of the center point coordinates, True value; The global loss function includes the above-mentioned face category loss and center point positioning loss, so the function is defined as the sum of face category loss and center point positioning loss. Substituting into equations (12) and (13) we can get Loss, which is calculated as follows: (14) (15)。