Artificial Intelligence-Based Face Recognition Method and Related Devices
By building a face recognition network, using the self-attention layer and feature extraction layer to extract common features, the problem of low face recognition accuracy in different postures and background environments is solved, and higher robustness and accuracy are achieved.
Patent Information
- Application Number
- CN202310150334.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-09
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2043-02-09
AI Technical Summary
In the prior art, the face recognition method has low recognition accuracy in different postures and background environments, resulting in insufficient robustness and accuracy.
By collecting images of the same face in different states, a face recognition network is constructed, and common features are extracted using the self-attention layer and feature extraction layer. Combined with teacher-student network training, attention maps and significant feature maps are generated to obtain facial recognition results.
It improves the robustness and accuracy of face recognition, and can effectively eliminate the impact of face posture and background environment on recognition results.
Smart Images

Figure CN116030525B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and particularly relates to a face recognition method, device, electronic device, and storage medium based on artificial intelligence. Background Art
[0002] In scenarios such as electronic payment and financial risk control that require face recognition, face recognition results are usually obtained based on the collected face images. However, the poses and background environments of the faces in the collected face images are often different. At the same time, multiple faces may appear in a single face image, affecting the accuracy of face recognition.
[0003] Currently, to address the above problems, face detection algorithms are usually used to obtain the regional images of each face in the face image, and then face recognition is performed on the regional images of each face to obtain face recognition results. However, this method cannot eliminate the influence of different poses and background environments of the faces in the regional images on the accuracy of face recognition and is not applicable to different face recognition scenarios, resulting in low robustness and accuracy of face recognition. Summary of the Invention
[0004] In view of the above, it is necessary to propose a face recognition method and related devices based on artificial intelligence to solve the technical problem of how to improve the robustness and accuracy of face recognition. Among them, the related devices include a face recognition device, an electronic device, and a storage medium based on artificial intelligence.
[0005] This application provides a face recognition method based on artificial intelligence. The method includes:
[0006] Collect face images of the same face in different states to obtain an image subset of the face, and store all image subsets as a training image set;
[0007] Build a face recognition network;
[0008] Train the face recognition network based on the training image set to obtain a target face recognition network. The input of the target recognition network is a face image, and the output is an attention map and a significant feature map of the face image. The significant feature map includes common features of the same face in different states;
[0009] Input the face image to be recognized into the target face recognition network to output an attention map to be recognized and a significant feature map to be recognized. Input a reference image into the target face recognition network to output a reference attention map and a reference significant feature map. The reference image includes a face recognition label;
[0010] Obtain the face recognition result of the to-be-recognized face image based on the to-be-recognized attention map, the benchmark attention map, the to-be-recognized salient feature map, the benchmark salient feature map, and the face recognition label.
[0011] In some embodiments, collecting face images of the same face in different states to obtain an image subset of the face and storing all image subsets as a training image set includes:
[0012] Collect multiple face images of the same face in different states, where the different states include at least one of different face postures and different background environments;
[0013] Store multiple face images of the same face as the image subset of the face;
[0014] Collect image subsets of different faces to obtain a training image set.
[0015] In some embodiments, building the face recognition network includes:
[0016] The input of the face recognition network is a face image, and the face recognition network includes a feature extraction layer, a self-attention layer, and a fusion layer;
[0017] The feature extraction layer extracts features from the input face image to obtain a feature map;
[0018] The self-attention layer divides the input face image into a preset number of equal-sized subgraphs and extracts features from all subgraphs to obtain an attention map, and the size of the attention map is the same as that of the feature map;
[0019] The fusion layer fuses the attention map and the feature map to obtain a salient feature map of the input face image, and the salient feature map is used to reflect the features of the face region in the input face image, and the salient feature map satisfies the relationship:
[0020] R = F ⊙ A
[0021] where F is the feature map, A is the self-attention map, F ⊙ A represents calculating the Hadamard product between F and A, and R is the salient feature map;
[0022] Use the attention map and the salient feature map as the output results of the face recognition network.
[0023] In some embodiments, training the face recognition network based on the training image set to obtain a target face recognition network includes:
[0024] A1. Build two face recognition networks. Connect a classification layer at the end of one face recognition network as the student network, and connect a bias layer and a classification layer in sequence at the end of the other face recognition network as the teacher network. The bias layer includes a bias matrix;
[0025] A2. Randomly select two face images from any image subset of the training image set without replacement as a training pair. Denote the two face images in the training pair as the first face image and the second face image respectively;
[0026] A3. Input the first face image into the student network and the teacher network respectively to obtain the first student output and the first teacher output. Input the second face image into the student network and the teacher network respectively to obtain the second student output and the second teacher output;
[0027] A4. Calculate the value of a preset loss function based on the first student output, the first teacher output, the second student output, and the second teacher output, and update the network parameters in the student network based on the preset loss function value and the gradient descent method;
[0028] A5. Update the network parameters in the teacher network except for the bias layer based on the updated network parameters in the student network. The update process of the network parameters satisfies the relational expression:
[0029]
[0030] where θ t is the network parameter in the teacher network except for the bias layer before update, is the network parameter in the teacher network except for the bias layer after update, θ s is the network parameter in the student network after update, and λ is the parameter update coefficient, with a value range of [0, 1];
[0031] A6. Update the bias matrix in the bias layer based on the first teacher output and the second teacher output. The update process of the bias matrix satisfies the relational expression:
[0032]
[0033] where C is the bias matrix before update, are the significant feature maps output by the face recognition network in the teacher network when the first face image and the second face image are input respectively, and C * is the bias matrix after update, and m is the bias update coefficient, with a value range of [0, 1];
[0034] A7. Repeat steps A2 to A6, continuously select new training pairs from the training image set, and continuously update the network parameters in the student network and the teacher network until the preset loss function value no longer changes, then stop the update to obtain the trained student network and teacher network;
[0035] A8. Respectively extract the face recognition networks in the trained student network and teacher network, and fuse the network parameters of the two face recognition networks to obtain a target face recognition network. The fusion process satisfies the relational expression:
[0036]
[0037] Wherein, are the network parameters of the face recognition network in the trained teacher network, are the network parameters of the face recognition network in the trained student network, θ final are the network parameters of the target face recognition network, and ε is a fusion coefficient with a value range of [0, 1].
[0038] In some embodiments, the bias layer is used to add the bias matrix to the significant feature map output by the face recognition network in the teacher network to obtain a bias feature map. The bias feature map satisfies the relational expression:
[0039]
[0040] Wherein, R t is the significant feature map output by the face recognition network in the teacher network, C is the bias matrix, and the size of the bias matrix is the same as that of the significant feature map, is the bias feature map; in the teacher network, the bias feature map is input into the classification layer to obtain the output result of the teacher network.
[0041] In some embodiments, the preset loss function value satisfies the relational expression:
[0042]
[0043] Wherein, are respectively the values of the i-th row in the first teacher output and the second teacher output, are respectively the values of the i-th row in the first student output and the second student output, N is the number of rows of the first student output, the first teacher output, the second student output, and the second teacher output, and Loss is the preset loss function value.
[0044] In some embodiments, the number of the reference images is at least one, and each reference image corresponds to a reference attention map and a reference salient feature map. Obtaining the face recognition result of the to-be-recognized face image based on the to-be-recognized attention map, the reference attention map, the to-be-recognized salient feature map, the reference salient feature map, and the face recognition label includes:
[0045] Calculating the similarity between the to-be-recognized attention map and the reference attention map as a first similarity;
[0046] Calculating the similarity between the to-be-recognized salient feature map and the reference salient feature map as a second similarity;
[0047] Performing weighted summation on the first similarity and the second similarity to obtain the target similarity between the to-be-recognized face image and each reference image, and obtaining the maximum similarity among all the target similarities;
[0048] Comparing the maximum similarity with a preset pre-threshold. If the maximum similarity is greater than the preset pre-threshold, using the face recognition label of the reference image corresponding to the maximum similarity as the face recognition result of the to-be-recognized face image; if the maximum similarity is not greater than the preset pre-threshold, sending an alarm in a preset manner.
[0049] An embodiment of the present application further provides a face recognition device based on artificial intelligence, and the device includes:
[0050] An acquisition unit, configured to acquire face images of the same face in different states to obtain an image subset of the face, and store all the image subsets as a training image set;
[0051] A building unit, configured to build a face recognition network;
[0052] A training unit, configured to train the face recognition network based on the training image set to obtain a target face recognition network. The input of the target recognition network is a face image, and the output is the attention map and the salient feature map of the face image. The salient feature map includes the common features of the same face in different states;
[0053] An input unit, configured to input the to-be-recognized face image into the target face recognition network to output a to-be-recognized attention map and a to-be-recognized salient feature map, and input the reference image into the target face recognition network to output a reference attention map and a reference salient feature map. The reference image includes a face recognition label;
[0054] A face recognition unit, configured to obtain a face recognition result of the to-be-recognized face image based on the to-be-recognized attention map, the reference attention map, the to-be-recognized salient feature map, the reference salient feature map, and the face recognition label.
[0055] An embodiment of the present application further provides an electronic device, which includes:
[0056] A memory, storing at least one instruction;
[0057] A processor, configured to execute the instruction stored in the memory to implement the artificial intelligence-based face recognition method.
[0058] An embodiment of the present application further provides a computer-readable storage medium, in which at least one instruction is stored, and the at least one instruction is executed by a processor in an electronic device to implement the artificial intelligence-based face recognition method.
[0059] In summary, through the self-attention layer and the feature extraction layer in the target face recognition network, the present application can extract the common features of the same face in different states from a face image, avoid the influence of state factors such as face pose and background environment on the face recognition result, and improve the robustness and accuracy of face recognition. BRIEF DESCRIPTION OF THE DRAWINGS
[0060] Figure 1 is a flowchart of a preferred embodiment of the artificial intelligence-based face recognition method involved in the present application.
[0061] Figure 2 is a schematic structural diagram of the face recognition network involved in the present application.
[0062] Figure 3 is a schematic structural diagram of the student network and the teacher network involved in the present application.
[0063] Figure 4 is a functional module diagram of a preferred embodiment of the artificial intelligence-based face recognition device involved in the present application.
[0064] Figure 5 is a schematic structural diagram of an electronic device of a preferred embodiment of the artificial intelligence-based face recognition method involved in the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0065] In order to more clearly understand the objectives, features, and advantages of the present application, the present application will be described in detail below with reference to the accompanying drawings and specific embodiments. It should be noted that, without conflict, the embodiments of the present application and the features in the embodiments may be combined with each other. In the following description, many specific details are set forth in order to fully understand the present application. The described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments.
[0066] In addition, the terms "first" and "second" are only used for descriptive purposes, and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more of the described features. In the description of the present application, the meaning of "a plurality" is two or more, unless otherwise specifically defined.
[0067] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs. The terms used in the description of the present application in this specification are only for the purpose of describing specific embodiments, and are not intended to limit the present application. The term "and / or" used herein includes any and all combinations of one or more of the related listed items.
[0068] An embodiment of the present application provides a face recognition method based on artificial intelligence, which can be applied to one or more electronic devices. An electronic device is a device that can automatically perform numerical calculations and / or information processing according to pre-set or stored instructions. Its hardware includes, but is not limited to, a microprocessor, an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a digital signal processor (DSP), an embedded device, etc.
[0069] An electronic device can be any electronic product that can perform human-computer interaction with a customer. For example, a personal computer, a tablet computer, a smart phone, a personal digital assistant (PDA), a game console, an Internet Protocol Television (IPTV), a smart wearable device, etc.
[0070] The electronic device may also include a network device and / or a client device. Among them, the network device includes, but is not limited to, a single network server, a server group composed of multiple network servers, or a cloud composed of a large number of hosts or network servers based on cloud computing.
[0071] The network where the electronic device is located includes, but is not limited to, the Internet, wide area network, metropolitan area network, local area network, virtual private network (VPN), etc.
[0072] As Figure 1 shown, it is a flowchart of a preferred embodiment of the face recognition method based on artificial intelligence in this application. According to different requirements, the order of steps in this flowchart can be changed, and some steps can be omitted.
[0073] The face recognition method based on artificial intelligence provided in the embodiments of this application can be applied to any scenario that requires face recognition, and then this method can be applied to products in these scenarios, such as, for example, electronic transactions, electronic payments, securities banks, and so on.
[0074] S10, collect face images of the same face in different states to obtain an image subset of the face, and store all image subsets as a training image set.
[0075] In an optional embodiment, the collecting face images of the same face in different states to obtain an image subset of the face, and storing all image subsets as a training image set includes:
[0076] Collect multiple face images of the same face in different states, where the different states include at least one of different face postures and different background environments;
[0077] Store multiple face images of the same face as the image subset of the face;
[0078] Collect image subsets of different faces to obtain a training image set.
[0079] In this optional embodiment, for the same face, multiple face images can be collected by changing the face posture, for example, raising the head, lowering the head, tilting the head, turning the head to the left, and so on; multiple face images can also be collected by changing the background environment, or the face posture and background environment can be changed simultaneously. All the collected face images are used as the image subset of the face. According to the same method, image subsets of different faces are collected, and all image subsets are used as a training image set.
[0080] It should be noted that all face images in the same image subset belong to the same face.
[0081] In this way, a training image set is obtained. The training image set includes multiple image subsets, and each image subset includes multiple face images of the same face in different states, providing a data basis for subsequent training of the face recognition network.
[0082] S11. Build a face recognition network.
[0083] In an optional embodiment, the building of the face recognition network includes:
[0084] The input of the face recognition network is a face image. The face recognition network includes a feature extraction layer, a self-attention layer, and a fusion layer;
[0085] The feature extraction layer extracts features from the input face image to obtain a feature map;
[0086] The self-attention layer divides the input face image into a preset number of equal-sized sub-images, and extracts features from all the sub-images to obtain an attention map, and the size of the attention map is the same as that of the feature map;
[0087] The fusion layer fuses the attention map and the feature map to obtain a significant feature map of the input face image. The significant feature map is used to reflect the features of the face region in the input face image, and the significant feature map satisfies the relational expression:
[0088] R = F ⊙ A
[0089] where F is the feature map, A is the self-attention map, F ⊙ A represents the calculation of the Hadamard product between F and A, and R is the significant feature map;
[0090] The attention map and the significant feature map are used as the output results of the face recognition network.
[0091] In this optional embodiment, the structural schematic diagram of the face recognition network is as Figure 2 shown. The feature extraction layer can adopt existing convolutional neural networks such as ResNet and ReXNet, and the self-attention layer adopts existing feature extraction networks based on the self-attention mechanism (Self-Attention) such as Non-Local Network and Vision Transformer, which are not limited in this application.
[0092] In this optional embodiment, in order to ensure that the size of the attention map is the same as that of the feature map, in the self-attention layer, when dividing the input face image into a preset number of equal-sized sub-images, the preset number is the same as the size of the feature map. Exemplarily, if the size of the feature map is 5×5, the preset number is 25.
[0093] In this way, the construction of the face recognition network is completed, providing a network foundation for subsequent face recognition.
[0094] S12. Based on the training image set, train the face recognition network to obtain a target face recognition network. The input of the target recognition network is a face image, and the output is the attention map and the significant feature map of the face image. The significant feature map includes the common features of the same face in different states.
[0095] In an optional embodiment, the training of the face recognition network based on the training image set to obtain a target face recognition network includes:
[0096] A1. Build two face recognition networks. Connect a classification layer at the end of one face recognition network as the student network, and connect a bias layer and a classification layer in sequence at the end of the other face recognition network as the teacher network. The bias layer includes a bias matrix.
[0097] In this optional embodiment, the bias layer is used to add the bias matrix to the significant feature map output by the face recognition network in the teacher network to obtain a bias feature map. The bias feature map satisfies the following relationship:
[0098]
[0099] where R t is the significant feature map output by the face recognition network in the teacher network, C is the bias matrix, the size of the bias matrix is the same as that of the significant feature map, is the bias feature map; in the teacher network, the bias feature map is input into the classification layer to obtain the output result of the teacher network.
[0100] It should be noted that adding a bias layer in the teacher network is to encourage the output result of the teacher network to be close to a uniform distribution and avoid the occurrence of degenerate solutions. The structural schematic diagrams of the student network and the teacher network are as Figure 3 shown.
[0101] A2. Randomly select two face images from any image subset of the training image set without replacement as a training pair. Denote the two face images in the training pair as the first face image and the second face image respectively.
[0102] Among them, the first face image and the second face image in the training pair are face images of the same face collected in different states.
[0103] A3. Input the first face image into the student network and the teacher network respectively to obtain a first student output and a first teacher output, and input the second face image into the student network and the teacher network respectively to obtain a second student output and a second teacher output;
[0104] Among them, the first student output, the first teacher output, the second student output, and the second teacher output are all category vectors of N rows and 1 column, and N is related to the structure of the classification layer.
[0105] A4. Calculate a preset loss function value based on the first student output, the first teacher output, the second student output, and the second teacher output, and update the network parameters in the student network based on the preset loss function value and the gradient descent method;
[0106] In this optional embodiment, the preset loss function value satisfies the relationship:
[0107]
[0108] Among them, are the values of the i-th row in the first teacher output and the second teacher output respectively, are the values of the i-th row in the first student output and the second student output respectively, N is the number of rows of the first student output, the first teacher output, the second student output, and the second teacher output, and Loss is the preset loss function value.
[0109] It should be noted that the preset loss function value constrains the face images in different states in the training pair to have the same output result, so that the face recognition networks in the student network and the teacher network can learn the significant common features of the same face in different states.
[0110] A5. Update the network parameters in the teacher network except the bias layer based on the updated network parameters in the student network;
[0111] In this optional embodiment, in the student network and the teacher network, the network structures except the bias layer are the same. Therefore, the network parameters in the teacher network except the bias layer can be updated based on the updated network parameters in the student network, and the update process of the network parameters satisfies the relationship:
[0112]
[0113] Among them, θ t is the network parameter in the teacher network except the bias layer before update, is the network parameter in the teacher network except the bias layer after update, θs is the updated network parameters in the student network, λ is the parameter update coefficient, and the value range is [0, 1]. Among them, the parameter update coefficient is used to control the update speed of the network parameters in the teacher network, and the parameter update coefficient λ takes the value of 0.5.
[0114] A6. Update the bias matrix in the bias layer based on the first teacher output and the second teacher output;
[0115] In this optional embodiment, the update process of the bias matrix satisfies the relational expression:
[0116]
[0117] where C is the bias matrix before update, are the significant feature maps output by the face recognition network in the teacher network when the first face image and the second face image are input respectively, C * is the updated bias matrix, m is the bias update coefficient, and the value range is [0, 1]. Among them, the bias update coefficient is used to control the update speed of the bias matrix in the bias layer, and the bias update coefficient m takes the value of 0.5.
[0118] A7. Repeat steps A2 to A6, continuously select new training pairs from the training image set, and continuously update the network parameters in the student network and the teacher network until the preset loss function value no longer changes, then stop the update to obtain the trained student network and teacher network;
[0119] A8. Extract the face recognition networks in the trained student network and teacher network respectively, and fuse the network parameters of the two face recognition networks to obtain the target face recognition network.
[0120] In this optional embodiment, the network structures of the face recognition networks in the trained student network and teacher network are the same, but the network parameters of the face recognition networks are different. Fuse the network parameters of the two face recognition networks to obtain the target face recognition network, and the fusion process satisfies the relational expression:
[0121]
[0122] where, are the network parameters of the face recognition network in the trained teacher network, are the network parameters of the face recognition network in the trained student network, θ final are the network parameters of the target face recognition network, ε is the fusion coefficient, and the value range is [0, 1]. Among them, in this optional embodiment, the fusion coefficient ε = 0.5.
[0123] In this optional embodiment, the target face recognition network can learn the remarkable common features of the same face in different states, thereby eliminating the influence of face pose and background environment in the face image on face recognition. In the attention map output by the target face recognition network, the pixel values at the positions related to the common features of the same face in different states are relatively high, and the pixel values at the positions unrelated to the common features are relatively low, which can reflect the spatial distribution characteristics of the face common features; the significant feature map output by the target face recognition network includes the remarkable common features of the same face in different states.
[0124] In this way, the training process of the face recognition network is completed to obtain the target face recognition network. The target face recognition network can extract the remarkable common features of the same face in different states, eliminate the influence of face pose and background environment in the face image on face recognition, and improve the accuracy and robustness of face recognition.
[0125] S13. Input the face image to be recognized into the target face recognition network to output the attention map to be recognized and the significant feature map to be recognized. Input the reference image into the target face recognition network to output the reference attention map and the reference significant feature map. The reference image includes a face recognition label.
[0126] In an optional embodiment, collect the face image to be recognized and input it into the target face recognition network to output the attention map to be recognized and the significant feature map to be recognized corresponding to the face image to be recognized; input the reference image into the target face recognition network to output the reference attention map and the reference significant feature map corresponding to the reference image. Among them, there is at least one reference image. Each reference image includes a frontal image of a pre-collected face, and the reference image includes a face recognition label, and the face recognition label is the identity information of the face in the reference image.
[0127] In this way, the attention map to be recognized and the significant feature map to be recognized corresponding to the face image to be recognized, and the reference attention map and the reference significant feature map corresponding to at least one reference image are obtained by means of the target recognition network, providing a data basis for face recognition.
[0128] S14. Obtain the face recognition result of the face image to be recognized based on the attention map to be recognized, the reference attention map, the significant feature map to be recognized, the reference significant feature map, and the face recognition label.
[0129] In an optional embodiment, the number of the reference images is at least one, and each reference image corresponds to a reference attention map and a reference salient feature map. Obtaining the face recognition result of the to-be-recognized face image based on the to-be-recognized attention map, the reference attention map, the to-be-recognized salient feature map, the reference salient feature map, and the face recognition label includes:
[0130] Calculating the similarity between the to-be-recognized attention map and the reference attention map as the first similarity;
[0131] Calculating the similarity between the to-be-recognized salient feature map and the reference salient feature map as the second similarity;
[0132] Performing weighted summation on the first similarity and the second similarity to obtain the target similarity between the to-be-recognized face image and each reference image, and obtaining the maximum similarity among all the target similarities;
[0133] Comparing the maximum similarity with a preset pre-threshold. If the maximum similarity is greater than the preset pre-threshold, using the face recognition label of the reference image corresponding to the maximum similarity as the face recognition result of the to-be-recognized face image; if the maximum similarity is not greater than the preset pre-threshold, sending an alarm in a preset manner.
[0134] Wherein, the preset manner includes at least voice reminder and phone reminder; the preset pre-threshold is 0.6.
[0135] In an optional embodiment, the first similarity satisfies the relational expression:
[0136] Sim1 = exp(-D(A′, A j ))
[0137] Wherein, A′ is the to-be-recognized attention map, and A j is the reference attention map corresponding to the reference image j, D(A′, A j ) is the distance between A′ and A j , and Sim1 is the first similarity. Wherein, the distance can be cosine distance, Euclidean distance, Hamming distance, etc., and the present application does not make any limitation.
[0138] In this optional embodiment, the second similarity satisfies the relational expression:
[0139] Sim2 = exp(-D(R′, R j ))
[0140] Wherein, R′ is the to-be-recognized salient feature map, and R j is the reference salient feature map corresponding to the reference image j, D(R′, R j ) is the distance between R′ and Rj The distance is Sim2 for the second similarity. Among them, the distance can be cosine distance, Euclidean distance, Hamming distance, etc., which is not limited in this application.
[0141] In this optional embodiment, the target similarity satisfies the relational expression:
[0142] Sim * = δSim1 + (1 - δ)Sim2
[0143] Among them, Sim1 is the first similarity, Sim2 is the second similarity, Sim * is the target similarity, δ is the weighting coefficient, and its value range is [0, 1]. Among them, in this optional embodiment, the value of the weighting coefficient is δ = 0.4.
[0144] In this way, based on the reference attention map and reference salient feature map of the reference image, as well as the to-be-recognized attention map and to-be-recognized salient feature map of the to-be-recognized face image, the face recognition result is obtained.
[0145] It can be seen from the above technical solutions that through the self-attention layer and feature extraction layer in the target face recognition network of this application, common features of the same face in different states can be extracted from the face image, avoiding the influence of state factors such as face pose and background environment on the face recognition result, and improving the robustness and accuracy of face recognition.
[0146] Please refer to Figure 4 , Figure 4 which is the functional module diagram of a preferred embodiment of the face recognition device based on artificial intelligence of this application. The face recognition device 11 based on artificial intelligence includes a collection unit 110, a construction unit 111, a training unit 112, an input unit 113, and a face recognition unit 114. The module / unit referred to in this application means a series of computer-readable instruction segments that can be executed by a processor 13 and can complete fixed functions, and are stored in a memory 12. In this embodiment, the functions of each module / unit will be described in detail in subsequent embodiments.
[0147] In an optional embodiment, the collection unit 110 is used to collect face images of the same face in different states to obtain an image subset of the face, and store all the image subsets as a training image set.
[0148] In an optional embodiment, the collecting face images of the same face in different states to obtain an image subset of the face, and storing all the image subsets as a training image set includes:
[0149] Collecting multiple face images of the same face in different states, where the different states include at least one of different face poses and different background environments;
[0150] Store multiple face images of the same face as an image subset of the face;
[0151] Collect image subsets of different faces to obtain a training image set.
[0152] In this optional embodiment, for the same face, multiple face images can be collected by changing the face pose, for example, raising the head, lowering the head, tilting the head, turning the head to the left, etc.; multiple face images can also be collected by changing the background environment, or the face pose and the background environment can be changed simultaneously. All the collected face images are used as the image subset of the face. Image subsets of different faces are collected in the same way, and all the image subsets are used as the training image set.
[0153] It should be noted that all the face images in the same image subset belong to the same face.
[0154] In an optional embodiment, the building unit 111 is used to build a face recognition network.
[0155] In an optional embodiment, building the face recognition network includes:
[0156] The input of the face recognition network is a face image. The face recognition network includes a feature extraction layer, a self-attention layer, and a fusion layer;
[0157] The feature extraction layer extracts features from the input face image to obtain a feature map;
[0158] The self-attention layer divides the input face image into a preset number of equal-sized sub-images, and extracts features from all the sub-images to obtain an attention map, and the size of the attention map is the same as that of the feature map;
[0159] The fusion layer fuses the attention map and the feature map to obtain a significant feature map of the input face image. The significant feature map is used to reflect the features of the face region in the input face image. The significant feature map satisfies the relational expression:
[0160] R = F ⊙ A
[0161] where F is the feature map, A is the self-attention map, F ⊙ A represents the Hadamard product between F and A, and R is the significant feature map;
[0162] The attention map and the significant feature map are used as the output results of the face recognition network.
[0163] In this optional embodiment, the structural schematic diagram of the face recognition network is as Figure 2As shown, the feature extraction layer can adopt existing convolutional neural networks such as ResNet and ReXNet, and the self-attention layer adopts existing feature extraction networks based on the self-attention mechanism (Self-Attention) such as Non-Local Network and Vision Transformer, which are not limited in this application.
[0164] In this optional embodiment, to ensure that the attention map has the same size as the feature map, in the self-attention layer, when the input face image is segmented into a preset number of equal-sized sub-images, the preset number is the same as the size of the feature map. Exemplarily, if the size of the feature map is 5×5, the preset number is 25.
[0165] In an optional embodiment, the training unit 112 is configured to train the face recognition network based on the training image set to obtain a target face recognition network. The input of the target recognition network is a face image, and the outputs are the attention map and the significant feature map of the face image. The significant feature map includes the common features of the same face in different states.
[0166] In an optional embodiment, training the face recognition network based on the training image set to obtain a target face recognition network includes:
[0167] A1. Build two face recognition networks. Connect a classification layer at the end of one face recognition network as the student network, and connect a bias layer and a classification layer in sequence at the end of the other face recognition network as the teacher network. The bias layer includes a bias matrix.
[0168] In this optional embodiment, the bias layer is used to add the bias matrix to the significant feature map output by the face recognition network in the teacher network to obtain a bias feature map. The bias feature map satisfies the following relationship:
[0169]
[0170] where R t is the significant feature map output by the face recognition network in the teacher network, C is the bias matrix, the bias matrix has the same size as the significant feature map, is the bias feature map; in the teacher network, the bias feature map is input to the classification layer to obtain the output result of the teacher network.
[0171] It should be noted that adding a bias layer in the teacher network is to encourage the output result of the teacher network to be close to a uniform distribution and avoid the occurrence of degenerate solutions. The structural diagrams of the student network and the teacher network are as Figure 3 shown.
[0172] A2. Randomly select two face images without replacement from any one image subset of the training image set as a training pair, and denote the two face images in the training pair as the first face image and the second face image respectively;
[0173] Among them, the first face image and the second face image in the training pair are face images of the same person collected under different states.
[0174] A3. Input the first face image into the student network and the teacher network respectively to obtain a first student output and a first teacher output, and input the second face image into the student network and the teacher network respectively to obtain a second student output and a second teacher output;
[0175] Among them, the first student output, the first teacher output, the second student output, and the second teacher output are all category vectors of N rows and 1 column, and N is related to the structure of the classification layer.
[0176] A4. Calculate a preset loss function value based on the first student output, the first teacher output, the second student output, and the second teacher output, and update the network parameters in the student network based on the preset loss function value and the gradient descent method;
[0177] In this optional embodiment, the preset loss function value satisfies the relational expression:
[0178]
[0179] Among them, are the values of the i-th row in the first teacher output and the second teacher output respectively, are the values of the i-th row in the first student output and the second student output respectively, N is the number of rows of the first student output, the first teacher output, the second student output, and the second teacher output, and Loss is the preset loss function value.
[0180] It should be noted that the preset loss function value constrains the face images in different states in the training pair to have the same output result, so that the face recognition networks in the student network and the teacher network can learn the significant common features of the same face in different states.
[0181] A5. Update the network parameters in the teacher network except for the bias layer based on the updated network parameters in the student network;
[0182] In this alternative embodiment, the network structures of the student network and the teacher network are the same except for the bias layer. Therefore, the network parameters of the teacher network except for the bias layer can be updated based on the updated network parameters in the student network. The update process of the network parameters satisfies the following relationship:
[0183]
[0184] where θ t is the network parameter of the teacher network except for the bias layer before update, is the network parameter of the teacher network except for the bias layer after update, θ s is the network parameter of the student network after update, and λ is the parameter update coefficient, whose value range is [0, 1]. Among them, the parameter update coefficient is used to control the update speed of the network parameters in the teacher network, and the parameter update coefficient λ takes the value of 0.5.
[0185] A6. Update the bias matrix in the bias layer based on the first teacher output and the second teacher output;
[0186] In this alternative embodiment, the update process of the bias matrix satisfies the following relationship:
[0187]
[0188] where C is the bias matrix before update, are the significant feature maps output by the face recognition network in the teacher network when the first face image and the second face image are input respectively, and C * is the bias matrix after update, and m is the bias update coefficient, whose value range is [0, 1]. Among them, the bias update coefficient is used to control the update speed of the bias matrix in the bias layer, and the bias update coefficient m takes the value of 0.5.
[0189] A7. Repeat steps A2 to A6, continuously select new training pairs from the training image set, and continuously update the network parameters in the student network and the teacher network until the preset loss function value no longer changes, then stop the update to obtain the trained student network and teacher network;
[0190] A8. Extract the face recognition networks in the trained student network and teacher network respectively, and fuse the network parameters of the two face recognition networks to obtain the target face recognition network.
[0191] In this optional embodiment, the network structures of the face recognition networks in the trained student network and teacher network are the same, but the network parameters of the face recognition networks are different. The network parameters of the two face recognition networks are fused to obtain a target face recognition network, and the fusion process satisfies the relational expression:
[0192]
[0193] Wherein, are the network parameters of the face recognition network in the trained teacher network, are the network parameters of the face recognition network in the trained student network, and θ final are the network parameters of the target face recognition network, and ε is a fusion coefficient with a value range of [0, 1]. In this optional embodiment, the fusion coefficient ε = 0.5.
[0194] In this optional embodiment, the target face recognition network can learn the significant common features of the same face in different states, thereby eliminating the influence of face pose and background environment in the face image on face recognition. In the attention map output by the target face recognition network, the pixel values at the positions related to the common features of the same face in different states are higher, and the pixel values at the positions unrelated to the common features are lower, which can reflect the spatial distribution characteristics of the face common features; the significant feature map output by the target face recognition network includes the significant common features of the same face in different states.
[0195] In an optional embodiment, the input unit 113 is configured to input the face image to be recognized into the target face recognition network to output an attention map to be recognized and a significant feature map to be recognized, and input a reference image into the target face recognition network to output a reference attention map and a reference significant feature map, where the reference image includes a face recognition label.
[0196] In an optional embodiment, a face image to be recognized is collected and input into the target face recognition network to output an attention map to be recognized and a significant feature map to be recognized corresponding to the face image to be recognized; a reference image is input into the target face recognition network to output a reference attention map and a reference significant feature map corresponding to the reference image. Wherein, there is at least one reference image, and each reference image includes a frontal image of a pre-collected face, and the reference image includes a face recognition label, and the face recognition label is the identity information of the face in the reference image.
[0197] In an optional embodiment, the face recognition unit 114 is configured to obtain the face recognition result of the face image to be recognized based on the attention map to be recognized, the reference attention map, the significant feature map to be recognized, the reference significant feature map, and the face recognition label.
[0198] In an alternative embodiment, the number of the reference images is at least one, and each reference image corresponds to a reference attention map and a reference salient feature map. Obtaining the face recognition result of the to-be-recognized face image based on the to-be-recognized attention map, the reference attention map, the to-be-recognized salient feature map, the reference salient feature map, and the face recognition label includes:
[0199] Calculating the similarity between the to-be-recognized attention map and the reference attention map as a first similarity;
[0200] Calculating the similarity between the to-be-recognized salient feature map and the reference salient feature map as a second similarity;
[0201] Performing weighted summation on the first similarity and the second similarity to obtain the target similarity between the to-be-recognized face image and each reference image, and obtaining the maximum similarity among all the target similarities;
[0202] Comparing the maximum similarity with a preset threshold. If the maximum similarity is greater than the preset threshold, using the face recognition label of the reference image corresponding to the maximum similarity as the face recognition result of the to-be-recognized face image; if the maximum similarity is not greater than the preset threshold, giving an alarm in a preset manner.
[0203] Wherein, the preset manner includes at least voice reminder and phone reminder; the preset threshold is 0.6.
[0204] In an alternative embodiment, the first similarity satisfies the relational expression:
[0205] Sim1 = exp(-D(A′, A j ))
[0206] Wherein, A′ is the to-be-recognized attention map, and A j is the reference attention map corresponding to the reference image j, D(A′, A j ) is the distance between A′ and A j , and Sim1 is the first similarity. Wherein, the distance can be cosine distance, Euclidean distance, Hamming distance, etc., and the present application does not make any limitation.
[0207] In this alternative embodiment, the second similarity satisfies the relational expression:
[0208] Sim2 = exp(-D(R′, R j ))
[0209] Wherein, R′ is the to-be-recognized salient feature map, and R jThe reference salient feature map corresponding to the reference image j, D(R′,R j ) is the distance between R′ and R j . The distance can be cosine distance, Euclidean distance, Hamming distance, etc., which is not limited in this application.
[0210] In this optional embodiment, the target similarity satisfies the relational expression:
[0211] Sim * =δSim1+(1 - δ)Sim2
[0212] where Sim1 is the first similarity, Sim2 is the second similarity, Sim * is the target similarity, and δ is the weighting coefficient with a value range of [0,1]. In this optional embodiment, the value of the weighting coefficient is δ = 0.4.
[0213] It can be seen from the above technical solutions that in this application, the self - attention layer and the feature extraction layer in the target face recognition network can extract the common features of the same face in different states from the face image, avoid the influence of state factors such as face pose and background environment on the face recognition result, and improve the robustness and accuracy of face recognition.
[0214] Please refer to Figure 5 , which is a schematic structural diagram of an electronic device provided by an embodiment of this application. The electronic device 1 includes a memory 12 and a processor 13. The memory 12 is used to store computer - readable instructions, and the processor 13 is used to execute the computer - readable instructions stored in the memory to implement the artificial - intelligence - based face recognition method described in any of the above embodiments.
[0215] In an optional embodiment, the electronic device 1 further includes a bus and a computer program stored in the memory 12 and executable on the processor 13, such as an artificial - intelligence - based face recognition program.
[0216] Figure 5 Only the electronic device 1 with the memory 12 and the processor 13 is shown. Those skilled in the art can understand that Figure 5 the shown structure does not limit the electronic device 1, and it may include fewer or more components than shown, or combine some components, or have different component arrangements.
[0217] Combined with Figure 1 , the memory 12 in the electronic device 1 stores multiple computer - readable instructions to implement an artificial - intelligence - based face recognition method, and the processor 13 can execute the multiple instructions to implement:
[0218] Collect face images of the same face in different states to obtain an image subset of the face, and store all image subsets as a training image set;
[0219] Build a face recognition network;
[0220] Train the face recognition network based on the training image set to obtain a target face recognition network. The input of the target recognition network is a face image, and the output is an attention map and a significant feature map of the face image. The significant feature map includes common features of the same face in different states;
[0221] Input the face image to be recognized into the target face recognition network to output an attention map to be recognized and a significant feature map to be recognized. Input a reference image into the target face recognition network to output a reference attention map and a reference significant feature map. The reference image includes a face recognition label;
[0222] Obtain the face recognition result of the face image to be recognized based on the attention map to be recognized, the reference attention map, the significant feature map to be recognized, the reference significant feature map, and the face recognition label.
[0223] Specifically, the specific implementation method of the processor 13 for the above instructions can refer to Figure 1 the description of the relevant steps in the corresponding embodiment, which will not be elaborated here.
[0224] Those skilled in the art can understand that the schematic diagram is only an example of the electronic device 1, and does not constitute a limitation on the electronic device 1. The electronic device 1 can be a bus structure or a star structure. The electronic device 1 can also include more or fewer other hardware or software than shown in the figure, or different component arrangements. For example, the electronic device 1 can also include input and output devices, network access devices, etc.
[0225] It should be noted that the electronic device 1 is only an example. Other existing or future electronic products that can be adapted to this application should also be included in the protection scope of this application and are included herein by reference.
[0226] Among them, the memory 12 includes at least one type of readable storage medium, which can be non-volatile or volatile. The readable storage medium includes flash memory, mobile hard disk, multimedia card, card-type memory (such as SD or DX memory, etc.), magnetic memory, magnetic disk, optical disc, etc. In some embodiments, the memory 12 can be an internal storage unit of the electronic device 1, such as the mobile hard disk of the electronic device 1. In other embodiments, the memory 12 can also be an external storage device of the electronic device 1, such as a plug-in mobile hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. equipped on the electronic device 1. The memory 12 can not only be used to store application software installed on the electronic device 1 and various types of data, such as the code of the face recognition program based on artificial intelligence, etc., but also be used to temporarily store the data that has been output or will be output.
[0227] In some embodiments, the processor 13 can be composed of integrated circuits. For example, it can be composed of a single packaged integrated circuit, or can be composed of multiple integrated circuits with the same or different functions packaged, including a combination of one or more Central Processing Units (CPUs), microprocessors, digital processing chips, graphics processors, and various control chips, etc. The processor 13 is the control core (Control Unit) of the electronic device 1, connecting various components of the entire electronic device 1 through various interfaces and lines, and by running or executing the programs or modules stored in the memory 12 (such as executing the face recognition program based on artificial intelligence, etc.), and calling the data stored in the memory 12, to execute various functions of the electronic device 1 and process data.
[0228] The processor 13 executes the operating system of the electronic device 1 and various installed application programs. The processor 13 executes the application programs to implement the steps in the above-mentioned various embodiments of the face recognition method based on artificial intelligence, such as Figure 1 the steps shown.
[0229] Exemplarily, the computer program can be divided into one or more modules / units, and the one or more modules / units are stored in the memory 12 and executed by the processor 13 to complete this application. The one or more modules / units can be a series of computer-readable instruction segments capable of completing specific functions, and these instruction segments are used to describe the execution process of the computer program in the electronic device 1. For example, the computer program can be divided into an acquisition unit 110, a construction unit 111, a training unit 112, an input unit 113, and a face recognition unit 114.
[0230] The integrated units implemented in the form of software functional modules can be stored in a computer-readable storage medium. The above-mentioned software functional modules stored in a storage medium include several instructions for causing a computer device (which can be a personal computer, a computer device, or a network device, etc.) or a processor to execute a part of the artificial intelligence-based face recognition method described in various embodiments of the present application.
[0231] If the modules / units integrated in the electronic device 1 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, to implement all or part of the processes in the above-mentioned embodiment methods of the present application, it can also be completed by a computer program instructing relevant hardware devices. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-mentioned various method embodiments can be implemented.
[0232] Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM, Read-Only Memory), random access memory, and other memories, etc.
[0233] Furthermore, the computer-readable storage medium mainly includes a storage program area and a storage data area. Among them, the storage program area can store an operating system, application programs required for at least one function, etc.; the storage data area can store data created according to the use of the blockchain node, etc.
[0234] The blockchain referred to in the present application is a new application mode of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanism, and encryption algorithm. A blockchain, essentially a decentralized database, is a string of data blocks generated by using cryptographic methods. Each data block contains information about a batch of network transactions, which is used to verify the validity (anti-counterfeiting) of the information and generate the next block. A blockchain can include a blockchain underlying platform, a platform product service layer, and an application service layer, etc.
[0235] The bus can be a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, or the like. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, in Figure 5 only one arrow is used to represent it, but it does not mean that there is only one bus or one type of bus. The bus is configured to enable connection communication between the memory 12 and at least one processor 13, etc.
[0236] The embodiment of the present application also provides a computer-readable storage medium (not shown in the figure). Computer-readable instructions are stored in the computer-readable storage medium, and the computer-readable instructions are executed by a processor in an electronic device to implement the face recognition method based on artificial intelligence described in any of the above embodiments.
[0237] In several embodiments provided by the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division, and there may be other division methods in actual implementation.
[0238] The modules described as separate components may or may not be physically separated. The components shown as modules may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0239] In addition, in each embodiment of the present application, the functional modules can be integrated in a processing unit, or each unit can exist physically separately, or two or more units can be integrated in one unit. The above integrated unit can be implemented in the form of hardware, or in the form of a combination of hardware and software functional modules.
[0240] Furthermore, obviously, the term "including" does not exclude other units or steps, and the singular does not exclude the plural. The multiple units or devices described in the specification can also be implemented by one unit or device through software or hardware. The terms such as first and second are used to represent names and do not indicate any specific order.
[0241] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application and not to limit them. Although the present application has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present application can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present application.
Claims
1. A face recognition method based on artificial intelligence, characterized in that, The method includes: Collecting face images of the same face in different states to obtain subsets of images of the face, and storing all subsets of images as a training image set; Constructing a face recognition network; Training the face recognition network based on the training image set to obtain a target face recognition network, where the input of the target face recognition network is a face image, and the output is an attention map and a significant feature map of the face image, and the significant feature map includes common features of the same face in different states; Inputting the face image to be recognized into the target face recognition network to output an attention map to be recognized and a significant feature map to be recognized, inputting a reference image into the target face recognition network to output a reference attention map and a reference significant feature map, and the reference image includes a face recognition label; Obtaining the face recognition result of the face image to be recognized based on the attention map to be recognized, the reference attention map, the significant feature map to be recognized, the reference significant feature map, and the face recognition label.
2. The face recognition method based on artificial intelligence according to claim 1, wherein The collecting face images of the same face in different states to obtain subsets of images of the face, and storing all subsets of images as a training image set includes: Collecting multiple face images of the same face in different states, where the different states include at least one of different face postures and different background environments; Storing multiple face images of the same face as subsets of images of the face; Collecting subsets of images of different faces to obtain a training image set.
3. The face recognition method based on artificial intelligence according to claim 1, characterized in that The constructing a face recognition network includes: The input of the face recognition network is a face image, and the face recognition network includes a feature extraction layer, a self-attention layer, and a fusion layer; The feature extraction layer extracts features from the input face image to obtain a feature map; The self-attention layer divides the input face image into a preset number of equal-sized sub-images, and extracts features from all sub-images to obtain an attention map, and the size of the attention map is the same as that of the feature map; The fusion layer fuses the attention map and the feature map to obtain a significant feature map of the input face image, and the significant feature map is used to reflect the features of the face region in the input face image, and the significant feature map satisfies the relationship: R = F ⊙ A where F is the feature map, A is the self-attention map, F ⊙ A represents calculating the Hadamard product between F and A, and R is the significant feature map; Taking the attention map and the significant feature map as the output results of the face recognition network.
4. The face recognition method based on artificial intelligence according to claim 1, characterized in that, The training the face recognition network based on the training image set to obtain a target face recognition network includes: A1. Constructing two face recognition networks, connecting a classification layer at the end of one face recognition network as a student network, and connecting a bias layer and a classification layer in sequence at the end of the other face recognition network as a teacher network, and the bias layer includes a bias matrix; A2. Randomly selecting two face images without replacement from any one subset of images in the training image set as a training pair, and respectively denoting the two face images in the training pair as a first face image and a second face image; A3. Input the first face image into the student network and the teacher network respectively to obtain a first student output and a first teacher output, and input the second face image into the student network and the teacher network respectively to obtain a second student output and a second teacher output; A4. Calculate a preset loss function value based on the first student output, the first teacher output, the second student output and the second teacher output, and update the network parameters in the student network based on the preset loss function value and the gradient descent method; A5. Update the network parameters in the teacher network except the bias layer based on the updated network parameters in the student network, and the update process of the network parameters satisfies the relational expression: where θ t is the network parameter of the teacher network except for the bias layer as described before the update, is the network parameter of the teacher network except for the bias layer after the update, θ s is the updated network parameter of the student network, and λ is the parameter update coefficient, whose value range is [0, 1]; A6. Update the bias matrix in the bias layer based on the first teacher output and the second teacher output, and the update process of the bias matrix satisfies the relational expression: Among them, C is the bias matrix before update, are the saliency feature maps output by the face recognition network in the teacher network when the first face image and the second face image are input respectively, and C * is the bias matrix after update, m is the bias update coefficient, and its value range is [0, 1]; A7. Repeat steps A2 to A6, continuously select new training pairs from the training image set, and continuously update the network parameters in the student network and the teacher network until the preset loss function value no longer changes, then stop the update to obtain the trained student network and teacher network; A8. Extract the face recognition networks in the trained student network and teacher network respectively, and fuse the network parameters of the two face recognition networks to obtain a target face recognition network, and the fusion process satisfies the relational expression: Among them, are the network parameters of the face recognition network in the trained teacher network, are the network parameters of the face recognition network in the trained student network, and θ final are the network parameters of the target face recognition network, and ε is the fusion coefficient with a value range of [0, 1].
5. The face recognition method based on artificial intelligence according to claim 4, wherein, The bias layer is used to add the bias matrix to the significant feature map output by the face recognition network in the teacher network to obtain a bias feature map, and the bias feature map satisfies the relational expression: where, R t is the significant feature map output by the face recognition network in the teacher network, C is the bias matrix, and the size of the bias matrix is the same as that of the significant feature map. is the bias feature map; in the teacher network, the bias feature map is input into the classification layer to obtain the output result of the teacher network.
6. The face recognition method based on artificial intelligence according to claim 4, wherein The preset loss function value satisfies the relational expression: wherein, are the values of the i-th row in the first teacher output and the second teacher output respectively, are the values of the i-th row in the first student output and the second student output respectively, N is the number of rows of the first student output, the first teacher output, the second student output and the second teacher output, and Loss is the value of the preset loss function.
7. The face recognition method based on artificial intelligence according to claim 1, wherein The number of the reference images is at least one, and each reference image corresponds to a reference attention map and a reference significant feature map. Obtaining the face recognition result of the face image to be recognized based on the attention map to be recognized, the reference attention map, the significant feature map to be recognized, the reference significant feature map and the face recognition label includes: Calculate the similarity between the attention map to be recognized and the reference attention map as a first similarity; Calculate the similarity between the significant feature map to be recognized and the reference significant feature map as a second similarity; Perform weighted summation on the first similarity and the second similarity to obtain the target similarity between the face image to be recognized and each reference image, and obtain the maximum similarity of all target similarities; Compare the maximum similarity with a preset pre-threshold. If the maximum similarity is greater than the preset pre-threshold, use the face recognition label of the reference image corresponding to the maximum similarity as the face recognition result of the face image to be recognized; if the maximum similarity is not greater than the preset pre-threshold, issue an alarm in a preset manner.
8. A face recognition device based on artificial intelligence, characterized in that, The device includes: An acquisition unit, configured to acquire face images of the same face in different states to obtain an image subset of the face, and store all image subsets as a training image set; A construction unit, configured to construct a face recognition network; A training unit for training the face recognition network based on the training image set to obtain a target face recognition network, wherein the input of the target recognition network is a face image, and the output is an attention map and a salient feature map of the face image, and the salient feature map includes common features of the same face in different states; An input unit for inputting the face image to be recognized into the target face recognition network to output an attention map to be recognized and a salient feature map to be recognized, and inputting a reference image into the target face recognition network to output a reference attention map and a reference salient feature map, wherein the reference image includes a face recognition label; A face recognition unit for obtaining a face recognition result of the face image to be recognized based on the attention map to be recognized, the reference attention map, the salient feature map to be recognized, the reference salient feature map and the face recognition label.
9. An electronic device, characterized in that, The electronic device includes: A memory storing computer-readable instructions; and A processor for executing the computer-readable instructions stored in the memory to implement the artificial intelligence-based face recognition method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, Computer-readable instructions are stored on the computer-readable storage medium, and when the computer-readable instructions are executed by a processor, the artificial intelligence-based face recognition method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Face feature extraction method and device, equipment and storage medium
CN114783019A
Systems and methods for secure face authentication
US20220414198A1