Face recognition model training method, system, device and storage medium
By constructing a convolutional neural network with multiple branches and training it with occluded and unoccluded face images, the problem of face recognition under occlusion conditions such as wearing masks was solved, and efficient recognition was achieved under different conditions.
Patent Information
- Application Number
- CN202110055204.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-01-15
- Publication Date
- 2025-11-21
- Estimated Expiration
- 2041-01-15
AI Technical Summary
Existing facial recognition technology has poor recognition rates when the face is covered by masks or other obstructions, and is prone to false recognition and rejection.
A convolutional neural network with at least two branches is used, with each branch connected to a loss function. The network is trained using sets of occluded and unoccluded face images. Backpropagation is performed by weighted summation of the loss functions to optimize the network until an optimization stopping threshold is reached.
It can effectively recognize faces in both occluded and unoccluded conditions, improving the recognition rate and making full use of the advantages of each network branch to enhance feature recognition capabilities.
Smart Images

Figure CN114764938B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of artificial intelligence, and in particular to a face recognition model training method and system, device and storage medium. BACKGROUND
[0002] In recent years, due to the rapid development of deep neural network technology, the face recognition technology based on it has also matured rapidly, which is specifically manifested in large-scale industrial applications. Although the existing face recognition technologies have shown excellent recognition rates, their generalization ability is still not good. Especially in the case of wearing a mask, the existing technologies are prone to misrecognition and rejection. SUMMARY
[0003] Therefore, it is necessary to provide a face recognition model training method for the case of incomplete face image presentation.
[0004] In order to achieve the purpose of the present application, the following technical solutions are adopted:
[0005] A face recognition model training method comprises:
[0006] obtaining an image set with occluded face images as a first training set and an image set without occluded face images as a second training set;
[0007] constructing a convolutional neural network with at least two network branches, and connecting a loss function to each network branch respectively;
[0008] inputting the first training set and the second training set into the convolutional neural network for training, and performing back propagation according to the weighted sum of all the loss functions to optimize the convolutional neural network until the weighted sum is less than a preset optimization stop threshold, and stopping optimization, and taking the optimized convolutional neural network at this time as a face recognition model.
[0009] A face recognition model training system comprises:
[0010] a training set acquisition module configured to obtain an image set with occluded face images as a first training set and an image set without occluded face images as a second training set;
[0011] a neural network construction module configured to construct a convolutional neural network with at least two network branches, and connect a loss function to each network branch respectively;
[0012] The training module is used to input the first training set and the second training set into the convolutional neural network for training, and to perform backpropagation based on the weighted sum of all the loss functions to optimize the convolutional neural network until the weighted sum is less than a preset optimization stopping threshold, and then use the optimized convolutional neural network as a face recognition model.
[0013] A face recognition model training device includes a memory, a processor, and a face recognition model training program stored in the memory and executable on the processor. When the face recognition model training program is executed by the processor, it implements the steps of the face recognition model training method described above.
[0014] A computer-readable storage medium storing a face recognition model training program, wherein the face recognition model training program, when executed by a processor, implements the steps of the face recognition model training method described above.
[0015] The aforementioned face recognition model training method, system, device, and storage medium employ at least two network branches, each connected to a loss function. During training, each network branch independently calculates its own loss, ultimately obtaining a weighted sum of losses to guide training. This fully leverages the characteristics of the different data processed by each network branch or the inherent advantages of each branch itself, significantly enhancing the features of key regions in the face. When the face is occluded, enhanced features from the unoccluded region are used to achieve face recognition; that is, identical features in both the occluded and unoccluded training sets are amplified. When the face is unoccluded, remaining features outside the key regions are combined to further enhance recognition, thus enabling recognition in both occluded and unoccluded situations. Attached Figure Description
[0016] Figure 1 This is a schematic diagram of a face recognition model training device in one embodiment;
[0017] Figure 2a This is a flowchart of a face recognition model training method in one embodiment;
[0018] Figure 2b This is a schematic diagram of a convolutional neural network structure in one embodiment;
[0019] Figure 3a for Figure 2a Flowchart of one implementation method of step S204;
[0020] Figure 3b This is a schematic diagram of a neural network structure with two branches.
[0021] Figure 4a for Figure 2aAnother embodiment method flow chart of step S204;
[0022] Figure 4b An embodiment full connection layer structure schematic diagram;
[0023] Figure 5 An embodiment full connection layer structure schematic diagram; Figure 2a One of the embodiment method flow charts of step S206;
[0024] Figure 6a An embodiment full connection layer structure schematic diagram; Figure 2a One of the embodiment method flow charts of step S202;
[0025] Figure 6b The processing result of each step; Figure 6a The processing result of each step;
[0026] Figure 7 An embodiment face recognition model training system module diagram. DETAILED DESCRIPTION
[0027] In order to facilitate the understanding of the present application, the present application will be described more fully below with reference to the accompanying drawings. The preferred embodiments of the present application are shown in the drawings. However, the present application can be realized in many different forms and is not limited to the embodiments described herein. On the contrary, these embodiments are provided so that the disclosure of the present application is more thorough and comprehensive.
[0028] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the present application belongs. The terminology used in the specification of the present application is only for the purpose of describing specific embodiments and is not intended to limit the present application.
[0029] Figure 1 The face recognition model training device 100 structure schematic diagram of the hardware running environment involved in the embodiment scheme of the present application.
[0030] The face recognition model training device of the embodiment of the present application, that is, the face recognition model training is carried out on the device carrying the data processing chip with data processing function, and the trained face recognition model is deployed in the corresponding intelligent device. The intelligent device can be, for example, a smart air conditioner, a smart lamp, a smart power supply, an AR / VR device, a smart sound box, an automatic driving car, a PC, a smart phone, a tablet computer, an electronic book reader, a portable computer, etc. As long as it has certain general data processing capability.
[0031] As shown in Figure 1 The face recognition model training device 100 includes a memory 104, a processor 102 and a network interface 106.
[0032] The processor 102 may, in some embodiments, be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chip, for running program codes stored in the memory 104 or processing data, such as executing a face recognition model training program, etc.
[0033] The memory 104 includes at least one type of readable storage medium, including a flash memory, a hard disk, a multimedia card, a card-type memory (e.g., an SD or DX memory, etc.), a magnetic memory, a magnetic disk, an optical disk, etc. The memory 104 may, in some embodiments, be an internal storage unit of the face recognition model training device 100, such as a hard disk of the face recognition model training device 100. The memory 104 may, in other embodiments, also be an external storage device of the face recognition model training device 100, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the face recognition model training device 100.
[0034] Further, the memory 104 can also include an internal storage unit of the face recognition model training device 100. The memory 104 can be used not only to store application software installed on the face recognition model training device 100 and various data, such as codes for face recognition model training, etc., but also to temporarily store data that has been output or will be output.
[0035] The network interface 106 can optionally include a standard wired interface, a wireless interface (e.g., a WI-FI interface), etc., and is generally used to establish a communication connection between the face recognition model training device 100 and other electronic devices.
[0036] The network can be the Internet, a cloud network, a wireless fidelity (Wi-Fi) network, a personal area network (PAN), a local area network (LAN), and / or a metropolitan area network (MAN). Various devices in the network environment can be configured to connect to the communication network according to various wired and wireless communication protocols. Examples of such wired and wireless communication protocols can include, but are not limited to, at least one of the following: transmission control protocol and internet protocol (TCP / IP), user datagram protocol (UDP), hypertext transfer protocol (HTTP), file transfer protocol (FTP), ZigBee, EDGE, IEEE 802.11, light fidelity (Li-Fi), 802.16, IEEE 802.11s, IEEE 802.11g, multi-hop communication, wireless access point (AP), device-to-device communication, cellular communication protocol, and / or BlueTooth communication protocol, or a combination thereof.
[0037] Figure 1 Only the face recognition model training device 100 with components 102-106 is shown, and those skilled in the art can understand that, Figure 1 The structure shown does not constitute a limitation on the face recognition model training device 100, and can include fewer or more components than shown, or combine certain components, or different component arrangements.
[0038] As Figure 2a shown, the face recognition model training method of an embodiment of the present application includes the following steps:
[0039] Step S202: Obtain an image set with occluded face images as a first training set, and an image set with non-occluded face images as a second training set.
[0040] An occluded face image refers to an image in which all features of a face are not fully presented in the captured image, such as when wearing sunglasses or a mask. Other situations where the face is partially occluded due to the shooting environment also belong to occluded face images. A non-occluded face image is an image in which all features of a face are fully presented. In most shooting portrait scenes, it is a non-occluded face image.
[0041] In the present application, the first training set includes images that occlude the same area in the face, for example, the first training set is images of people wearing masks, or images of people wearing sunglasses. In addition, the occluded face images in the first training set can also be obtained by technically processing a non-occluded face image to occlude a part of it. The technical processing includes erasing the part to be occluded, or adding an occlusion object such as sunglasses, a mask, etc. to the image.
[0042] Step S204: Construct a convolutional neural network with at least two network branches, and connect a loss function to each network branch respectively.
[0043] Referring to Figure 2b , a convolutional neural network 200 with at least two network branches is constructed. It includes an input layer 202, a base network layer 204, a first network branch 206, a second network branch 208, …, an Nth network branch 210. A first loss function L1 connected to the first network branch 206, a second loss function L2 connected to the second network branch 208, and an Nth loss function LN connected to the Nth network branch 210.
[0044] The first network branch 206 outputs a first face feature 212, the second network branch 208 outputs a second face feature 214, and the Nth network branch 210 outputs an Nth face feature 216. The first face feature 212, the second face feature 214, and the Nth face feature 216 are finally fused into a fused feature 218. The fused feature 218 is used to output a face recognition result when the face recognition model is applied after being trained. In this application, the training process of the face recognition model will be mainly discussed.
[0045] The input layer 202 is used to receive an input image. Generally, the input layer 202 has a fixed size, for example, 128*128. When the input image is input, the size of the input image must be the same as the input layer 202.
[0046] The base network layer 204 includes at least one convolutional layer. Each convolutional layer shrinks in the length and width directions and increases in the channel dimension when processing the result of the previous layer (the input image of the input layer 202 is processed by the first convolutional layer). For example, in the MobileNet network, the input layer input is 224*224*3, that is, a 3-channel image with a resolution of 224*224, and after the first layer of convolution processing, a feature map of 112*112*32 is obtained. In the convolutional neural network, each convolutional layer is generally followed by batch processing (Batch Normalization, BN) and activation processing (such as ReLU function activation), which will not be described here, and mainly the convolutional layer will be described.
[0047] On the basis of the same base network layer 204, multiple network branches are respectively extended and connected. That is, the first network branch 206, the second network branch 208, …, and the Nth network branch 210 all take the output feature map of the base network layer 204 as input and respectively process them. In this application, the first network branch 206, the second network branch 208, …, and the Nth network branch 210 respectively process the features of different parts of the same feature map. For example, the first network branch 206 processes the feature map in a predetermined region, for example, the predetermined region is the upper feature or the lower feature of the same feature map, and the second network branch 208 processes all the features of the same feature map. Each network branch also includes at least one convolutional layer, and outputs a respective face feature vector through a respective fully connected layer. The multiple network branches can be homogeneous networks or heterogeneous networks. For example, the first network branch 206 can be a residual network, and the second network branch 208 can be a non-residual network.
[0048] The loss function is used to evaluate the difference between the face features output by the network and the actual face features, and whether the network parameters are reasonable. The greater the difference, the more unreasonable the network parameters, and further optimization is needed. Face recognition includes recognizing images as faces and recognizing the identity of the face. In the training set, the face image can contain label information of the face category and the identity. Therefore, the loss function of the present application can be a classification loss function or a contrast loss function.
[0049] Step S206: input the first training set and the second training set into the convolutional neural network respectively for training, and perform back propagation according to the weighted sum of all the loss functions to optimize the convolutional neural network until the weighted sum is less than a preset optimization stop threshold, and the convolutional neural network optimized at this time is taken as a face recognition model.
[0050] The images in the first training set and the second training set are input into the convolutional neural network 200 in turn, and the loss of each image in the training set can be calculated after being processed by the convolutional neural network 200. Therefore, the weighted sum of all the loss functions is:
[0051] L = λ1L1+ λ2L2+ … + λ N L N
[0052] wherein L1, L2, … L N are the loss functions of the first network branch 206, the second network branch 208, …, and the Nth network branch 210, respectively, and λ1, λ2, … λ N are the corresponding weights.
[0053] When training each input image, the weighted sum L of all the loss functions is used for back propagation to optimize the convolutional neural network. After a sufficient number of training, the loss weighted sum L is small to a certain extent (for example, 10 -6 ), the model training is completed, and the required face recognition model is obtained.
[0054] During the training process, forward propagation is to calculate the features of the lower layer layer by layer according to the initial parameters in the network. Back propagation is to adjust the parameters step by step in the reverse direction from the end of the network to minimize the loss under the current input. Then, according to the updated parameters, forward propagation is performed again, and the cycle is repeated until the optimized parameters make the loss weighted sum L of the input image small to a certain extent, i.e., less than a preset optimization stop threshold, for example, 1 × 10 -6 .
[0055] In the above embodiment, since at least two network branches are adopted and each network branch is connected with a loss function, each network branch independently calculates its own loss during training, and finally a weighted sum of the losses is obtained to guide the training. The characteristics of different data processed by each network branch or the advantages of the network branch itself can be fully utilized to fully enhance the features of the key areas in the face, and in the case of face occlusion, the enhanced features of the unoccluded areas are used to realize face recognition, while in the case of unoccluded face, the remaining features outside the key areas are combined to strengthen the recognition, so that the face can be recognized in both face occlusion and unocclusion.
[0056] As shown in the figure, in one embodiment, step S204 specifically adopts two network branches, which include: Figure 3a
[0057] Step S302: constructing a complete convolutional neural network; the complete convolutional neural network includes an input layer, an intermediate layer, a fully connected layer, and an output layer.
[0058] This step can construct a complete convolutional neural network by using a conventional method, and details are not described herein. A ready-made convolutional neural network can also be used.
[0059] Step S304: constructing a first network branch based on any intermediate layer of the complete convolutional neural network and a convolutional layer before the any intermediate layer as a base network layer; the remaining part extending from the same intermediate layer in the complete convolutional neural network is defined as a second network branch.
[0060] Generally, each convolutional neural network includes a certain number of convolutional layers, from several layers to tens of layers, hundreds of layers, or even thousands of layers. Referring to Figure 3b A first network branch 304 is constructed based on any intermediate layer 302 of the convolutional neural network and a convolutional layer before the any intermediate layer 302 as a base network layer 308. The any intermediate layer 302 can be the 2nd layer, the 3rd layer, …, or the (N-1)th layer. In a preferred embodiment, the intermediate layer at the branch can be selected from a layer relatively in the middle, such as a layer before and after (N-1) / 2. Wherein, N is the total number of layers of the complete convolutional neural network. The remaining part extending from the same intermediate layer 302 is defined as a second network branch 306. The intermediate layer 302 and the convolutional layer before the intermediate layer are the base network layer 308, which is shared by the two network branches. The base network layer 308 and the second network branch 306 belong to a complete convolutional neural network, so the second network branch 306 inputs all the features of the intermediate layer 302. The first network branch 304 inputs part of the features of the intermediate layer 302.
[0061] In this embodiment, the first network branch is extended based on the complete convolutional neural network or even the ready-made convolutional neural network, so that the convolutional neural network can be quickly constructed and the efficiency of establishing the model can be improved.
[0062] As shown in FIG. 4, in one embodiment, step S204 specifically adopts two network branches, which can include:
[0063] Step S402: starting from the input layer, a base network layer is constructed. The base network layer includes at least one convolutional layer, which is used to perform at least one layer of processing on the input image data to obtain a feature map. When constructing the base network layer, convolutional layers are added layer by layer, and the input image of each convolutional layer is the output image of the previous convolutional layer, and the size of the output image gradually decreases and the number of channels gradually increases after each convolutional processing. It should be noted that each convolutional layer can further include batch normalization (BN) processing and activation processing (for example, using a ReLU activation function).
[0064] Step S404: based on the base network layer, a first network branch and a second network branch are respectively constructed.
[0065] In combination Figure 3b , the difference between the present embodiment and the previous embodiment lies in the construction process, that is, the base network layer 308 is first constructed, and then the first network branch 304 and the second network branch 306 are respectively constructed. The advantage is that different structures of network branches can be selected, which is not limited by existing neural networks.
[0066] In the above embodiments, the first network branch 304 takes part of the features of the feature map output by the base network layer 308 as input, and the second network branch 306 takes all the features of the feature map output by the base network layer 308 as input. Specifically, in different scenarios, the part of the features will be different. For the scenario of identifying wearing a mask, the part of the features is the upper half of the feature map of the middle layer 302, and the input of the first network branch 304 can be obtained by cutting off 50% of the upper half. For the scenario of identifying wearing glasses, the part of the features is the remaining part excluding about 20% distance from the top and occupying about 20% of the band-shaped part. According to the application requirements, each feature map is input or output with a preset size, such as 224x224, 112x112 pixel number, etc. The pixels in the feature map can be saved in a matrix, for example, the feature map of 224x224 can be represented as:
[0067] a[0,0], a[0,1], a[0,2], …, a[0,223],
[0068] a[1,0], a[1,1], a[1,2], …, a[1,223],
[0069] …
[0070] a[112,0], a[112,1], a[112,2], …, a[112,223],
[0071] a[113,0], a[113,1], a[113,2], …, a[113,223],
[0072] …
[0073] a[223,0], a[223,1], a[223,2], …, a[223,223].
[0074] When 50% of the lower half needs to be shielded, the values of the lower half (such as the following pixels) in the above feature map can be set to 0 or 255.
[0075] a[112,0], a[112,1], a[112,2], …, a[112,223],
[0076] a[113,0], a[113,1], a[113,2], …, a[113,223],
[0077] …
[0078] a[223,0], a[223,1], a[223,2], …, a[223,223].
[0079] Since the first training set and the second training set are both input into the neural network of the above embodiment for training, and the training target is to be able to simultaneously recognize the occluded face image and the non-occluded face image, the feature map after removing the occluded area is input into the first network branch 304, which can simultaneously use the occluded face image and the non-occluded face image data to enhance the ability of the first network branch 304 to recognize the non-occluded area of the face; and the non-occluded face image can train the second network branch 306 to recognize the ability of the feature of the whole face. When calculating the weighted sum, by adjusting the appropriate weight, the face recognition model can reliably balance the ability to recognize part of the face and the whole face, and finally achieve reliable recognition of the face in all scenes.
[0080] In the above embodiment, the first network branch 304 and the second network branch 306 output feature vectors of the same dimension through respective fully connected layers. Since the first network branch 304 only inputs part of the features of the intermediate layer 302, the final feature quantity will be relatively small. In order to realize feature fusion, the dimension of the output feature of the first network branch 304 needs to be expanded, such as shown in Figure 4b , and an additional fully connected layer is added, which has the same dimension as the second network branch 306.
[0081] In the above embodiments, the feature vector is also subjected to length normalization. This is to normalize the data and reduce the complexity of processing the data.
[0082] As shown in Figure 5 Fig. 2, still taking two network branches as an example, step S206 includes:
[0083] Step S502: input a set proportion of the first training set and the second training set into the convolutional neural network respectively.
[0084] That is, the number of occluded face images and the number of non-occluded face images in the images input into the convolutional neural network have a set proportion. This proportion can enhance the ability to recognize faces through non-occluded areas, and will not weaken the recognition ability of the whole face due to the input of occluded area images. The recommended value of the proportion is 1:1, of course, it can also be adjusted according to the situation of the training set and the feedback of the training process to obtain a better proportion.
[0085] Step S504: calculate the loss of the first network branch and the loss of the second network branch respectively.
[0086] The loss is calculated according to the loss function connected to the first network branch and the second network branch respectively. The loss function can be a classification loss function or a contrast loss function.
[0087] Step S506: calculate the weighted sum of the loss of the first network branch and the loss of the second network branch.
[0088] The loss weighted sum is calculated according to the formula L = λ1L1 + λ2L2.
[0089] Step S508: back-propagate the weighted sum to update the parameters of the face recognition model.
[0090] Back-propagation is a method of using local minimum to find the parameter combination that makes the loss function minimum under the current input. Generally, the partial derivative of the loss function with respect to each parameter is taken and made to be 0 under the current input to obtain the parameter, so as to realize the one-by-one update of the parameter. The loss weighted sum will affect the parameters of the first network branch and the second network branch.
[0091] As shown in Figure 6a Fig. 2, step S202 can use existing non-occluded face images to perform data augmentation to obtain an image set with occluded face images. Specifically, it can include:
[0092] Step S602: obtain existing non-occluded face images.
[0093] Specifically, a public face-labeled image dataset can be acquired. For example, IMDB-WIKI, CelebA, CASIA-WebFace, etc. Of course, it can also be collected and labeled by itself.
[0094] Step S604: performing face detection and face key point detection on the existing unoccluded face image.
[0095] Based on the shooting angle, position or other reasons, the unoccluded face image may not be a standard image. For example, Figure 6b As shown in the figure, the unoccluded face image is subjected to face detection and face key point detection. Among them, face detection identifies the face region in the image, and then the irrelevant region can be removed. Face key point detection identifies the key points in the face image, such as eyes, nose, and corners of the mouth.
[0096] Step S606: using the face key point detection result to perform face affine transformation alignment to obtain a standard size face image.
[0097] According to the detection result of the key points, the angle at which the face image is placed and the perspective direction can be understood, so that the image can be rotated, enlarged and affine transformed accordingly to obtain a complete front standard image. Reduce the noise brought by the image.
[0098] Step S608: erasing a set region of the standard size face image.
[0099] The pixels in the set region can be replaced with pixels of the same color, such as black pixels. The size of the entire image is still standard size to match the input layer of the convolutional neural network. For example, the image of wearing a mask is obtained by erasing the lower half of the image. At this point, the image simulating the partially occluded face can be obtained.
[0100] Since there is no unified collection and labeling standard dataset for occluded face data, the existing unoccluded face image is used to form an occluded face image, which can reduce the work of data collection and sorting. The processed occluded face image can also well simulate the occlusion situation.
[0101] As shown in the figure, Figure 7 It is a face recognition model training system module diagram. The face recognition model training system comprises:
[0102] The training set acquisition module 702 is used to acquire an image set with occluded face images as a first training set, and an image set with unoccluded face images as a second training set.
[0103] The neural network construction module 704 is used to construct a convolutional neural network with at least two network branches, and connect a loss function to each network branch respectively.
[0104] a training module 706, configured to input the first training set and the second training set into the convolutional neural network respectively for training, and perform back propagation according to a weighted sum of all the loss functions to optimize the convolutional neural network until the weighted sum is less than a preset optimization stop threshold, and stop the optimization, and take the convolutional neural network optimized at this time as the face recognition model.
[0105] The system corresponds to the method flow described in FIG. 2, and details are not described herein.
[0106] In addition, the embodiment of the present application further provides a computer readable storage medium, and the computer readable storage medium stores the face recognition model training program described above, and the face recognition model training program is executed by a processor to realize the steps of the face recognition model training method described above.
[0107] The computer readable storage medium of the present application has basically the same implementation as the face recognition model training method described above, and details are not described herein.
[0108] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can adopt a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can adopt a computer program product in the form of being implemented on one or more computer readable storage media containing computer usable program codes, including but not limited to disk memory, CD-ROM, optical memory, etc.
[0109] The present application is described with reference to flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of the flows and / or blocks in the flowcharts and / or block diagrams can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a means for implementing the functions specified in the flowcharts and / or block diagrams. Figure One The functions specified in one flow or multiple flows and / or blocks Figure One The means for implementing the functions specified in one flow or multiple flows and / or blocks.
[0110] These computer program instructions can also be stored in a computer readable memory capable of guiding the computer or other programmable data processing devices to work in a specific way, so that the instructions stored in the computer readable memory produce a product including instruction means, which implements the functions specified in the flowcharts and / or block diagrams. Figure Oneone or more processes and / or blocks Figure One the function specified in the one or more blocks.
[0111] These computer program instructions can also be loaded into computer or other programmable data processing devices, so that a series of operation steps are performed on the computer or other programmable data processing devices to generate computer-implemented processes, thus the instructions executed on the computer or other programmable data processing devices provide processes for implementing the flow Figure One one or more processes and / or blocks Figure One Figure One the function specified in the one or more blocks.
[0112] While the preferred embodiments of the application have been described, additional variations and modifications can be made to the preferred embodiments by those skilled in the art once they learn of the basic inventive concepts. Therefore, the appended claims are intended to cover all such additional variations and modifications as fall within the scope of the application. What is claimed is:
[0113] Obviously, numerous modifications and variations of the present application are possible in light of the above teachings. It is therefore to be understood that within the scope of the appended claims and their equivalents, the application can be practiced otherwise than as specifically described.
Claims
1. A method for training a face recognition model, comprising: The first training set is obtained by acquiring an image set containing occluded face images, and the second training set is obtained by acquiring an image set containing unoccluded face images. Construct a convolutional neural network with at least two network branches, and connect a loss function to each network branch separately; The first training set and the second training set are respectively input into the convolutional neural network for training, and backpropagation is performed according to the weighted sum of all the loss functions to optimize the convolutional neural network until the weighted sum is less than a preset optimization stopping threshold, and the optimized convolutional neural network is used as the face recognition model. The construction of a convolutional neural network with at least two network branches includes, as a method for constructing a convolutional neural network structure with two network branches: Construct a complete convolutional neural network; the complete convolutional neural network includes an input layer, intermediate layers, fully connected layers, and an output layer; A first network branch is constructed using any intermediate layer of the complete convolutional neural network and the convolutional layer preceding that intermediate layer as the base network layer; the remaining portion extending from the same intermediate layer in the complete convolutional neural network is defined as the second network branch; or The construction of a convolutional neural network with at least two network branches includes, as a method for constructing a convolutional neural network structure with two network branches: Starting from the input layer, a base network layer is constructed; the base network layer includes at least one convolutional layer for processing the input image data at least once. Based on the basic network layer, the first network branch and the second network branch are constructed respectively; The first network branch takes a portion of the feature map output by the base network layer at a preset ratio as input, while the second network branch takes all the features of the feature map output by the base network layer as input.
2. The face recognition model training method according to claim 1, characterized in that, The first and second network branches output feature vectors of the same dimension through their respective fully connected layers.
3. The face recognition model training method according to claim 2, characterized in that, Also includes: The feature vector is normalized by its magnitude.
4. The face recognition model training method according to claim 1, characterized in that, The step of inputting the first training set and the second training set into the convolutional neural network for training includes: The first and second training sets, with predetermined proportions, are respectively input into the convolutional neural network; Calculate the loss of the first network branch and the loss of the second network branch respectively; Calculate the weighted sum of the loss of the first network branch and the loss of the second network branch; The weighted sum is backpropagated to update the parameters of the face recognition model.
5. The face recognition model training method according to claim 1, characterized in that, The acquisition of the image set containing occluded faces as the first training set includes: Data augmentation is performed using existing unoccluded face images to obtain an image set with occluded face images.
6. The face recognition model training method according to claim 5, characterized in that, The data augmentation using existing unobstructed facial images includes: Obtain existing unobstructed face images; Face detection and facial landmark detection are performed on the existing unobstructed face images; Affine transformation alignment of the face is performed using the results of facial landmark detection to obtain a standard-sized face image; The specified area of the standard-sized face image is erased.
7. A face recognition model training system, characterized in that, include: The training set acquisition module is used to acquire an image set with occluded face images as the first training set and an image set with unoccluded face images as the second training set. The neural network building block is used to build convolutional neural networks with at least two network branches and connect a loss function to each network branch. The training module is used to input the first training set and the second training set into the convolutional neural network for training, and to perform backpropagation based on the weighted sum of all the loss functions to optimize the convolutional neural network until the weighted sum is less than a preset optimization stopping threshold, and to stop the optimization, and use the optimized convolutional neural network at this time as a face recognition model. The construction of a convolutional neural network with at least two network branches includes, as a method for constructing a convolutional neural network structure with two network branches: Construct a complete convolutional neural network; the complete convolutional neural network includes an input layer, intermediate layers, fully connected layers, and an output layer; A first network branch is constructed using any intermediate layer of the complete convolutional neural network and the convolutional layer preceding that intermediate layer as the base network layer; the remaining portion extending from the same intermediate layer in the complete convolutional neural network is defined as the second network branch; or The construction of a convolutional neural network with at least two network branches includes, as a method for constructing a convolutional neural network structure with two network branches: Starting from the input layer, a base network layer is constructed; the base network layer includes at least one convolutional layer for processing the input image data at least once. Based on the basic network layer, the first network branch and the second network branch are constructed respectively; The first network branch takes a portion of the feature map output by the base network layer at a preset ratio as input, while the second network branch takes all the features of the feature map output by the base network layer as input.
8. A face recognition model training device, characterized in that, The system includes a memory, a processor, and a face recognition model training program stored in the memory and executable on the processor. When executed by the processor, the face recognition model training program implements the steps of the face recognition model training method as described in any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a face recognition model training program, which, when executed by a processor, implements the steps of the face recognition model training method as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Neural network training method, device and system, and storage medium
CN108875934A
Face detection and recognition method and device based on face key point correction
CN109800648A