Face recognition method, device, electronic device and computer-readable storage medium

By constructing a parallel replica network group and training the face recognition model in sample groups, the problem of parameter update failure and the failure to fit the optimal solution caused by only one prototype of each class in the prior art is solved, and the accuracy of face recognition is improved.

CN114140845BActive Publication Date: 2025-07-11SHENZHEN XUMI YUNTU SPACE TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111373975.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-19
Publication Date
2025-07-11
Estimated Expiration
2041-11-19

AI Technical Summary

Technical Problem

现有技术中,每个类只有一个原型的人脸识别模型在训练过程中无法兼容多个样本的多样性,导致参数更新失败,无法拟合到最优解。

Method used

Build the main network, sliding replica network, multi-convolution replica network, multi-activation replica network and multi-code replica network, which are divided into the first sample group and the second sample group, and train the parallel replica network group, generate multiple prototypes, and train the face recognition model.

Benefits of technology

The accuracy of the face recognition model is improved, and the problem of failed parameter updates and inability to fit the optimal solution is solved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114140845B_ABST
    Figure CN114140845B_ABST
Patent Text Reader

Abstract

The present disclosure relates to the field of artificial intelligence technology, and provides a face recognition method, apparatus, electronic device, and computer-readable storage medium. The method includes: dividing multiple samples in each class in the input image into a first sample group and a second sample group; respectively training a multi-convolution replication network, a multi-activation replication network, and a multi-encoding replication network using the first sample group of each class; inputting the second sample group of each class into a parallel replication network group to output multiple generated prototypes of each class; training a face recognition model according to the first sample group, the second sample group, and the multiple generated prototypes of each class, and performing face recognition using the trained neural network model. By adopting the above technical means, in the prior art, in the training of a face recognition model, there is only one prototype for each class, which may cause the parameter update of the face recognition model to fail and the face recognition model cannot be fitted to the optimal solution.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technology, and in particular to a face recognition method, device, electronic device and computer-readable storage medium. Background Art

[0002] In existing face recognition algorithms, each class has only one prototype. When training a face recognition model, the entire model can only be updated or optimized by maintaining one prototype for each class. The update of the face recognition model essentially relies on only one prototype for each class. However, since each class has only one prototype, it cannot accommodate the diversity among multiple samples in each class, which will lead to the failure of updating the parameters of the face recognition model and the failure of fitting the face recognition model to the optimal solution.

[0003] In the process of implementing the concept of the present disclosure, the inventors found that there are at least the following technical problems in the related technology: in the training of the face recognition model, each class has only one prototype, which will lead to the failure of parameter update of the face recognition model and the problem that the face recognition model cannot fit the optimal solution. Summary of the invention

[0004] In view of this, the embodiments of the present disclosure provide a face recognition method, device, electronic device and computer-readable storage medium to solve the problem in the prior art that, in the training of the face recognition model, each class has only one prototype, which will cause the parameter update of the face recognition model to fail and the face recognition model cannot be fitted to the optimal solution.

[0005] According to a first aspect of an embodiment of the present disclosure, a face recognition method is provided, including: constructing a main network, a sliding replica network, a multi-convolution replica network, a multi-activation replica network and a multi-coding replica network; using the sliding replica network, the multi-convolution replica network, the multi-activation replica network and the multi-coding replica network to form a parallel replica network group, and using the main network and the parallel replica network group to form a face recognition model; obtaining an input image, and dividing multiple samples in each class in the input image into a first sample group and a second sample group; using the first sample group of each class to respectively train the multi-convolution replica network, the multi-activation replica network and the multi-coding replica network; inputting the second sample group of each class into the parallel replica network group, and outputting multiple generated prototypes of each class; training the face recognition model according to the first sample group, the second sample group and the multiple generated prototypes of each class, and using the trained neural network model to perform face recognition.

[0006] In a second aspect of the embodiments of the present disclosure, a face recognition device is provided, including: a construction module configured to construct a main network, a sliding replication network, a multi-convolution replication network, a multi-activation replication network, and a multi-encoding replication network; a composition module configured to use the sliding replication network, the multi-convolution replication network, the multi-activation replication network, and the multi-encoding replication network to form a parallel replication network group, and use the main network and the parallel replication network group to form a face recognition model; an acquisition module configured to acquire an input image and divide multiple samples in each class in the input image into a first sample group and a second sample group; a first training module configured to use the first sample group of each class to train the multi-convolution replication network, the multi-activation replication network, and the multi-encoding replication network respectively; a network group module configured to input the second sample group of each class into the parallel replication network group and output multiple generated prototypes of each class; a second training module configured to train the face recognition model according to the first sample group, the second sample group, and the multiple generated prototypes of each class, and perform face recognition using the trained neural network model.

[0007] In a third aspect of the embodiments of the present disclosure, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the above method are implemented.

[0008] In a fourth aspect of the embodiments of the present disclosure, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above method are implemented.

[0009] The beneficial effects of the embodiments of the present disclosure compared with the prior art are as follows: Since the embodiments of the present disclosure divide multiple samples in each class in the input image into a first sample group and a second sample group; use the first sample group of each class to train the multi-convolution replication network, the multi-activation replication network, and the multi-encoding replication network respectively; input the second sample group of each class into the parallel replication network group and output multiple generated prototypes of each class; train the face recognition model according to the first sample group, the second sample group, and the multiple generated prototypes of each class, and perform face recognition using the trained neural network model. Therefore, by adopting the above technical means, the problem in the prior art that in the training of a face recognition model, there is only one prototype for each class, which may cause the parameter update of the face recognition model to fail and the face recognition model cannot be fitted to the optimal solution can be solved, and further the accuracy of the face recognition model in recognizing faces can be improved. Description of the Drawings

[0010] To more clearly illustrate the technical solutions in the embodiments of the present disclosure, the following will briefly introduce the accompanying drawings required for use in the embodiments or the description of the prior art. Obviously, the accompanying drawings in the following description are only some embodiments of the present disclosure. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on these drawings.

[0011] Figure 1 is a schematic diagram of the application scenario of the embodiments of the present disclosure;

[0012] Figure 2 is a schematic flowchart of a face recognition method provided by the embodiments of the present disclosure;

[0013] Figure 3 is a schematic structural diagram of a face recognition device provided by the embodiments of the present disclosure;

[0014] Figure 4 is a schematic structural diagram of an electronic device provided by the embodiments of the present disclosure. Specific Embodiments

[0015] In the following description, specific details such as specific system structures and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the embodiments of the present disclosure. However, those skilled in the art should clearly understand that the present disclosure can also be implemented in other embodiments without these specific details. In other cases, the detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present disclosure.

[0016] A face recognition method and device according to the embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings.

[0017] Figure 1 is a schematic diagram of the application scenario of the embodiments of the present disclosure. The application scenario may include terminal devices 1, 2, and 3, a server 4, and a network 5.

[0018] The terminal devices 1, 2, and 3 can be hardware or software. When the terminal devices 1, 2, and 3 are hardware, they can be various electronic devices with a display screen and supporting communication with the server 4, including but not limited to smartphones, tablets, laptop computers, and desktop computers, etc.; when the terminal devices 1, 2, and 3 are software, they can be installed in the above-mentioned electronic devices. The terminal devices 1, 2, and 3 can be implemented as multiple software or software modules, or can be implemented as a single software or software module, and the embodiments of the present disclosure do not limit this. Further, various applications can be installed on the terminal devices 1, 2, and 3, such as data processing applications, instant messaging tools, social platform software, search applications, shopping applications, etc.

[0019] The server 4 can be a server that provides various services. For example, it can be a background server that receives requests sent by terminal devices with which it establishes a communication connection. This background server can receive and analyze requests sent by terminal devices and generate processing results. The server 4 can be a single server, a server cluster composed of several servers, or a cloud computing service center. The embodiments of the present disclosure do not limit this.

[0020] It should be noted that the server 4 can be hardware or software. When the server 4 is hardware, it can be various electronic devices that provide various services for the terminal devices 1, 2, and 3. When the server 4 is software, it can be multiple software or software modules that provide various services for the terminal devices 1, 2, and 3, or a single software or software module that provides various services for the terminal devices 1, 2, and 3. The embodiments of the present disclosure do not limit this.

[0021] The network 5 can be a wired network connected by coaxial cables, twisted pairs, and optical fibers, or a wireless network that can interconnect various communication devices without wiring, such as Bluetooth, Near Field Communication (NFC), Infrared, etc. The embodiments of the present disclosure do not limit this.

[0022] Users can establish a communication connection between the terminal devices 1, 2, and 3 and the server 4 via the network 5 to receive or send information, etc. It should be noted that the specific types, quantities, and combinations of the terminal devices 1, 2, and 3, the server 4, and the network 5 can be adjusted according to the actual requirements of the application scenario. The embodiments of the present disclosure do not limit this.

[0023] Figure 2 It is a schematic flowchart of a face recognition method provided by the embodiments of the present disclosure. Figure 2 The face recognition method can be executed by Figure 1 the terminal device or the server. As Figure 2 shown, the face recognition method includes:

[0024] S201, construct a main network, a sliding replication network, a multi-convolution replication network, a multi-activation replication network, and a multi-encoding replication network;

[0025] S202, use the sliding replication network, the multi-convolution replication network, the multi-activation replication network, and the multi-encoding replication network to form a parallel replication network group, and use the main network and the parallel replication network group to form a face recognition model;

[0026] S203, obtain an input image, and divide multiple samples in each class in the input image into a first sample group and a second sample group;

[0027] S204. Use the first sample group of each class to train the multi-convolution replication network, the multi-activation replication network, and the multi-encoding replication network respectively;

[0028] S205. Input the second sample group of each class into the parallel replication network group to output multiple generated prototypes of each class;

[0029] S206. Train the face recognition model according to the first sample group, the second sample group, and the multiple generated prototypes of each class, and perform face recognition using the trained neural network model.

[0030] It should be noted that the input images obtained in the embodiments of the present disclosure are used to train the neural network model. The input images include multiple classes. One generated prototype of each class is equivalent to a class center corresponding to each class, and each class has multiple samples. In the field of face recognition, one class can be a person, and one sample can be a picture of a person. The class center corresponding to one class can be understood as the average value of the features of all pictures of a person.

[0031] The number of samples of each class in the input images is equal to the number of samples in the first sample group and the second sample group of each class. Dividing the multiple samples in each class of the input images into the first sample group and the second sample group can be randomly assigned by a program. Among them, the ratio of the assigned first sample group and the second sample group can be preset in advance. The first sample group is used to train the model or network, and the second sample group is used to generate or determine the generated prototype and the standard prototype. Inputting the second sample group of each class into the parallel replication network group to output multiple generated prototypes of each class can be understood as inputting each sample of the second sample group of each class into the parallel replication network group. Because in the parallel replication network group, the sliding replication network, the multi-convolution replication network, the multi-activation replication network, and the multi-encoding replication network are arranged in parallel, so inputting a sample into the parallel replication network group can obtain 4 generated prototypes (each replication network outputs 1).

[0032] According to the technical solution provided by the embodiments of the present disclosure, because the embodiments of the present disclosure divide the multiple samples in each class of the input images into the first sample group and the second sample group; use the first sample group of each class to train the multi-convolution replication network, the multi-activation replication network, and the multi-encoding replication network respectively; input the second sample group of each class into the parallel replication network group to output multiple generated prototypes of each class; train the face recognition model according to the first sample group, the second sample group, and the multiple generated prototypes of each class, and perform face recognition using the trained neural network model. Therefore, by adopting the above technical means, it can solve the problem in the prior art that in the training of the face recognition model, there is only one prototype for each class, which may cause the parameter update of the face recognition model to fail and the face recognition model cannot fit to the optimal solution, and further improve the accuracy of the face recognition model in recognizing faces.

[0033] In step S204, the multi-convolution replication network, the multi-activation replication network, and the multi-encoding replication network are trained respectively using the first sample group of each class, including: extracting the first feature, the second feature, the third feature, and the fourth feature of each sample in the first sample group of each class through the main network, the multi-convolution replication network, the multi-activation replication network, and the multi-encoding replication network respectively; training the multi-convolution replication network according to the first difference between the first feature and the second feature of each sample in the first sample group of each class; training the multi-activation replication network according to the second difference between the first feature and the third feature of each sample in the first sample group of each class; training the multi-encoding replication network according to the third difference between the first feature and the fourth feature of each sample in the first sample group of each class.

[0034] The features in the present disclosure can be understood as feature maps or matrices corresponding to the feature maps. Specifically, in the first sample group of each class, the first difference of each sample can be the subtraction of the first feature of each sample from the second feature, and the first difference is used as the loss to train the multi-convolution replication network; the subtraction of the first feature and the third feature of each sample obtains the second difference of each sample, and the second difference is used as the loss to train the multi-activation replication network; the subtraction of the first feature and the fourth feature of each sample obtains the third difference of each sample, and the third difference is used as the loss to train the multi-encoding replication network.

[0035] It should be noted that since the main network, the multi-convolution replication network, the multi-activation replication network, and the multi-encoding replication network all have four stages, the extraction of the first feature, the second feature, the third feature, and the fourth feature of each sample in the first sample group of each class by the main network, the multi-convolution replication network, the multi-activation replication network, and the multi-encoding replication network all includes the features of the four stages of each network. For example, the first difference of each sample can be the subtraction of the first feature extracted by the third stage of the main network of each sample from the second feature extracted by the third stage of the multi-convolution replication network, or the subtraction of the first feature extracted by the fourth stage of the main network of each sample from the second feature extracted by the fourth stage of the multi-convolution replication network, etc.

[0036] In step S206, according to the first sample group, the second sample group of each class, and multiple generated prototypes, the face recognition model is trained, including: determining the standard prototype of each class according to the second sample group of each class; extracting the fifth feature of each sample in the first sample group of each class through the face recognition model; training the face recognition model according to the fourth difference between the fifth feature of each sample in the first sample group of each class and the standard prototype of each class; and / or training the face recognition model according to the fifth difference between the fifth feature of each sample in the first sample group of each class and each generated prototype of each class.

[0037] One generated prototype and one standard prototype for each class are equivalent to a class center corresponding to each class. The generated prototype is generated by the parallel replication network group, and the standard prototype is the prototype corresponding to each class under the general face algorithm. Each class can have multiple generated prototypes, but each class has only one standard prototype. When training the face recognition model for the first time, the standard prototype of each class can be determined according to the second sample group of each class. Since the standard prototype of each class determined when training the face recognition model for the first time is saved in the prototype database, when training the face recognition model for non-first time, only the standard prototype of each class needs to be obtained from the prototype database. Determining the standard prototype of each class according to the second sample group of each class can be understood as calculating the class center of each class based on each sample in the second sample group of each class, and taking the class center of each class as the standard prototype of each class.

[0038] Because the embodiments of the present disclosure train the face recognition model based on multiple generated prototypes and one standard prototype for each class, by using the above technical means, it can solve the problems in the prior art that in the training of the face recognition model, each class has only one prototype, which may lead to the failure of parameter update of the face recognition model and the problem that the face recognition model cannot fit to the optimal solution.

[0039] Before training the face recognition model according to the fourth difference between each sample in the first sample group of each class and the standard prototype of each class, the method further includes: setting a positive queue and a negative queue for the standard prototype of each class; identifying each sample in the first sample group of each class as a positive sample or a negative sample through the face recognition model; adding the positive samples in the first sample group of each class to the positive queue of the standard prototype of each class, and adding the negative samples in the first sample group of each class to the negative queue of the standard prototype of each class; updating the standard prototype of each class according to the positive samples in the positive queue of the standard prototype of each class and / or the negative samples in the negative queue of the standard prototype of each class.

[0040] Since the first sample group and the second sample group of each class belong to the same class, the standard prototype of each class determined according to the second sample group of each class is also valid for the first sample group of each class. Similarly, inputting the second sample group of each class into the parallel replication network group and outputting multiple generated prototypes of each class is also valid for the first sample group of each class. Identifying each sample in the first sample group of each class as a positive sample or a negative sample through the face recognition model can be understood as calculating the similarity between the features of each sample in the first sample group of each class and the standard prototype of each class through the face recognition model. When the similarity is greater than a certain threshold, the corresponding sample in the first sample group of each class is identified as a positive sample, and when the similarity is less than a certain threshold, the corresponding sample in the first sample group of each class is identified as a negative sample.

[0041] In an alternative embodiment, updating the standard prototype of each class based on the positive samples in the positive queue of the standard prototype of each class and / or the negative samples in the negative queue of the standard prototype of each class includes: updating the standard prototype of each class using the positive samples in the positive queue of the standard prototype of each class when at least one of the following conditions is met: when the number of positive samples in the positive queue of the standard prototype of each class is greater than the first preset threshold, when there are positive samples in the positive queue of the standard prototype of each class that have existed for more than a preset duration, when there are positive samples in the positive queue of the standard prototype of each class whose similarity to the standard prototype of each class is less than the third preset threshold; updating the standard prototype of each class using the negative samples in the negative queue of the standard prototype of each class when at least one of the following conditions is met: when the number of negative samples in the negative queue of the standard prototype of each class is greater than the fourth preset threshold, when there are negative samples in the negative queue of the standard prototype of each class that have existed for more than a preset duration, when there are negative samples in the negative queue of the standard prototype of each class whose similarity to the standard prototype of each class is greater than the fifth preset threshold.

[0042] Because when determining a sample as a positive sample, the similarity between the features of each sample in the first sample group of each class and the standard prototype of each class is calculated by the face recognition model, and when the similarity is greater than a certain threshold, the corresponding sample in the first sample group of each class is recognized as a positive sample. When updating the standard prototype of each class using the positive samples in the positive queue of the standard prototype of each class, one of the conditions that is met is when there are positive samples in the positive queue of the standard prototype of each class whose similarity to the standard prototype of each class is less than the third preset threshold. Combining these two conditions or judgment bases, it can be understood that when there are samples whose similarity to the standard prototype of each class is within the preset threshold range, the corresponding standard prototype is updated according to this sample. Among them, the preset threshold range is a range greater than the above-mentioned certain threshold and less than the third preset threshold. The reason for using the positive samples whose similarity to the standard prototype of each class is less than the third preset threshold to update the corresponding standard prototype is that the positive samples whose similarity to the standard prototype of each class is less than the third preset threshold have a better effect on updating the corresponding standard prototype.

[0043] Updating the standard prototype of each class using the negative samples in the negative queue of the standard prototype of each class is similar to the above-mentioned updating the standard prototype of each class using the positive samples in the positive queue of the standard prototype of each class, and will not be elaborated again here.

[0044] In an alternative embodiment, the standard prototype can also be optimized or updated through the following formula:

[0045]

[0046] Let \(L\) be the loss function, which can be the cross - entropy loss function. Let \(W\) be the standard prototype, \(i\) be the label of the feature of the sample, \(j\) be the label of the standard prototype, \(x\) be the feature of the sample, and \(P\) ij is the similarity between the \(i\) - th sample and the \(j\) - th standard prototype, \(Q\) is the first preset threshold, \(T\) is the preset duration, \(t\) is the current duration, and \(f\) is the sample.

[0047] In an alternative embodiment, both the main network and the sliding replication network include four stages: both the first stage and the fourth stage are composed of a first preset number of core modules, the second stage is composed of a second preset number of core modules, and the third stage is composed of a third preset number of core modules; each core module is sequentially composed of a normalization layer, a convolutional layer, an activation layer, a convolutional layer, and a normalization layer; the main network updates network parameters using the loss - derivative gradient backpropagation technique; the sliding replication network updates network parameters using the sliding - average optimization technique.

[0048] The network structures of the main network and the sliding replication network are the same, but the methods for updating their respective network parameters are different. For example, in the 4 stages of the main network and the sliding replication network, there are 3, 4, 14, and 3 core modules respectively. Each core module is sequentially composed of a normalization layer, a convolutional layer, an activation layer, a convolutional layer, and a normalization layer, and the convolutional kernel is \(3\times3\). The main network updates network parameters using the loss - derivative gradient backpropagation technique, and the sliding replication network updates network parameters using the sliding - average optimization technique. The loss - derivative gradient backpropagation technique is a prior art and will not be elaborated here. The sliding - average optimization technique can be implemented by the following formula:

[0049] \(N_1=\alpha N_1+(1 - \alpha)N_0\)

[0050] \(\alpha\) is a parameter that can be set, generally 0.99. \(N_1\) is the network parameter of the sliding replication network, \(N_0\) is the network parameter of the main network. The \(N_1\) on the left - hand side of the equation is the network parameter of the sliding replication network at the current moment, and the \(N_1\) on the right - hand side of the equation is the network parameter of the sliding replication network at the previous moment of the current moment.

[0051] In an alternative embodiment, both the multi - convolutional replication network, the multi - activation replication network, and the multi - coding replication network include four stages: both the first stage and the second stage are composed of a fourth preset number of core modules, the third stage is composed of a fifth preset number of core modules, and the fourth stage is composed of a sixth preset number of core modules; each core module of the multi - convolutional replication network and the multi - activation replication network is composed of a seventh preset number of branches and a weighted - sum module, and each core module of the multi - coding replication network is composed of an eighth preset number of branches and a weighted - sum module; among them, the branches are arranged in parallel, and the seventh preset number of branches and the weighted - sum module are connected in a cascaded manner; each branch includes: a normalization layer, a convolutional layer, and an activation layer.

[0052] The number of core modules in each stage of the multi-convolution replication network, the multi-activation replication network, and the multi-encoding replication network is different. However, the branches that make up the core modules are different. Specifically, the number of branches in the core modules of the multi-convolution replication network and the multi-activation replication network is the same, but the functions of the branches are different. The number of branches in the core module of the multi-encoding replication network is different from that of the multi-convolution replication network and the multi-activation replication network, and the functions of the branches are also different. The above-mentioned seventh preset quantity and the eighth preset quantity are the numbers of branches, and only one weighted sum module is needed. The weighted sum module can multiply the multiple input results by the corresponding weights, and then add the multiple multiplication results of the multiple input results multiplied by the corresponding weights.

[0053] For example, in the four stages of the multi-convolution replication network, there are 2, 2, 5, and 1 core modules respectively. Each core module has three branches: the first branch can be composed of a normalization layer, a convolutional layer, an activation layer, a convolutional layer, and a normalization layer in sequence; the second branch can be composed of a normalization layer, a convolutional layer, an activation layer, a convolutional layer, an activation layer, a convolutional layer, and a normalization layer in sequence; the third branch can be composed of a normalization layer, a convolutional layer, an activation layer, a convolutional layer, an activation layer, a convolutional layer, an activation layer, a convolutional layer, and a normalization layer in sequence. The normalization layer of the multi-convolution replication network uses the bn normalization function, and the activation layer uses the relu activation function.

[0054] In the four stages of the multi-activation replication network, there are 2, 2, 5, and 1 core modules respectively. Each core module has three branches: the first branch can be composed of a normalization layer composed of the bn normalization function, a convolutional layer, an activation layer composed of the Mish activation function, a convolutional layer, and a normalization layer composed of the bn normalization function in sequence; the second branch can be composed of a normalization layer composed of the IN normalization function, a convolutional layer, an activation layer composed of the Tanh activation function, a convolutional layer, and a normalization layer composed of the IN normalization function in sequence; the third branch can be composed of a normalization layer composed of the GN normalization function, a convolutional layer, an activation layer composed of the Swish activation function, a convolutional layer, and a normalization layer composed of the GN normalization function in sequence.

[0055] In the four stages of the multi-encoding replication network, there are 2, 2, 5, and 1 core modules respectively. Each core module has two branches: the first branch can be composed of a normalization layer composed of the bn normalization function, a convolutional layer, an activation layer composed of the relu activation function, a convolutional layer, and a normalization layer composed of the bn normalization function in sequence; the second branch can be composed of a normalization layer composed of the IN normalization function, a multi-head self-attention network layer, a convolutional layer, and a normalization layer composed of the IN normalization function and a multi-layer perceptron network layer in sequence.

[0056] All of the above optional technical solutions can be combined arbitrarily to form optional embodiments of the present application, which will not be elaborated here one by one.

[0057] The following are embodiments of the disclosed device, which can be used to execute the method embodiments of the present disclosure. For details not disclosed in the device embodiments of the present disclosure, please refer to the method embodiments of the present disclosure.

[0058] Figure 3 It is a schematic diagram of a face recognition device provided by an embodiment of the present disclosure. As Figure 3 shown, the face recognition device includes:

[0059] A construction module 301, configured to construct a main network, a sliding replication network, a multi-convolution replication network, a multi-activation replication network, and a multi-encoding replication network;

[0060] A composition module 302, configured to use the sliding replication network, the multi-convolution replication network, the multi-activation replication network, and the multi-encoding replication network to form a parallel replication network group, and use the main network and the parallel replication network group to form a face recognition model;

[0061] An acquisition module 303, configured to acquire an input image, and divide multiple samples in each class in the input image into a first sample group and a second sample group;

[0062] A first training module 304, configured to use the first sample group of each class to train the multi-convolution replication network, the multi-activation replication network, and the multi-encoding replication network respectively;

[0063] A network group module 305, configured to input the second sample group of each class into the parallel replication network group, and output multiple generated prototypes of each class;

[0064] A second training module 306, configured to train the face recognition model according to the first sample group, the second sample group, and the multiple generated prototypes of each class, and perform face recognition using the trained neural network model.

[0065] It should be noted that the input image obtained in the embodiments of the present disclosure is for training the neural network model. The input image includes multiple classes. One generated prototype of each class is equivalent to a class center corresponding to each class, and each class has multiple samples. In the field of face recognition, one class can be a person, and one sample can be a picture of a person. The class center corresponding to one class can be understood as the average value of the features of all pictures of a person.

[0066] The number of samples of each class in the input image is equal to the number of samples in the first sample group and the second sample group of each class. Dividing the multiple samples in each class in the input image into the first sample group and the second sample group can be randomly assigned by a program, where the ratio of the assigned first sample group and the second sample group can be preset in advance. The first sample group is used to train the model or network, and the second sample group is used to generate or determine the generated prototypes and the standard prototypes. Inputting the second sample group of each class into the parallel replication network group and outputting multiple generated prototypes of each class can be understood as inputting each sample of the second sample group of each class into the parallel replication network group. Since in the parallel replication network group, the sliding replication network, the multi-convolution replication network, the multi-activation replication network, and the multi-encoding replication network are arranged in parallel, when inputting a sample into the parallel replication network group, 4 generated prototypes can be obtained (each replication network outputs 1).

[0067] According to the technical solution provided by the embodiments of the present disclosure, since the embodiments of the present disclosure divide the multiple samples in each class in the input image into the first sample group and the second sample group; use the first sample group of each class to train the multi-convolution replication network, the multi-activation replication network, and the multi-encoding replication network respectively; input the second sample group of each class into the parallel replication network group and output multiple generated prototypes of each class; train the face recognition model according to the first sample group, the second sample group, and the multiple generated prototypes of each class, and perform face recognition using the trained neural network model. Therefore, by adopting the above technical means, it is possible to solve the problem in the prior art that in the training of the face recognition model, there is only one prototype for each class, which may cause the parameter update of the face recognition model to fail and the face recognition model cannot be fitted to the optimal solution, thereby improving the accuracy of the face recognition model in recognizing faces.

[0068] Optionally, the first training module 404 is further configured to extract the first feature, the second feature, the third feature, and the fourth feature of each sample in the first sample group of each class through the main network, the multi-convolution replication network, the multi-activation replication network, and the multi-encoding replication network respectively; train the multi-convolution replication network according to the first difference between the first feature and the second feature of each sample in the first sample group of each class; train the multi-activation replication network according to the second difference between the first feature and the third feature of each sample in the first sample group of each class; train the multi-encoding replication network according to the third difference between the first feature and the fourth feature of each sample in the first sample group of each class.

[0069] The features in the present disclosure can be understood as feature maps or matrices corresponding to the feature maps. Specifically, in the first sample group of each class, the first difference of each sample can be obtained by subtracting the second feature from the first feature of each sample, and the first difference is used as the loss to train the multi-convolution replication network; the second difference of each sample is obtained by subtracting the third feature from the first feature of each sample, and the second difference is used as the loss to train the multi-activation replication network; the third difference of each sample is obtained by subtracting the fourth feature from the first feature of each sample, and the third difference is used as the loss to train the multi-encoding replication network.

[0070] It should be noted that since the main network, the multi-convolution replication network, the multi-activation replication network, and the multi-encoding replication network all have four stages, the first feature, the second feature, the third feature, and the fourth feature of each sample in the first sample group of each class extracted by the main network, the multi-convolution replication network, the multi-activation replication network, and the multi-encoding replication network all include the features of the four stages of each network. For example, the first difference of each sample can be obtained by subtracting the second feature extracted by the third stage of the multi-convolution replication network from the first feature extracted by the third stage of the main network, or by subtracting the second feature extracted by the fourth stage of the multi-convolution replication network from the first feature extracted by the fourth stage of the main network, etc.

[0071] Optionally, the second training module 406 is further configured to determine the standard prototype of each class according to the second sample group of each class; extract the fifth feature of each sample in the first sample group of each class through the face recognition model; train the face recognition model according to the fourth difference between the fifth feature of each sample in the first sample group of each class and the standard prototype of each class; and / or train the face recognition model according to the fifth difference between the fifth feature of each sample in the first sample group of each class and each generated prototype of each class.

[0072] One generated prototype and the standard prototype of each class are equivalent to a class center corresponding to each class. Only the generated prototype is generated by the parallel replication network group, and the standard prototype is the prototype corresponding to each class under the general face algorithm. Each class can have multiple generated prototypes, but each class has only one standard prototype. When training the face recognition model for the first time, the standard prototype of each class can be determined according to the second sample group of each class. Since the standard prototype of each class determined when training the face recognition model for the first time is stored in the prototype database, when training the face recognition model for non-first time, only the standard prototype of each class needs to be obtained from the prototype database. Determining the standard prototype of each class according to the second sample group of each class can be understood as calculating the class center of each class according to each sample in the second sample group of each class, and taking the class center of each class as the standard prototype of each class.

[0073] Since the embodiments of the present disclosure train the face recognition model based on multiple generated prototypes and a standard prototype for each class, by adopting the above technical means, it is possible to solve the problems in the prior art that in the training of the face recognition model, there is only one prototype for each class, which may lead to the failure of parameter update of the face recognition model and the failure of the face recognition model to fit to the optimal solution.

[0074] Optionally, the second training module 406 is further configured to set a positive queue and a negative queue for the standard prototype of each class; identify each sample in the first sample group of each class as a positive sample or a negative sample through the face recognition model; add the positive samples in the first sample group of each class to the positive queue of the standard prototype of each class, and add the negative samples in the first sample group of each class to the negative queue of the standard prototype of each class; update the standard prototype of each class according to the positive samples in the positive queue of the standard prototype of each class and / or the negative samples in the negative queue of the standard prototype of each class.

[0075] Since the first sample group and the second sample group of each class belong to the same class, the standard prototype of each class determined according to the second sample group of each class is also effective for the first sample group of each class. Similarly, inputting the second sample group of each class into the parallel replication network group and outputting multiple generated prototypes of each class is also effective for the first sample group of each class. Identifying each sample in the first sample group of each class as a positive sample or a negative sample through the face recognition model can be understood as calculating the similarity between the features of each sample in the first sample group of each class and the standard prototype of each class through the face recognition model. When the similarity is greater than a certain threshold, the corresponding sample in the first sample group of each class is identified as a positive sample, and when the similarity is less than a certain threshold, the corresponding sample in the first sample group of each class is identified as a negative sample.

[0076] Optionally, the second training module 406 is further configured to update the standard prototype of each class using the positive samples in the positive queue of the standard prototype of each class when at least one of the following conditions is met: when the number of positive samples in the positive queue of the standard prototype of each class is greater than the first preset threshold, when there are positive samples in the positive queue of the standard prototype of each class that have exceeded the preset duration, when there are positive samples in the positive queue of the standard prototype of each class whose similarity to the standard prototype of each class is less than the third preset threshold; update the standard prototype of each class using the negative samples in the negative queue of the standard prototype of each class when at least one of the following conditions is met: when the number of negative samples in the negative queue of the standard prototype of each class is greater than the fourth preset threshold, when there are negative samples in the negative queue of the standard prototype of each class that have exceeded the preset duration, when there are negative samples in the negative queue of the standard prototype of each class whose similarity to the standard prototype of each class is greater than the fifth preset threshold.

[0077] When determining a sample as a positive sample, the similarity between the features of each sample in the first sample group of each class and the standard prototype of each class is calculated through a face recognition model. When the similarity is greater than a certain threshold, the corresponding sample in the first sample group of each class is recognized as a positive sample. When updating the standard prototype of each class with the positive samples in the positive queue of the standard prototype of each class, one of the conditions that is satisfied is that when there are positive samples in the positive queue of the standard prototype of each class with a similarity less than the third preset threshold to the standard prototype of each class. These two conditions or judgment bases, when combined, can be understood as that when there are samples with similarities to the standard prototype of each class within a preset threshold range, the corresponding standard prototype is updated according to this sample. Among them, the preset threshold range is an interval greater than the above-mentioned certain threshold and less than the third preset threshold. The reason for using positive samples with similarities less than the third preset threshold to the standard prototype of each class to update the corresponding standard prototype is that positive samples with similarities less than the third preset threshold to the standard prototype of each class have a better effect on updating the corresponding standard prototype.

[0078] Updating the standard prototype of each class with the negative samples in the negative queue of the standard prototype of each class and the above-mentioned updating the standard prototype of each class with the positive samples in the positive queue of the standard prototype of each class are understood similarly and will not be elaborated again here.

[0079] Optionally, the second training module 406 is further configured to optimize or update the standard prototype through the following formula:

[0080]

[0081] L is the loss function, which can be a cross-entropy loss function, W is the standard prototype, i is the label of the feature of the sample, j is the label of the standard prototype, x is the feature of the sample, P ij is the similarity between the i-th sample and the j-th standard prototype, Q is the first preset threshold, T is the preset duration, t is the current duration, and f is the sample.

[0082] Optionally, both the main network and the sliding replica network include four stages: both the first stage and the fourth stage are composed of a first preset number of core modules, the second stage is composed of a second preset number of core modules, and the third stage is composed of a third preset number of core modules; each core module is successively composed of a normalization layer, a convolutional layer, an activation layer, a convolutional layer, and a normalization layer; the main network updates the network parameters using the loss derivative gradient backpropagation technique; the sliding replica network updates the network parameters using the sliding average optimization technique.

[0083] The network structures of the main network and the sliding replica network are the same, but the methods for updating their respective network parameters are different. For example, in the four stages of the main network and the sliding replica network, there are 3, 4, 14, and 3 core modules respectively. Each core module is sequentially composed of a normalization layer, a convolutional layer, an activation layer, a convolutional layer, and a normalization layer. The convolutional kernel is 3*3. The main network uses the loss derivative gradient backpropagation technique to update the network parameters, and the sliding replica network uses the sliding average optimization technique to update the network parameters. The loss derivative gradient backpropagation technique is an existing technique and will not be elaborated here. The sliding average optimization technique can be implemented through the following formula:

[0084] N1 = αN1+(1 - α)N0

[0085] α is a parameter that can be set, generally 0.99. N1 is the network parameter of the sliding replica network, and N0 is the network parameter of the main network. N1 on the left side of the equation is the network parameter of the sliding replica network at the current moment, and N1 on the right side of the equation is the network parameter of the sliding replica network at the previous moment of the current moment.

[0086] Optionally, the multi-convolution replica network, the multi-activation replica network, and the multi-encoding replica network all include four stages: The first stage and the second stage are both composed of a fourth preset number of core modules, the third stage is composed of a fifth preset number of core modules, and the fourth stage is composed of a sixth preset number of core modules; Each core module of the multi-convolution replica network and the multi-activation replica network is composed of a seventh preset number of branches and a weighted sum module, and each core module of the multi-encoding replica network is composed of an eighth preset number of branches and a weighted sum module; Among them, the branches are arranged in parallel, and the seventh preset number of branches and the weighted sum module are connected in a cascaded manner; Each branch includes: a normalization layer, a convolutional layer, and an activation layer.

[0087] The number of core modules in each stage of the multi-convolution replica network, the multi-activation replica network, and the multi-encoding replica network is the same, but the branches that make up the core modules are different. The main difference is that the number of branches in the core modules of the multi-convolution replica network and the multi-activation replica network is the same, but the functions of the branches are different. The number of branches in the core modules of the multi-encoding replica network is different from that of the multi-convolution replica network and the multi-activation replica network, and the functions of the branches are also different.

[0088] For example, the four stages of the multi-convolution replication network have 2, 2, 5, and 1 core modules respectively. Each core module has three branches: the first branch can be successively composed of a normalization layer, a convolutional layer, an activation layer, a convolutional layer, and a normalization layer; the second branch can be successively composed of a normalization layer, a convolutional layer, an activation layer, a convolutional layer, an activation layer, a convolutional layer, and a normalization layer; the third branch can be successively composed of a normalization layer, a convolutional layer, an activation layer, a convolutional layer, an activation layer, a convolutional layer, an activation layer, a convolutional layer, and a normalization layer. The normalization layer of the multi-convolution replication network uses the bn normalization function, and the activation layer uses the relu activation function.

[0089] The four stages of the multi-activation replication network have 2, 2, 5, and 1 core modules respectively. Each core module has three branches: the first branch can be successively composed of a normalization layer composed of the bn normalization function, a convolutional layer, an activation layer composed of the Mish activation function, a convolutional layer, and a normalization layer composed of the bn normalization function; the second branch can be successively composed of a normalization layer composed of the IN normalization function, a convolutional layer, an activation layer composed of the Tanh activation function, a convolutional layer, and a normalization layer composed of the IN normalization function; the third branch can be successively composed of a normalization layer composed of the GN normalization function, a convolutional layer, an activation layer composed of the Swish activation function, a convolutional layer, and a normalization layer composed of the GN normalization function.

[0090] The four stages of the multi-encoding replication network have 2, 2, 5, and 1 core modules respectively. Each core module has two branches: the first branch can be successively composed of a normalization layer composed of the bn normalization function, a convolutional layer, an activation layer composed of the relu activation function, a convolutional layer, and a normalization layer composed of the bn normalization function; the second branch can be successively composed of a normalization layer composed of the IN normalization function, a multi-head self-attention network layer, a convolutional layer, and a normalization layer composed of the IN normalization function and a multi-layer perceptron network layer.

[0091] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not mean the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present disclosure.

[0092] Figure 4 is a schematic diagram of the electronic device 4 provided by the embodiment of the present disclosure. As Figure 4 shown, the electronic device 4 of this embodiment includes: a processor 401, a memory 402, and a computer program 403 stored in the memory 402 and executable on the processor 401. When the processor 401 executes the computer program 403, the steps in the above-mentioned method embodiments are implemented. Alternatively, when the processor 401 executes the computer program 403, the functions of each module / unit in the above-mentioned device embodiments are implemented.

[0093] Exemplarily, the computer program 403 can be divided into one or more modules / units. The one or more modules / units are stored in the memory 402 and executed by the processor 401 to implement the present disclosure. The one or more modules / units can be a series of computer program instruction segments capable of performing specific functions, and these instruction segments are used to describe the execution process of the computer program 403 in the electronic device 4.

[0094] The electronic device 4 can be a desktop computer, a notebook, a palm computer, a cloud server, or other electronic devices. The electronic device 4 can include, but is not limited to, the processor 401 and the memory 402. Those skilled in the art can understand that Figure 4 merely examples of the electronic device 4, which do not constitute a limitation on the electronic device 4. It may include more or fewer components than shown in the figure, or combine certain components, or different components. For example, the electronic device may also include input / output devices, network access devices, a bus, etc.

[0095] The processor 401 can be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc.

[0096] The memory 402 can be an internal storage unit of the electronic device 4. For example, the hard disk or memory of the electronic device 4. The memory 402 can also be an external storage device of the electronic device 4. For example, a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the electronic device 4. Further, the memory 402 can also include both the internal storage unit and the external storage device of the electronic device 4. The memory 402 is used to store computer programs and other programs and data required by the electronic device. The memory 402 can also be used to temporarily store data that has been output or will be output.

[0097] Those skilled in the art can clearly understand that, for the convenience and conciseness of description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiments can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of the functional units and modules are only for the convenience of mutual distinction and do not limit the protection scope of this application. The specific working processes of the units and modules in the above system can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0098] In the above embodiments, the descriptions of the various embodiments have their own emphases. For the parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0099] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or by a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this disclosure.

[0100] In the embodiments provided by this disclosure, it should be understood that the disclosed device / electronic device and method can be implemented in other ways. For example, the device / electronic device embodiments described above are only illustrative. For example, the division of modules or units is only a logical function division, and there can be other division methods in actual implementation. Multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed mutual coupling or direct coupling or communication connection can be through some interfaces, and the indirect coupling or communication connection of the device or unit can be in electrical, mechanical or other forms.

[0101] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0102] In addition, in each of the embodiments of the present disclosure, the functional units may be integrated into one processing unit, or each unit may exist physically alone, or two or more units may be integrated into one unit. The above integrated unit may be implemented in the form of hardware or in the form of a software functional unit.

[0103] If the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it may be stored in a computer-readable storage medium. Based on such understanding, to implement all or part of the processes in the above method embodiments of the present disclosure, it may also be completed by instructing relevant hardware through a computer program. The computer program may be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above method embodiments may be implemented. The computer program may include computer program code, and the computer program code may be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the computer-readable medium may be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.

[0104] The above embodiments are only used to illustrate the technical solutions of the present disclosure, rather than to limit them; although the present disclosure has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present disclosure, and should all be included in the protection scope of the present disclosure.

Claims

1. A face recognition method, characterized in that, Including: Construct a main network, a sliding replication network, a multi-convolution replication network, a multi-activation replication network, and a multi-encoding replication network; Use the sliding replication network, the multi-convolution replication network, the multi-activation replication network, and the multi-encoding replication network to form a parallel replication network group, and use the main network and the parallel replication network group to form a face recognition model; Obtain an input image, and divide multiple samples in each class of the input image into a first sample group and a second sample group; Use the first sample group of each class to train the multi-convolution replication network, the multi-activation replication network, and the multi-encoding replication network respectively; Input the second sample group of each class into the parallel replication network group, and output multiple generated prototypes of each class; one generated prototype of each class is equivalent to a class center corresponding to each class; Train the face recognition model according to the first sample group, the second sample group, and the multiple generated prototypes of each class, and use the trained face recognition model for face recognition; Wherein it includes: Determine the standard prototype of each class according to the second sample group of each class; the standard prototype of each class is the class center corresponding to each class under the general face algorithm; Extract the fifth feature of each sample in the first sample group of each class through the face recognition model; Train the face recognition model according to the fourth difference between the fifth feature of each sample in the first sample group of each class and the standard prototype of each class; and / or Train the face recognition model according to the fifth difference between the fifth feature of each sample in the first sample group of each class and each generated prototype of each class; Both the main network and the sliding replication network include four stages: both the first stage and the fourth stage are composed of a first preset number of first core modules, the second stage is composed of a second preset number of the first core modules, and the third stage is composed of a third preset number of the first core modules; each first core module is successively composed of a normalization layer, a convolutional layer, an activation layer, a convolutional layer, and a normalization layer; the main network updates network parameters using the loss derivative gradient backpropagation technique; the sliding replication network updates network parameters using the sliding average optimization technique; The multi-convolution replication network, the multi-activation replication network, and the multi-encoding replication network all include four stages: The first stage and the second stage are both composed of a fourth preset number of second core modules, the third stage is composed of a fifth preset number of the second core modules, and the fourth stage is composed of a sixth preset number of the second core modules; Each of the second core modules of the multi-convolution replication network and the multi-activation replication network is composed of a seventh preset number of branches and a weighted sum module, and each of the second core modules of the multi-encoding replication network is composed of an eighth preset number of branches and a weighted sum module; Wherein, the branches in each of the second core modules are arranged in parallel, and the parallel branches and the weighted sum module are connected in a cascaded manner to perform weighted addition on the multiple results input by the parallel branches to the weighted sum module; Each branch includes: a normalization layer, a convolutional layer, and an activation layer.

2. The method according to claim 1, wherein The step of using the first sample group of each class to train the multi-convolution replication network, the multi-activation replication network, and the multi-encoding replication network respectively includes: Extract the first feature, second feature, third feature, and fourth feature of each sample in the first sample group of each class through the main network, the multi-convolution replication network, the multi-activation replication network, and the multi-encoding replication network respectively; Train the multi-convolution replication network according to the first difference between the first feature and the second feature of each sample in the first sample group of each class; Train the multi-activation replication network according to the second difference between the first feature and the third feature of each sample in the first sample group of each class; Train the multi-encoding replication network according to the third difference between the first feature and the fourth feature of each sample in the first sample group of each class.

3. The method according to claim 1, wherein Before training the face recognition model according to the fourth difference between the fifth feature of each sample in the first sample group of each class and the standard prototype of each class, the method further includes: Set a positive queue and a negative queue for the standard prototype of each class; Identify whether each sample in the first sample group of each class is a positive sample or a negative sample through the face recognition model; Add the positive samples in the first sample group of each class to the positive queue of the standard prototype of each class, and add the negative samples in the first sample group of each class to the negative queue of the standard prototype of each class; Update the standard prototype of each class according to the positive samples in the positive queue of the standard prototype of each class and / or the negative samples in the negative queue of the standard prototype of each class.

4. The method according to claim 3, wherein The step of updating the standard prototype of each class according to the positive samples in the positive queue of the standard prototype of each class and / or the negative samples in the negative queue of the standard prototype of each class includes: Update the standard prototype of each class with the positive samples in the positive queue of the standard prototype of each class when at least one of the following conditions is satisfied: when the number of positive samples in the positive queue of the standard prototype of each class is greater than the first preset threshold, when there are positive samples in the positive queue of the standard prototype of each class that have exceeded the preset duration, when there are positive samples in the positive queue of the standard prototype of each class whose similarity to the standard prototype of each class is less than the third preset threshold; Update the standard prototype of each class with the negative samples in the negative queue of the standard prototype of each class when at least one of the following conditions is satisfied: when the number of negative samples in the negative queue of the standard prototype of each class is greater than the fourth preset threshold, when there are negative samples in the negative queue of the standard prototype of each class that have exceeded the preset duration, when there are negative samples in the negative queue of the standard prototype of each class whose similarity to the standard prototype of each class is greater than the fifth preset threshold.

5. A face recognition device, characterized in that, Comprising: A construction module configured to construct a main network, a sliding replication network, a multi-convolution replication network, a multi-activation replication network, and a multi-encoding replication network; A composition module configured to use the sliding replication network, the multi-convolution replication network, the multi-activation replication network, and the multi-encoding replication network to form a parallel replication network group, and use the main network and the parallel replication network group to form a face recognition model; An acquisition module configured to acquire an input image and divide multiple samples in each class in the input image into a first sample group and a second sample group; A first training module configured to train the multi-convolution replication network, the multi-activation replication network, and the multi-encoding replication network respectively using the first sample group of each class; A network group module configured to input the second sample group of each class into the parallel replication network group and output multiple generated prototypes of each class; where one generated prototype of each class is equivalent to a class center corresponding to each class; A second training module configured to train the face recognition model according to the first sample group, the second sample group, and the multiple generated prototypes of each class, and perform face recognition using the trained face recognition model; Including: determining a standard prototype for each class according to the second sample group of each class; the standard prototype of each class is the class center corresponding to each class under the general face algorithm; extracting a fifth feature of each sample in the first sample group of each class through the face recognition model; training the face recognition model according to a fourth difference between the fifth feature of each sample in the first sample group of each class and the standard prototype of each class; and / or training the face recognition model according to a fifth difference between the fifth feature of each sample in the first sample group of each class and each generated prototype of each class; Both the main network and the sliding replication network include four stages: both the first stage and the fourth stage are composed of a first preset number of first core modules, the second stage is composed of a second preset number of the first core modules, and the third stage is composed of a third preset number of the first core modules; each of the first core modules is sequentially composed of a normalization layer, a convolutional layer, an activation layer, a convolutional layer, and a normalization layer; the main network updates network parameters by using a loss derivative gradient backpropagation technique; the sliding replication network updates network parameters by using a sliding average optimization technique; Both the multi-convolution replication network, the multi-activation replication network, and the multi-encoding replication network include four stages: both the first stage and the second stage are composed of a fourth preset number of second core modules, the third stage is composed of a fifth preset number of the second core modules, and the fourth stage is composed of a sixth preset number of the second core modules; each of the second core modules of the multi-convolution replication network and the multi-activation replication network is composed of a seventh preset number of branches and a weighted sum module, and each of the core modules of the multi-encoding replication network is composed of an eighth preset number of branches and a weighted sum module; wherein, the branches in each of the second core modules are arranged in parallel, and the parallel branches and the weighted sum module are connected in a cascaded manner to perform weighted addition on multiple results input to the weighted sum module by the parallel branches; each branch includes: a normalization layer, a convolutional layer, and an activation layer.

6. An electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 4.

7. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the method according to any one of claims 1 to 4.

Citation Information

Patent Citations

  • Identity authentication method and apparatus

    CN108491805A

  • Design method of multi-class center classification network model

    CN111242245A