Face recognition method and device in presence of occlusion

By acquiring and generating occluded images, extracting features and calculating similarity using a face recognition model, the problem of low accuracy of face recognition models trained by traditional models under occlusion conditions is solved, achieving higher recognition accuracy.

CN115880756BActive Publication Date: 2026-07-14BEIJING LONGZHI DIGITAL TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING LONGZHI DIGITAL TECH CO LTD
Filing Date
2022-12-20
Publication Date
2026-07-14

AI Technical Summary

Technical Problem

Current face recognition models trained using traditional methods have low accuracy in recognizing images with occlusion.

Method used

By acquiring the base database dataset, generating occluded images, extracting features using a face recognition model, and calculating similarity, the target base database image most similar to the image to be identified is determined.

Benefits of technology

It improves the accuracy of face recognition models in occluded image conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115880756B_ABST
    Figure CN115880756B_ABST
Patent Text Reader

Abstract

The present disclosure relates to the technical field of face recognition, and provides a face recognition method and device in the presence of occlusion. The method comprises: acquiring a base data set, wherein the base data set contains multiple base pictures; generating an occlusion for each base picture to obtain an occluded picture corresponding to each base picture; when a face picture to be recognized is detected, extracting a first feature of the face picture, a second feature of each base picture, and a third feature of the occluded picture corresponding to each base picture by using a face recognition model; calculating the similarity of the first feature and the second feature of each base picture and the third feature of the occluded picture corresponding to the base picture respectively; determining a target base picture with the highest similarity to the face picture according to the similarity of the first feature and the second feature of each base picture and the third feature of the occluded picture corresponding to the base picture, and recognizing the face picture as a target object corresponding to the target base picture.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of facial recognition technology, and in particular to a facial recognition method and apparatus for faces with occlusion. Background Technology

[0002] Currently, because people need to wear masks when going out, the images to be recognized during facial recognition are often obscured. However, facial recognition systems that handle occlusion do not adequately account for this situation, leading to low accuracy in recognizing obscured images after training.

[0003] In realizing the present invention, the inventors discovered at least the following technical problems in the related technology: the face recognition model trained based on traditional model training methods has low accuracy in recognizing images with occlusion. Summary of the Invention

[0004] In view of this, the present disclosure provides a face recognition method, apparatus, electronic device, and computer-readable storage medium with occlusion, to solve the problem in the prior art that the face recognition model trained based on traditional model training methods has low accuracy in recognizing images with occlusion.

[0005] A first aspect of this disclosure provides a face recognition method with occlusion, comprising: acquiring a base database dataset, wherein the base database dataset contains multiple base database images, each of which is free from occlusion; generating an occlusion for each base database image to obtain an occluded image corresponding to each base database image; when a face image to be recognized is detected, extracting a first feature of the face image, a second feature of each base database image, and a third feature of the occluded image corresponding to each base database image using a face recognition model; calculating the similarity between the first feature and the second feature of each base database image and the third feature of the occluded image corresponding to the base database image, respectively; determining a target base database image with the highest similarity to the face image based on the similarity between the first feature and the second feature of each base database image and the third feature of the occluded image corresponding to the base database image, and recognizing the face image as the target object corresponding to the target base database image.

[0006] A second aspect of this disclosure provides a face recognition device with occlusion, comprising: an acquisition module configured to acquire a database dataset, wherein the database dataset contains multiple database images, each database image being free of occlusion; a generation module configured to generate an occlusion for each database image, obtaining an occluded image corresponding to each database image; a model module configured to, upon detecting a face image to be recognized, extract a first feature of the face image, a second feature of each database image, and a third feature of the occluded image corresponding to each database image using a face recognition model; a calculation module configured to calculate the similarity between the first feature and the second feature of each database image and the third feature of the occluded image corresponding to the database image, respectively; and a determination module configured to, based on the similarity between the first feature and the second feature of each database image and the third feature of the occluded image corresponding to the database image, determine a target database image with the highest similarity to the face image, and identify the face image as a target object corresponding to the target database image.

[0007] A third aspect of this disclosure provides an electronic device including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the method described above.

[0008] A fourth aspect of this disclosure provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the above-described method.

[0009] The beneficial effects of this disclosure embodiment compared with the prior art are as follows: This disclosure embodiment obtains a database dataset containing multiple database images, each without occlusion; generates occlusions for each database image, obtaining an occluded image corresponding to each database image; when a face image to be identified is detected, a face recognition model is used to extract the first feature of the face image, the second feature of each database image, and the third feature of the occluded image corresponding to each database image; the similarity between the first feature and the second feature of each database image and the third feature of the occluded image corresponding to that database image is calculated respectively; based on the similarity between the first feature and the second feature of each database image and the third feature of the occluded image corresponding to that database image, the target database image with the highest similarity to the face image is determined, and the face image is identified as the target object corresponding to the target database image. Therefore, by adopting the above technical means, the problem of low accuracy in recognizing images with occlusions by face recognition models trained based on traditional model training methods in the prior art can be solved, thereby improving the accuracy of the model in recognizing images with occlusions. Attached Figure Description

[0010] To more clearly illustrate the technical solutions in the embodiments of this disclosure, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this disclosure. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0011] Figure 1 This is a schematic diagram illustrating an application scenario of an embodiment of this disclosure;

[0012] Figure 2 This is a schematic flowchart of a face recognition method with occlusion provided in an embodiment of this disclosure;

[0013] Figure 3 This is a schematic diagram of the structure of a face recognition device with occlusion provided in an embodiment of this disclosure;

[0014] Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. Detailed Implementation

[0015] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, so as to provide a thorough understanding of the embodiments of this disclosure. However, those skilled in the art will understand that this disclosure may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this disclosure with unnecessary detail.

[0016] A face recognition method and apparatus with occlusion according to an embodiment of the present disclosure will now be described in detail with reference to the accompanying drawings.

[0017] Figure 1 This is a schematic diagram illustrating an application scenario of an embodiment of this disclosure. The application scenario may include terminal devices 101, 102, and 103, server 104, and network 105.

[0018] Terminal devices 101, 102, and 103 can be hardware or software. When terminal devices 101, 102, and 103 are hardware, they can be various electronic devices with displays that support communication with server 104, including but not limited to smartphones, tablets, laptops, and desktop computers. When terminal devices 101, 102, and 103 are software, they can be installed in the aforementioned electronic devices. Terminal devices 101, 102, and 103 can be implemented as multiple software programs or software modules, or as a single software program or software module; this disclosure does not impose any limitations on this. Furthermore, various applications can be installed on terminal devices 101, 102, and 103, such as data processing applications, instant messaging tools, social platform software, search applications, shopping applications, etc.

[0019] Server 104 can be a server that provides various services, such as a backend server that receives requests sent by terminal devices with which it has established communication connections. This backend server can receive and analyze the requests sent by the terminal devices and generate processing results. Server 104 can be a single server, a server cluster consisting of several servers, or a cloud computing service center. This embodiment of the disclosure does not impose any limitations on these aspects.

[0020] It should be noted that server 104 can be either hardware or software. When server 104 is hardware, it can be various electronic devices that provide various services to terminal devices 101, 102, and 103. When server 104 is software, it can be multiple software programs or software modules that provide various services to terminal devices 101, 102, and 103, or it can be a single software program or software module that provides various services to terminal devices 101, 102, and 103. This disclosure does not limit the scope of the embodiments.

[0021] Network 105 can be a wired network using coaxial cable, twisted pair, and fiber optic connection, or it can be a wireless network that enables interconnection of various communication devices without wiring, such as Bluetooth, Near Field Communication (NFC), Infrared, etc. This disclosure does not limit the scope of the network.

[0022] Users can establish a communication connection with server 104 via network 105 through terminal devices 101, 102, and 103 to receive or send information, etc. It should be noted that the specific types, quantities, and combinations of terminal devices 101, 102, and 103, server 104, and network 105 can be adjusted according to the actual needs of the application scenario, and this disclosure embodiment does not impose any limitations on this.

[0023] Figure 2This is a schematic flowchart of a face recognition method with occlusion provided in an embodiment of this disclosure. Figure 2 Face recognition methods with occlusion can be derived from Figure 1 The computer or server, or the software on the computer or server, executes the command. For example... Figure 2 As shown, the face recognition method with occlusion includes:

[0024] S201, Obtain the base database dataset, which contains multiple base database images, each of which is free of occlusion;

[0025] S202, generate occlusion for each base image, and obtain the occlusion image corresponding to each base image;

[0026] S203, when a face image to be identified is detected, the face recognition model is used to extract the first feature of the face image, the second feature of each base image, and the third feature of the occluded image corresponding to each base image.

[0027] S204, calculate the similarity between the first feature and the second feature of each base image and the third feature of the occluded image corresponding to the base image;

[0028] S205, based on the similarity between the first feature and the second feature of each base image and the third feature of the occluded image corresponding to the base image, determine the target base image with the highest similarity to the face image, and identify the face image as the target object corresponding to the target base image.

[0029] The base dataset consists of a large number of pre-stored face images, also known as the base image set. When a face image to be identified is detected, calculations show that the face image to be identified has the highest similarity to the target image in the base dataset. Therefore, the face image to be identified is determined to be the image of the target object corresponding to the target image in the base dataset. The face recognition model can be a residual neural network or similar model. Features and class centers can be understood as vectors or matrices, and the similarity can be cosine similarity. Occlusion is generated for each base image, simulating face images of people wearing masks, etc., i.e., occluded images.

[0030] According to the technical solution provided in this disclosure, a base database dataset is obtained, wherein the base database dataset contains multiple base database images, and each base database image is free from occlusion; an occlusion is generated for each base database image to obtain an occluded image corresponding to each base database image; when a face image to be identified is detected, a face recognition model is used to extract a first feature of the face image, a second feature of each base database image, and a third feature of the occluded image corresponding to each base database image; the similarity between the first feature and the second feature of each base database image and the third feature of the occluded image corresponding to that base database image is calculated respectively; based on the similarity between the first feature and the second feature of each base database image and the third feature of the occluded image corresponding to that base database image, a target base database image with the highest similarity to the face image is determined, and the face image is identified as the target object corresponding to the target base database image. Therefore, by adopting the above technical means, the problem of low accuracy in recognizing images with occlusion in face recognition models trained based on traditional model training methods in the prior art can be solved, thereby improving the accuracy of the model in recognizing images with occlusion.

[0031] Before extracting the first feature of the face image, the second feature of each base image, and the third feature of the occluded image corresponding to each base image using a face recognition model when a face image to be identified is detected, the method further includes: obtaining a training dataset, wherein the training dataset includes: multiple classes, each class is divided into occluded subclasses and non-occluded subclasses, and each subclass contains multiple training samples; modifying the loss function of the face recognition model accordingly based on whether each class is divided into occluded and non-occluded subclasses; and training the face recognition model using the modified loss function based on the training dataset.

[0032] A class can be viewed as a person, and each person has multiple occluded training samples from the occluded subclass and multiple unoccluded training samples from the unoccluded subclass. The class center is the matrix or vector that best represents the features of the samples in that class or subclass. Modifying the loss function of the face recognition model based on whether each class is divided into occluded and unoccluded subclasses means that the modified loss function can comprehensively consider images from both occluded and unoccluded subclasses. Existing loss functions do not classify each class into occluded and unoccluded subclasses; that is, they do not distinguish between occluded and unoccluded images, but only process all images in a general way.

[0033] It should be noted that training samples are images. In this disclosure, the images used in training are referred to as training samples, and the images used in recognition are referred to as images. In fact, training samples are images.

[0034] Based on the division of each class into occluded and non-occluded subclasses, the loss function of the face recognition model is modified accordingly, including: based on the division of each class into occluded and non-occluded subclasses, the target formula for calculating features and class centers is modified so that the modified target formula takes into account both cases where there is occlusion and cases where there is no occlusion in each class, and each subclass corresponds to a class center; the loss function is modified using the modified target formula.

[0035] In one alternative embodiment, it includes:

[0036] The revised target formula is

[0037]

[0038] The modified loss function is

[0039]

[0040] Among them, W jk Let x be the class center of the k-th subclass of the j-th class, T be the transpose symbol, and x be the class center of the k-th subclass of the j-th class. i Let θ be the feature of the i-th training sample. j ′ for x i With W jk The angle between them, W jk Not x i The corresponding class center, For x i With W ik The angle between them, W ik It is x i The corresponding class centers are s, where s is the feature scaling parameter, m is the additive margin parameter, and N is the total number of training samples.

[0041] calculate The target formula will also be used; you only need to use W. ik Replace W jk The calculated result is When k is 0, the 0th subclass is the occluded subclass; when k is 1, the 1st subclass is the non-occluded subclass. The feature scaling parameter and the additive margin parameter can both be adjusted in advance. The max() function takes the maximum of the two or more parameters; exp is an exponential function with the natural constant e as its base; and log is the sign of the logarithmic function.

[0042] In one alternative embodiment, it includes:

[0043] The original target formula was:

[0044]

[0045] The loss function before modification was

[0046]

[0047] W j Let θ be the class center of the j-th class. j For x i With W j The angle between them, W j Not x i The corresponding class center, For x i With W i The angle between them.

[0048] Based on the similarity between the first feature and the second feature of each database image, and the third feature of the corresponding occluded image of the database image, the target database image with the highest similarity to the face image is determined, including: taking the largest similarity value between the first feature and the second feature of each database image, and the first feature and the third feature of the corresponding occluded image of the database image, as the similarity between the face image and the database image; and determining the target database image with the highest similarity to the face image based on the similarity between the face image and each database image.

[0049] The similarity between a face image and each image in the database is calculated using the following formula:

[0050] s1 = cos < f1, f 2,1 >

[0051] s2 = cos < f1, f 2,2 >

[0052] s = max(s1, s2)

[0053] f1 is the first feature, f 2,1 f is the second feature of each base image. 2,2 Let s1 be the third feature of the occluded image corresponding to the base image, s2 be the similarity between the first feature and the second feature of the base image, s2 be the similarity between the first feature and the third feature of the occluded image corresponding to the base image, and s be the similarity between the face image and the base image.

[0054] In one optional embodiment, the method includes: using a UVMap algorithm to generate occlusion for each base image, thereby obtaining an occlusion image corresponding to each base image.

[0055] The UVMap algorithm, also known as UV mapping, precisely maps each point on an image to the surface of a model object. Software then performs smooth interpolation at the gaps between these points. This embodiment utilizes the UVMap algorithm to occlude certain features in an image.

[0056] All of the above-mentioned optional technical solutions can be combined in any way to form the optional embodiments of this application, and will not be described in detail here.

[0057] The following are embodiments of the apparatus disclosed herein, which can be used to execute embodiments of the method disclosed herein. For details not disclosed in the apparatus embodiments of this disclosure, please refer to the embodiments of the method disclosed herein.

[0058] Figure 3 This is a schematic diagram of a face recognition device with occlusion provided in an embodiment of this disclosure. Figure 3 As shown, the occluded facial recognition device includes:

[0059] The acquisition module 301 is configured to acquire the base database dataset, which contains multiple base database images, each of which is free from occlusion.

[0060] The generation module 302 is configured to generate occlusion for each base image, thereby obtaining an occlusion image corresponding to each base image;

[0061] Model module 303 is configured to extract the first feature of the face image, the second feature of each base image, and the third feature of the occluded image corresponding to each base image when a face image to be identified is detected using a face recognition model.

[0062] The calculation module 304 is configured to calculate the similarity between the first feature and the second feature of each base image and the third feature of the occluded image corresponding to the base image, respectively.

[0063] The determination module 305 is configured to determine the target database image with the highest similarity to the face image based on the similarity between the first feature and the second feature of each database image and the third feature of the occluded image corresponding to the database image, and to identify the face image as the target object corresponding to the target database image.

[0064] The residual neural network sequentially comprises a zero-stage network, a first-stage network, a second-stage network, a third-stage network, and a fourth-stage network. For example, a ResNet50 residual neural network includes a zero-stage network (Stage 0), a first-stage network (Stage 1), a second-stage network (Stage 2), a third-stage network (Stage 3), and a fourth-stage network (Stage 4). In this embodiment, the multiple stages of the residual neural network refer to the second-stage network, the third-stage network, and the fourth-stage network.

[0065] The base dataset consists of a large number of pre-stored face images, also known as the base image set. When a face image to be identified is detected, calculations show that the face image to be identified has the highest similarity to the target image in the base dataset. Therefore, the face image to be identified is determined to be the image of the target object corresponding to the target image in the base dataset. The face recognition model can be a residual neural network or similar model. Features and class centers can be understood as vectors or matrices, and the similarity can be cosine similarity. Occlusion is generated for each base image, simulating face images of people wearing masks, etc., i.e., occluded images.

[0066] According to the technical solution provided in this disclosure, a base database dataset is obtained, wherein the base database dataset contains multiple base database images, and each base database image is free from occlusion; an occlusion is generated for each base database image to obtain an occluded image corresponding to each base database image; when a face image to be identified is detected, a face recognition model is used to extract a first feature of the face image, a second feature of each base database image, and a third feature of the occluded image corresponding to each base database image; the similarity between the first feature and the second feature of each base database image and the third feature of the occluded image corresponding to that base database image is calculated respectively; based on the similarity between the first feature and the second feature of each base database image and the third feature of the occluded image corresponding to that base database image, a target base database image with the highest similarity to the face image is determined, and the face image is identified as the target object corresponding to the target base database image. Therefore, by adopting the above technical means, the problem of low accuracy in recognizing images with occlusion in face recognition models trained based on traditional model training methods in the prior art can be solved, thereby improving the accuracy of the model in recognizing images with occlusion.

[0067] Optionally, the model module 303 is further configured to acquire a training dataset, wherein the training dataset includes: multiple classes, each class being divided into occluded subclasses and non-occluded subclasses, each subclass containing multiple training samples; modifying the loss function of the face recognition model accordingly based on each class being divided into occluded and non-occluded subclasses; and training the face recognition model using the modified loss function based on the training dataset.

[0068] A class can be viewed as a person, and each person has multiple occluded training samples from the occluded subclass and multiple unoccluded training samples from the unoccluded subclass. The class center is the matrix or vector that best represents the features of the samples in that class or subclass. Modifying the loss function of the face recognition model based on whether each class is divided into occluded and unoccluded subclasses means that the modified loss function can comprehensively consider images from both occluded and unoccluded subclasses. Existing loss functions do not classify each class into occluded and unoccluded subclasses; that is, they do not distinguish between occluded and unoccluded images, but only process all images in a general way.

[0069] It should be noted that training samples are images. In this disclosure, the images used in training are referred to as training samples, and the images used in recognition are referred to as images. In fact, training samples are images.

[0070] Optionally, model module 303 is also configured to modify the target formula for calculating features and class centers based on each class being divided into occluded and non-occluded subclasses, so that the modified target formula takes into account both cases where there is occlusion and no occlusion in each class, and each subclass corresponds to a class center; and modify the loss function using the modified target formula.

[0071] In one alternative embodiment, it includes:

[0072] The revised target formula is

[0073]

[0074] The modified loss function is

[0075]

[0076] Among them, W jk Let x be the class center of the k-th subclass of the j-th class, T be the transpose symbol, and x be the class center of the k-th subclass of the j-th class. i Let θ be the feature of the i-th training sample. j ′ for x i With W jk The angle between them, W jk Not x i The corresponding class center, For x i With W ik The angle between them, W ik It is x i The corresponding class centers are s, where s is the feature scaling parameter, m is the additive margin parameter, and N is the total number of training samples.

[0077] calculate The target formula will also be used; you only need to use W. ik Replace W jk The calculated result is When k is 0, the 0th subclass is the occluded subclass; when k is 1, the 1st subclass is the non-occluded subclass. The feature scaling parameter and the additive margin parameter can both be adjusted in advance. The max() function takes the maximum of the two or more parameters; exp is an exponential function with the natural constant e as its base; and log is the sign of the logarithmic function.

[0078] In one alternative embodiment, it includes:

[0079] The original target formula was:

[0080]

[0081] The loss function before modification was

[0082]

[0083] W j Let θ be the class center of the j-th class. j For x i With W j The angle between them, W j Not x i The corresponding class center, For x i With W i The angle between them.

[0084] Optionally, the determining module 305 is further configured to take the largest similarity between the first feature and the second feature of each base image and the first feature and the third feature of the occluded image corresponding to the base image as the similarity between the face image and the base image; and determine the target base image with the highest similarity to the face image based on the similarity between the face image and each base image.

[0085] Optionally, the calculation module 304 is also configured to calculate the similarity between the face image and each corresponding image in the database using the following formula:

[0086] s1 = cos < f1, f 2,1 >

[0087] s2 = cos < f1, f 2,2 >

[0088] s = max(s1, s2)

[0089] f1 is the first feature, f 2,1 f is the second feature of each base image. 2,2 Let s1 be the third feature of the occluded image corresponding to the base image, s2 be the similarity between the first feature and the second feature of the base image, s2 be the similarity between the first feature and the third feature of the occluded image corresponding to the base image, and s be the similarity between the face image and the base image.

[0090] Optionally, the generation module 302 is also configured to generate occlusion for each base image using the UVMap algorithm, thereby obtaining an occlusion image corresponding to each base image.

[0091] The UVMap algorithm, also known as UV mapping, precisely maps each point on an image to the surface of a model object. Software then performs smooth interpolation at the gaps between these points. This embodiment utilizes the UVMap algorithm to occlude certain features in an image.

[0092] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this disclosure.

[0093] Figure 4 This is a schematic diagram of the electronic device 4 provided in an embodiment of this disclosure. Figure 4 As shown, the electronic device 4 of this embodiment includes: a processor 401, a memory 402, and a computer program 403 stored in the memory 402 and executable on the processor 401. When the processor 401 executes the computer program 403, it implements the steps in the various method embodiments described above. Alternatively, when the processor 401 executes the computer program 403, it implements the functions of each module / unit in the various device embodiments described above.

[0094] Electronic device 4 can be a desktop computer, laptop, handheld computer, cloud server, or other electronic device. Electronic device 4 may include, but is not limited to, processor 401 and memory 402. Those skilled in the art will understand that... Figure 4 This is merely an example of electronic device 4 and does not constitute a limitation on electronic device 4. It may include more or fewer components than shown, or different components.

[0095] The processor 401 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.

[0096] The memory 402 can be an internal storage unit of the electronic device 4, such as a hard disk or RAM of the electronic device 4. The memory 402 can also be an external storage device of the electronic device 4, such as a plug-in hard disk, Smart Media Card (SMC), Secure Digital (SD) card, Flash Card, etc., equipped on the electronic device 4. The memory 402 can also include both internal and external storage units of the electronic device 4. The memory 402 is used to store computer programs and other programs and data required by the electronic device.

[0097] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0098] If an integrated module / unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments can also be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program may include computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. A computer-readable medium may include: any entity or device capable of carrying computer program code, recording media, USB flash drives, portable hard drives, magnetic disks, optical disks, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content included in a computer-readable medium may be appropriately added to or subtracted according to the requirements of legislation and patent practice in a jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer-readable media may not include electrical carrier signals and telecommunication signals.

[0099] The above embodiments are only used to illustrate the technical solutions of this disclosure, and are not intended to limit it. Although this disclosure has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this disclosure, and should all be included within the protection scope of this disclosure.

Claims

1. A face recognition method with occlusion, characterized in that, include: Obtain the base database dataset, wherein the base database dataset contains multiple base database images, and there are no occlusions on each base database image; Generate occlusion for each base image, and obtain the occlusion image corresponding to each base image; When a face image to be identified is detected, the face recognition model is used to extract the first feature of the face image, the second feature of each base image, and the third feature of the occluded image corresponding to each base image. Calculate the similarity between the first feature and the second feature of each base image, as well as the third feature of the occluded image corresponding to that base image. Based on the similarity between the first feature and the second feature of each base image and the third feature of the occluded image corresponding to the base image, the target base image with the highest similarity to the face image is determined, and the face image is identified as the target object corresponding to the target base image. Before extracting the first feature of the face image, the second feature of each base image, and the third feature of the occluded image corresponding to each base image using a face recognition model when a face image to be identified is detected, the method further includes: Obtain a training dataset, wherein the training dataset includes: multiple classes, each class is divided into occluded subclasses and non-occluded subclasses, and each subclass contains multiple training samples; Based on the fact that each class is divided into occluded subclasses and non-occluded subclasses, the loss function of the face recognition model is modified accordingly; Based on the training dataset, the face recognition model is trained using the modified loss function. The loss function of the face recognition model is modified accordingly based on the fact that each class is divided into occluded and non-occluded subclasses, including: Based on the division of each class into occluded and non-occluded subclasses, the target for calculating features and class centers is modified. The formula ensures that the modified target formula takes into account both cases where there is occlusion and cases where there is no occlusion in each class, and each subclass corresponds to a class center; The loss function is modified using the revised objective formula; The revised target formula is The modified loss function is Among them, W jk Let x be the class center of the k-th subclass of the j-th class, T be the transpose symbol, and x be the class center of the k-th subclass of the j-th class. i For the features of the i-th training sample, For x i With W jk The angle between them, W jk Not x i The corresponding class center, For x i With W ik The angle between them, W ik It is x i The corresponding class centers are s, where s is the feature scaling parameter, m is the additive margin parameter, and N is the total number of training samples.

2. The method according to claim 1, characterized in that, The step of determining the target database image with the highest similarity to the face image based on the similarity between the first feature and the second feature of each database image, as well as the third feature of the corresponding occluded image in the database image, includes: The similarity between the first feature and the second feature of each database image and the first feature and the third feature of the occluded image corresponding to the database image is the largest of the two values. Based on the similarity between the face image and each corresponding image in the database, the target image in the database with the highest similarity to the face image is determined.

3. The method according to claim 2, characterized in that, The similarity between the face image and each image in the database is calculated using the following formula: f1 is the first feature, f 2,1 f is the second feature of each base image. 2,2 s1 is the third feature of the occluded image corresponding to the base image, s2 is the similarity between the first feature and the second feature of the base image, s3 is the similarity between the first feature and the third feature of the occluded image corresponding to the base image, and s4 is the similarity between the face image and the base image.

4. The method according to claim 1, characterized in that, include: The UVMap algorithm is used to generate occlusions for each base image, resulting in an occlusion image for each base image.

5. A face recognition device with occlusion, used to perform the method according to any one of claims 1-4, characterized in that, include: The acquisition module is configured to acquire a base database dataset, wherein the base database dataset contains multiple base database images, and there are no occlusions on each base database image; The generation module is configured to generate occlusion for each base image, resulting in an occlusion image for each base image. The model module is configured to extract the first feature of the face image, the second feature of each base image, and the third feature of the occluded image corresponding to each base image when a face image to be identified is detected using a face recognition model. The calculation module is configured to calculate the similarity between the first feature and the second feature of each base image and the third feature of the occluded image corresponding to the base image, respectively. The determination module is configured to determine the target database image with the highest similarity to the face image based on the similarity between the first feature and the second feature of each database image and the third feature of the occluded image corresponding to the database image, and to identify the face image as the target object corresponding to the target database image.

6. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method as described in any one of claims 1 to 4.

7. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method as described in any one of claims 1 to 4.

Citation Information

Patent Citations

  • Face recognition method and device, terminal equipment and medium

    CN111597910A