Face recognition method and device

By using the reconstruction model in the face recognition system to encode the background area picture and determine whether the face area is a living area, the complex problem of feature recognition in static detection in the prior art is solved, and the effect of simplifying the process and reducing costs is achieved.

CN113569806BActive Publication Date: 2025-05-16ZHEJIANG DAHUA TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110950403.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-08-18
Publication Date
2025-05-16
Estimated Expiration
2041-08-18

AI Technical Summary

Technical Problem

The existing facial recognition system needs to accurately identify subtle features of the face or background in static detection, resulting in complex technology implementation, high computing power consumption and high cost, and it is impossible to directly expand and upgrade on the original system.

Method used

By obtaining the target image collected by the face recognition device, determining the background area picture in the target image, and inputting it into the reconstruction model, obtaining the target encoding corresponding to the background area picture. Based on the target encoding and background encoding, it is determined whether the face area in the target picture is a living area, and then face recognition is performed.

Benefits of technology

It reduces the data processing volume, simplifies the identification process, reduces the difficulty of comparatively, solves the problem of accurate identification of subtle features in static detection, and realizes direct expansion and upgrading on the original system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113569806B_ABST
    Figure CN113569806B_ABST
Patent Text Reader

Abstract

The present invention discloses a face recognition method and device. The method comprises: obtaining a target image with a face to be recognized collected by a face recognition device; determining a background area image in the target image; inputting the background area image into a reconstruction model to obtain a target code corresponding to the background area image, wherein the reconstruction model is trained by multiple sets of training data; determining whether the face area in the target image is a living area or a non-living area according to the target code and the background code of the face recognition device, wherein the background code is a background code corresponding to the background image of the area where the face recognition device is collected; when the face area in the target image is a living area, face recognition is performed according to the face area. The present invention solves the technical problem in the related art that static detection of face recognition using photos or videos requires accurate recognition of subtle features of the face or background, which is complex and difficult to achieve.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data security protocols, and in particular to a face recognition method and device. Background Art

[0002] At present, face recognition is increasingly widely used in access control, mobile payment, device authentication, etc., and it faces more and more threats. Among them, the non-live attack method of using photos or videos to deceive the face recognition system is a common one. Liveness detection or non-liveness attack detection technology can be divided into dynamic detection and static detection. The implementation principle of dynamic detection is to make the test subject make corresponding actions according to the random prompts of the detection system, and judge whether it is a non-liveness attack by judging whether the action is accurate; static detection does not require the test subject to make special actions. Most of them use facial detail feature analysis and extraction, optical flow analysis, stereo feature analysis (binocular camera or multiple cameras), biometric feature analysis (such as skin elasticity, heat map, etc.) and different combinations or improved methods of different methods. The basic principle is to extract the detailed features of the face as much as possible, and compare the facial features in the video with the facial features in the database to determine whether the current target to be tested is forged; on the other hand, through the assistance of additional equipment, such as using binocular cameras, multiple cameras, thermal imaging cameras or fill-in lighting devices, the biometric features of the current target to be tested can be detected as accurately as possible to determine whether there is a non-liveness attack.

[0003] In existing face recognition systems, the technologies for liveness detection or non-liveness attack detection can be summarized as dynamic detection technology and static detection technology. Among them, dynamic detection technology requires users to make corresponding facial expressions or movements according to the random prompts of the system. Although it is very secure, the user experience is poor; and in existing static detection technology, for plane face recognition, it is necessary to accurately identify the subtle features of the face or background, and then it is necessary to build a complex face feature model, a full background model or extract iris features, etc. The technical implementation is relatively complex and consumes a lot of computing power; for three-dimensional face recognition, it is necessary to use binocular cameras, multiple cameras, fill light equipment or ultrasonic ranging, thermal imaging and other additional hardware support, the technical implementation is also relatively complex, and the cost is high, and it cannot be directly expanded and upgraded on the original face recognition system.

[0004] A related technology provides a liveness detection method for face recognition, which uses the average frame method for background modeling, that is, using more than 10 background images for denoising and grayscale processing, and calculating parameters such as mean and variance as the benchmark for structural similarity testing. Since moving targets will inevitably appear in the background and the background light will also change, the above factors will affect the detection accuracy of this method; on this basis, if a strategy of real-time updating of the background image is adopted, it is necessary to calculate the mean and variance of more than 10 full background images in real time, which consumes a lot of computing power; in addition, the method proposed in this scheme also requires the assistance of ultrasonic ranging equipment, which is not conducive to direct development and deployment on the original equipment.

[0005] In another face recognition liveness detection method based on facial feature points and optical flow field, the movement direction of the real face at the facial feature points is obviously inconsistent with the overall movement direction of the face. It has a good detection effect on non-liveness attacks using photos, but is not suitable for the detection of non-liveness attacks using videos.

[0006] There is also a method of liveness detection for face recognition, which requires a large number of auxiliary equipment, such as iris sensors, infrared temperature sensors, and electronic behavior sensors. In addition, users are required to operate the interface and perform prescribed actions such as blinking and shaking their heads. Although it is very safe, on the one hand, the user experience is poor and the detection efficiency is low; on the other hand, it is not conducive to expansion and upgrading on low-end face detection equipment.

[0007] To address the above-mentioned problems, no effective solution has been proposed yet. Summary of the invention

[0008] The embodiments of the present invention provide a face recognition method and device to at least solve the technical problem that static detection of face recognition using photos or videos in the related art requires accurate recognition of subtle features of the face or background, which is complex and difficult to implement.

[0009] According to one aspect of an embodiment of the present invention, a face recognition method is provided, comprising: obtaining a target image with a face to be recognized, which is collected by a face recognition device; determining a background area image in the target image; inputting the background area image into a reconstruction model to obtain a target code corresponding to the background area image, wherein the reconstruction model comprises a multi-layer coding network trained by multiple groups of training data, each group of training data comprising a background image and a code corresponding to the background image; determining whether the face area in the target image is a living area or a non-living area based on the target code and the background code of the face recognition device, wherein the background code is a background code corresponding to the background image of the area collected by the face recognition device; and when the face area in the target image is a living area, performing face recognition based on the face area.

[0010] Optionally, before determining whether the face area in the target image is a living area or a non-living area based on the target code and the background code of the face recognition device, the method further includes: using a photo of the background of the area captured when the face recognition device is in an idle state as the background image; and inputting the background image into the reconstruction model to obtain a background code corresponding to the background image.

[0011] Optionally, before inputting the background image into the reconstruction model to obtain the background code corresponding to the background image, it also includes: obtaining multiple background training images without a first motion detection area, wherein the first motion detection area is an area in the background training image where motion is detected; training the reconstruction model according to the background training image; inputting the background training image into the reconstruction model, and the reconstruction model outputs the corresponding training code; decoding according to the training code to obtain a reconstructed image, and determining the error between the background training image and the reconstructed image; when the error meets an error threshold, or the number of training times reaches a preset number, determining that the reconstruction model training is completed.

[0012] Optionally, before inputting the background image into the reconstruction model to obtain the background code corresponding to the background image, it also includes: deleting the second motion detection area in the background image, wherein the second motion detection area is the area in the background image where motion is detected; preprocessing the deleted background image, wherein the preprocessing includes denoising and normalization; inputting the background image into the reconstruction model to obtain the background code corresponding to the background image includes: inputting the processed background image into the reconstruction model to obtain the code corresponding to the background image as the background code.

[0013] Optionally, determining the background area picture in the target picture also includes: identifying the face area in the target picture, deleting the face area in the target picture to obtain a non-face area; determining a third motion detection area in the non-face area, wherein the third motion detection area is the motion detection area in the target picture; deleting the third motion detection area to obtain a selected background area; determining an image in the selected background area within a predetermined range from the face area as the background area picture.

[0014] Optionally, deleting the third motion detection area to obtain the background area to be selected includes: when the position of the third motion detection area coincides with the position of the second motion detection area of ​​the background image, deleting the third motion detection area to obtain the background area to be selected in the target image; when the position of the third motion detection area coincides with the position of the non-second motion detection area in the background image, replacing the image of the third motion detection area with the corresponding background image in the background image to obtain the background area to be selected in the target image.

[0015] Optionally, inputting the background area picture into a reconstruction model to obtain a target code corresponding to the background area picture includes: dividing the background area picture into multiple unit areas; using each of the multiple unit areas as an input to the reconstruction model, and having the reconstruction model output a reconstructed code of each unit area; and generating the target code according to the reconstruction code of each unit area.

[0016] Optionally, determining whether the face area in the target image is a living area or a non-living area based on the target code and the background code of the face recognition device includes: determining background reconstruction codes in the background image corresponding to multiple unit areas of the background area image, wherein the background image is divided into multiple unit areas in the same way when generating the corresponding background code according to the reconstruction model, and determining background reconstruction codes corresponding to the multiple unit areas; calculating reconstruction codes of multiple unit areas in the background area image and error parameters of the corresponding background reconstruction codes, wherein the error parameters include likelihood ratio and cumulative error; when the error parameters all meet preset error conditions, determining that the face area in the target image is a living area; when the error parameters all do not meet the preset error conditions, determining that the face area in the target image is a non-living area.

[0017] According to another aspect of an embodiment of the present invention, a face recognition device is also provided, including: an acquisition module, used to acquire a target image with a face to be recognized collected by the face recognition device; a first determination module, used to determine a background area image in the target image; a reconstruction module, used to input the background area image into a reconstruction model to obtain a target code corresponding to the background area image, wherein the reconstruction model includes a multi-layer encoding network trained by multiple groups of training data, each group of training data includes a background image and a code corresponding to the background image; a second determination module, used to determine whether the face area in the target image is a living area or a non-living area based on the target code and the background code of the face recognition device, wherein the background code is the background code corresponding to the background image of the area where the face recognition device is collected; an identification module, used to perform face recognition based on the face area when the face area in the target image is a living area.

[0018] According to another aspect of an embodiment of the present invention, a processor is further provided, wherein the processor is used to run a program, wherein the program executes any one of the above-mentioned face recognition methods when running.

[0019] According to another aspect of an embodiment of the present invention, a computer storage medium is further provided, wherein the computer storage medium includes a stored program, wherein when the program is executed, the device where the computer storage medium is located is controlled to execute any one of the above-mentioned face recognition methods.

[0020] In an embodiment of the present invention, a target image with a face to be recognized collected by a face recognition device is obtained; a background area image in the target image is determined; the background area image is input into a reconstruction model to obtain a target code corresponding to the background area image, wherein the reconstruction model includes a multi-layer coding network trained by multiple groups of training data, each group of training data includes a background image and a code corresponding to the background image; based on the target code and the background code of the face recognition device, it is determined whether the face area in the target image is a living area or a non-living area, wherein the background code is a background code corresponding to the background image of the area where the face recognition device is collected; when the face area in the target image is a living area, face recognition is performed based on the face area. When comparing the background code of the background image with the target code of the background area of ​​the target image of the face to be identified, if the face to be identified in the target image is a living face, the background code and the target code should be consistent. If the face to be identified in the target image is a photo or video, the background around the face is inconsistent with the real background. The target code and background code of the target image are used to determine whether the face in the target image is alive, thereby achieving the technical effect of reducing the amount of data processing, simplifying the recognition process, and reducing the difficulty of comparison. This solves the technical problem of static detection of face recognition using photos or videos in related technologies, which requires accurate identification of subtle features of the face or background, and is complex and difficult to achieve. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of this application. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:

[0022] Figure 1 is a flow chart of a face recognition method according to an embodiment of the present invention;

[0023] Figure 2 is a schematic diagram of an identification system architecture according to an embodiment of the present invention;

[0024] Figure 3 is a schematic diagram of a SAE model according to an embodiment of the present invention;

[0025] Figure 4 is a flow chart of a face detection method according to an embodiment of the present invention;

[0026] Figure 5 is a schematic diagram of a face recognition device according to an embodiment of the present invention. DETAILED DESCRIPTION

[0027] In order to enable those skilled in the art to better understand the scheme of the present invention, the technical scheme in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present invention.

[0028] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0029] According to an embodiment of the present invention, a method embodiment of a face recognition method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0030] Figure 1 is a flow chart of a face recognition method according to an embodiment of the present invention. Figure 1 As shown, the method comprises the following steps:

[0031] Step S102, obtaining a target image with a face to be recognized collected by a face recognition device;

[0032] Step S104, determining the background area image in the target image;

[0033] Step S106, inputting the background area image into a reconstruction model to obtain a target code corresponding to the background area image, wherein the reconstruction model includes a multi-layer coding network trained by multiple sets of training data, each set of training data includes a background image and a code corresponding to the background image;

[0034] Step S108, determining whether the face area in the target image is a living area or a non-living area according to the target code and the background code of the face recognition device, wherein the background code is the background code corresponding to the background image of the area where the face recognition device is collected;

[0035] Step S110: When the face region in the target image is a living body region, face recognition is performed based on the face region.

[0036] Through the above steps, a target image with a face to be recognized collected by a face recognition device is obtained; a background area image in the target image is determined; the background area image is input into a reconstruction model to obtain a target code corresponding to the background area image, wherein the reconstruction model includes a multi-layer coding network trained by multiple groups of training data, each group of training data includes a background image and a code corresponding to the background image; according to the target code and the background code of the face recognition device, it is determined whether the face area in the target image is a living area or a non-living area, wherein the background code is a background code corresponding to the background image of the area where the face recognition device is collected; when the face area in the target image is a living area, face recognition is performed according to the face area. When comparing the background code of the background image with the target code of the background area of ​​the target image of the face to be identified, if the face to be identified in the target image is a living face, the background code and the target code should be consistent. If the face to be identified in the target image is a photo or video, the background around the face is inconsistent with the real background. The target code and background code of the target image are used to determine whether the face in the target image is alive, thereby achieving the technical effect of reducing the amount of data processing, simplifying the recognition process, and reducing the difficulty of comparison. This solves the technical problem of static detection of face recognition using photos or videos in related technologies, which requires accurate identification of subtle features of the face or background, and is complex and difficult to achieve.

[0037] The above-mentioned face recognition device may include a camera device for taking a photo of the face to be recognized. In static face recognition, that is, a recognition method that does not require the face to be verified according to a predetermined action, the predetermined action may be nodding or turning the head, etc. The above-mentioned static recognition is often easily deceived by faces in photos or videos. In the related art, it is usually necessary to use external ultrasonic equipment or infrared detection equipment to assist in recognition, or to recognize subtle features of the face. It has the problems of high cost, complex implementation, and difficult algorithm.

[0038] This embodiment determines the background area image in the target image, reconstructs the background area image, obtains the target code, compares the target code with the background code of the real background image, and determines whether the face area is alive. The actual principle is as follows: during the recognition process of a real living face, its target image will only include the face area and the real background area outside the face area. However, when attacking and identifying through pictures or videos, a small part of the background closely adjacent to the face area outside the recognized face area is not actually a real background image, but an area outside the face area in the picture or video used to carry the face information. When it is relative to the live face recognition, the background of the area near the face contour will be completely different from the real background, and its corresponding code is also completely different. Therefore, the target image is recognized alive through the background code and the target code.

[0039] During recognition, first obtain the target image collected by the face recognition device, and determine the background area image in the target image. The background area image may be the area in the target image other than the face area. In this embodiment, it may be the area at a predetermined distance outside the outline of the face area in the target image. Since the area far from the face area is generally more likely to be a real background image, during static recognition, it is only necessary to determine whether the background near the face outline is the real background, and then determine whether the face is the face in the image or video. Selecting the area at a predetermined distance outside the outline of the above-mentioned face as the background area image for comparison can not only reduce the amount of data and improve data processing efficiency, but also improve the accuracy of the determination to a certain extent.

[0040] The above background area picture is input into the reconstruction model to obtain the target code corresponding to the background area picture. The purpose is to obtain the background features in the background area picture. Since the above target picture and background picture are collected at different times, the light and sensitivity of the external environment will affect the picture. It is easy to make mistakes by directly comparing the pixels of the image. Therefore, the reconstruction model is used to determine the code corresponding to the image as the image feature for comparison, which can largely avoid the probability of error.

[0041] The reconstruction model may be a SAE model, including a multi-layer encoding network and a multi-layer decoding network. The reconstruction model is trained by multiple sets of training data, each set of training data including a background image and a code corresponding to the background image. Each layer of the encoding network of the reconstruction model can output a code for the background image. In this embodiment, only the code output by the last layer of the encoding network is used to reduce the model complexity and increase the accuracy.

[0042] According to the target code and the background code of the face recognition device, it is determined whether the face area in the target image is a living area or a non-living area, that is, whether the image features in the target code and the background code are consistent. If they are consistent, it means that the area at a predetermined distance outside the face area in the target image is consistent with the real background, that is, the face is alive. Otherwise, if they are inconsistent, it means that the area at a predetermined distance outside the face area in the target image is inconsistent with the real background, that is, the face is non-living.

[0043] It should be noted that the above-mentioned background picture is obtained by taking a photo of the background before the face recognition device obtains the above-mentioned target picture, in the idle time when there is no face recognition. The background code is the background code corresponding to the background picture of the area where the face recognition device is collecting, that is, the background code obtained by reconstructing the background picture through the reconstruction model.

[0044] When the face area of ​​the target image is a living area, face recognition is performed based on the face area, and the recognition method can be an existing technology, specifically, it can be recognized through a deep learning network model or a machine learning network model. Identify the feature information of the face, or use the feature information to match whether the face is a face with permission.

[0045] Optionally, before determining whether the face area in the target image is a living area or a non-living area based on the target code and the background code of the face recognition device, the method also includes: taking a photo of the background of the area captured when the face recognition device is in an idle state as a background image; and inputting the background image into a reconstruction model to obtain a background code corresponding to the background image.

[0046] The above-mentioned face recognition device collects a photo of the background of the area in an idle state as a background picture. Considering that the real background changes with time and its light changes to a certain extent, the background picture can be obtained at a predetermined frequency to avoid a large difference in the collection time of the background picture and the target picture, which leads to changes in their light intensity or background, and then causes errors during comparison, thereby reducing the probability of errors and further improving the accuracy of face liveness detection.

[0047] Optionally, before inputting the background image into the reconstruction model to obtain the background code corresponding to the background image, it also includes: obtaining multiple background training images without the first motion detection area, wherein the first motion detection area is the area in the background training image where motion is detected; training the reconstruction model according to the background training images; inputting the background training images into the reconstruction model, and the reconstruction model outputs the corresponding training code; decoding according to the training code to obtain the reconstructed image, and determining the error between the background training image and the reconstructed image; when the error meets the error threshold, or the number of training times reaches a preset number, determining that the reconstruction model training is completed.

[0048] Before inputting the background image into the reconstruction model and obtaining the background code corresponding to the background image, the reconstruction model needs to be created and trained to ensure that the reconstruction model has a high accuracy. During training, considering that the moving objects in the background image are unlikely to exist in the target image during face recognition, in order to avoid dynamic objects in the background training image, such as moving people or equipment, from interfering with the background image, the background image is dynamically monitored to determine the first motion detection area in the background training image, that is, the image area of ​​the moving person or equipment mentioned above, and delete it. It is also possible to obtain multiple background training images without the first motion detection area, or directly obtain background training images with the first motion detection area, and delete the first motion detection area to improve the effectiveness of the background training images and improve the effect and accuracy of the reconstruction model training.

[0049] During training, the first motion detection area in the above-mentioned background training picture can be deleted and preprocessed, for example, denoising and normalization. Then the background training picture is divided into M*N small picture units, the image coordinates and number of each small picture unit are recorded, the M*N small picture units are input into the reconstruction model, each small picture unit is reconstructed to obtain the reconstruction code, and the reconstruction code is decoded by the decoding network corresponding to the encoding network of the reconstruction model to obtain the reconstruction unit of each small picture unit. According to the image of the small picture unit and the image of the reconstruction unit, the reconstruction error of the reconstruction of the small picture unit can be calculated, and according to the reconstruction error in the reconstruction process of multiple small picture units, the cumulative reconstruction error of the background training image can be calculated. According to the loss of the cumulative reconstruction error sum of squares and the sparse penalty term of the reconstruction model's own encoding, a loss function is constructed to perform iterative training of the parameters of the reconstruction model. When the error between the training code output by the self-encoding and the reconstruction model training meets the preset error threshold, it is considered that the accuracy of the reconstruction model meets the use requirements, and the training of the reconstruction model is completed.

[0050] Optionally, before inputting the background image into the reconstruction model to obtain the background code corresponding to the background image, the method further includes: deleting the second motion detection area in the background image, wherein the second motion detection area is the area in the background image where motion is detected; preprocessing the deleted background image, wherein the preprocessing includes denoising and normalization; inputting the background image into the reconstruction model to obtain the background code corresponding to the background image includes: inputting the processed background image into the reconstruction model to obtain the code corresponding to the background image as the background code.

[0051] When the background image is input into the reconstruction model to obtain the background code, the second motion detection area of ​​the background image also needs to be deleted to avoid the interference of the second motion detection area on the result of the liveness detection. Its preprocessing, including denoising and normalization, is also to improve the image quality, reduce the error of the background image itself, and improve the accuracy of the reconstruction model in extracting image features to avoid errors in detection and recognition.

[0052] Then the second motion detection area is deleted, and the pre-processed background image is input into the reconstruction model, and the reconstruction model outputs the background code corresponding to the background image.

[0053] Optionally, determining the background area picture in the target picture also includes: identifying the face area in the target picture, deleting the face area in the target picture to obtain a non-face area; determining a third motion detection area in the non-face area, wherein the third motion detection area is the motion detection area in the target picture; deleting the third motion detection area to obtain a background area to be selected; determining an image within a predetermined range from the face area in the background area to be selected as the background area picture.

[0054] Similar to the background image, when the target image determines the background area image, the face area is first removed, and then the non-face area is detected to see if there is a third motion detection area. The third motion detection area is deleted to avoid the situation where the third motion detection area just happens to exist around the face contour. This will cause the comparison to think that the background area image is different from the actual background image, and the face is considered non-liveable, which will result in misjudgment and reduce the accuracy of liveness detection.

[0055] The selected background area images of the face area and the third motion detection area are taken out, and the above background area images are determined according to the contour of the face area. It should be noted that the above background area images can also be directly input into the reconstruction model and compared with the background code of the background image. However, the selected background area images not only have a large range, but also have the problems of large data volume and slow processing efficiency. Moreover, regardless of whether the face is alive or not, the background far away from the face area in the selected background image is generally the same as the real background. There is no need to compare them, which will not only increase the amount of data, but also have no effect on liveness recognition.

[0056] Optionally, deleting the third motion detection area to obtain the background area to be selected includes: when the position of the third motion detection area coincides with the position of the second motion detection area of ​​the background image, deleting the third motion detection area to obtain the background area to be selected in the target image; when the position of the third motion detection area coincides with the position of the non-second motion detection area in the background image, replacing the image of the third motion detection area with the corresponding background image in the background image to obtain the background area to be selected in the target image.

[0057] When deleting the third motion detection area, if the second motion detection area exists in the background image and the position of the third motion detection area is the same, it means that the position is not only deleted in the background image, but also in the target image, so it is directly deleted without comparing the area. When comparing the third motion detection area with the second motion detection area, it can be determined in units of image pixels or divided small image units. Specifically, if the pixel points or small image units corresponding to the third motion detection area overlap with the pixel points or small image units corresponding to the second motion detection area, the overlapping pixel points or small image units of the third motion detection area are directly deleted.

[0058] When the third motion detection area coincides with the non-second motion detection area in the background picture, it means that there is a real background in the background picture of that position, and the position in the target picture is blocked by the third motion detection area, and the actual background of that position cannot be determined. The third motion detection area in the target picture can be replaced by the area of ​​the background picture of that position to obtain the real background of the third motion detection area in the target picture.

[0059] Optionally, inputting the background area image into the reconstruction model to obtain the target code corresponding to the background area image includes: dividing the background area image into multiple unit areas; using each unit area of ​​the multiple unit areas as an input to the reconstruction model, and the reconstruction model outputs the reconstructed code of each unit area; and generating the target code according to the reconstruction code of each unit area.

[0060] The unit area can be divided in a preset manner, for example, divided into M*N small image units in a grid manner, and the small image unit is also the unit area. The images of multiple unit areas are respectively input into the reconstruction model, and the reconstruction model is as follows: Figure 3 As shown, multiple inputs may be included. The reconstruction model outputs reconstruction codes corresponding to multiple unit regions respectively, and the multiple reconstruction codes of the multiple unit regions are combined to obtain the above target code.

[0061] It should be noted that when the reconstruction model is trained with background training pictures, and when the background pictures are input into the reconstruction model to obtain the background code, it is also executed according to similar steps as above. It is first divided into multiple unit areas, and the images of the multiple unit areas are input into the reconstruction model. The reconstruction model outputs the reconstruction codes of the multiple unit areas, which are combined to obtain the corresponding training code or background code.

[0062] When dividing multiple unit areas, the coordinates and numbers of the unit areas in the original image are recorded. When the reconstructed codes of the multiple unit areas are subsequently combined, the reconstructed codes corresponding to the unit areas are combined according to the coordinates and numbers of the unit areas in the original image.

[0063] Optionally, determining whether the face area in the target image is a living area or a non-living area based on the target code and the background code of the face recognition device includes: determining the background reconstruction code in the background image corresponding to multiple unit areas of the background area image, wherein the background image is divided into multiple unit areas in the same way when generating the corresponding background code according to the reconstruction model, and determining the background reconstruction codes corresponding to the multiple unit areas; calculating the error parameters between the reconstruction codes of the multiple unit areas in the background area image and the corresponding background reconstruction codes, wherein the error parameters include a likelihood ratio and an accumulated error; when the error parameters all meet the preset error conditions, determining that the face area in the target image is a living area; when the error parameters all do not meet the preset error conditions, determining that the face area in the target image is a non-living area.

[0064] When the error parameters all meet the preset error conditions, it is considered that the background reconstruction code of the background image is the same as the reconstruction code corresponding to the unit area, that is, the unit area in the background area image is an image of the real background, and the face area in the target image can be determined to be a living area. When the error parameters all do not meet the preset error conditions, it is considered that the background reconstruction code of the background image is different from the reconstruction code corresponding to the unit area, that is, the unit area in the background area image is not an image of the real background, and the face area in the target image can be determined to be a non-living area. Thus, according to the target code and the background code of the face recognition device, it is determined that the face area in the target image is a living area or a non-living area.

[0065] It should be noted that the embodiment of the present application also provides an optional implementation, which is described in detail below.

[0066] This embodiment provides a real-time anti-counterfeiting detection system solution for face recognition based on SAE network local background reconstruction, which not only ensures the accuracy and real-time performance of detection, but also saves computing power, and can be effectively developed and deployed on low-end devices or existing face recognition systems.

[0067] The non-living attack detection system designed in this embodiment relies on the face recognition system and the motion detection system. The information interaction between the systems is as follows: Figure 2 As shown, Figure 2 is a schematic diagram of an identification system architecture according to an embodiment of the present invention.

[0068] The motion detection system detects the moving area in the background in real time and sends its coordinates to the liveness detection system. The detection system removes it when reconstructing the local background to reduce the false alarm rate of non-liveness attacks; the face recognition system sends the coordinates of the detected face area to the liveness detection system. The detection system obtains the background area to be detected at the face boundary based on the coordinates, reconstructs it and detects it, and gives the detection result; finally, the face recognition system gives prompts such as recognition success or failure based on the returned detection result, and links the corresponding linkage items.

[0069] The specific implementation of the liveness detection system includes the training process of the SAE reconstruction model and the real-time detection process. For the model training process, the specific steps are as follows:

[0070] (1) Data preparation and preprocessing: After the system is initialized, the on-site video data is collected first, and background data without abnormalities is selected, that is, the above-mentioned background training pictures. Each background training picture is divided into M*N small picture units, and the image coordinates and number of each small picture unit are recorded. The data is denoised and normalized.

[0071] (2) Model construction and training: Figure 3 is a schematic diagram of a SAE model according to an embodiment of the present invention, constructed as follows Figure 3 The SAE network shown uses the data processed in step (1) to perform end-to-end training, that is, taking each small image unit as a unit, the reconstruction model reconstructs each frame of input data (M*N small images), and constructs a loss function by accumulating the loss of the sum of squares of the reconstruction error and the sparse penalty of the autoencoder. The loss function is shown in formula (1-1), and the network parameters are updated until the overall training loss is less than the predefined threshold or the number of iterations reaches the set value.

[0072] L(Y, W, b, W′, b′)=||X′-X|| 2 +γ∑ i |Y i |

[0073] =||δ′(W′δ(WX+b)+b′)-X|| 2 +γ∑ i |Y i | (1-1)

[0074] Among them, the corresponding encoding network Y = δ (WX + b) and the decoding network X' = δ' (W'Y + b'), X is the network input, Y iis the network sparse coding output of the i-th neuron, i is the number of the neuron, W and W′ are the weight coefficients between the input to the hidden layer code and the hidden layer code to the decoding output, b and b′ are the corresponding bias terms, δ and δ′ are the activation functions of the encoding network and the decoding network, and γ is the penalty coefficient of the linear regularization sparse penalty restriction term.

[0075] Figure 4 is a flow chart of a face detection method according to an embodiment of the present invention. Figure 4 As shown, for the real-time detection process, the specific implementation steps are as follows:

[0076] (1) Obtaining a background image and performing data preprocessing on the background image, that is, performing denoising and normalization processing on the real-time collected video frame data and feeding it into the reconstruction model;

[0077] (2) Reconstruct the background image, that is, reconstruct the data of the M*N areas that do not trigger motion detection, and calculate the reconstruction error mean, variance and error accumulation sum of multiple small image units respectively;

[0078] (3) Obtaining a target image, after detecting a face in the target image, obtaining a background area image of the face boundary (in small images) according to the face range, respectively eliminating the third motion detection area in the background area image where motion detection exists according to the time of step (2) and the current motion detection information, and then reconstructing the local background of the remaining area, and calculating the reconstruction error mean, variance and error accumulation sum of multiple small image units;

[0079] (4) Combine the statistics obtained in step (2) to detect and judge the statistical value obtained in step (3), that is, calculate the likelihood ratio and error accumulation sum of each small image unit in (3) and the corresponding small image unit in (2). When the likelihood ratio and the error accumulation sum are both less than the predefined threshold, that is, the preset error condition is met, and it is judged to be normal, that is, the face is alive; when both are greater than the predefined threshold, that is, the preset error condition is not met, it is judged to be abnormal, that is, there is a non-live attack; when either the likelihood ratio or the error accumulation sum is greater than the threshold, the detection is considered invalid.

[0080] Compared with the existing technology, this implementation can solve the problem that lower-end face recognition devices are deceived by photos or videos, and can be directly developed on the original system, that is: using local background images for false face detection (photos or videos, etc.), without the need for additional auxiliary equipment, can reduce detection costs, and the detection system is easy to embed into the existing face recognition system; only detecting the features of the background around the face is more targeted and reduces computing power loss, on the one hand improving the detection speed, on the other hand, it is easy to develop or deploy on low-end face detection devices; using the SAE network for background feature extraction, the model is simple, the feature dimension can change according to the complexity of the actual scene background, the effective feature extraction is accurate, and the detection accuracy is high.

[0081] This implementation method combines a face recognition system, a motion detection system, and a non-living attack detection system, effectively utilizing face recognition data and motion detection data to assist in the non-living attack detection proposed in this solution, and reducing the false alarm rate of detection; updating the reference background map in real time according to the motion detection or the presence or absence of a face target in the area to be detected, effectively reducing the false alarm rate caused by background light changes or moving targets, etc.; dividing the video window into M*N areas, and performing detection management on each area as a unit, to achieve detection of a specific background area to be detected, effectively saving computing power, and facilitating development or deployment on low-end face recognition devices; using the same SAE network to reconstruct the updated background and area to be detected data. After the model training is completed, only the reconstruction difference between the area to be detected and the corresponding area of ​​the background by the same reconstruction model is concerned, and the dual difference of the reconstruction error and statistical characteristics of the local specific area is used to determine whether it is a non-living attack.

[0082] Figure 5 is a schematic diagram of a face recognition device according to an embodiment of the present invention. Figure 5 As shown, according to another aspect of an embodiment of the present invention, a face recognition device is also provided, including: an acquisition module 50, a first determination module 52, a reconstruction module 54, a second determination module 56 and a recognition module 58. The device is described in detail below.

[0083] An acquisition module 50 is used to acquire a target image with a face to be recognized collected by a face recognition device; a first determination module 52 is connected to the acquisition module 50 and is used to determine a background area image in the target image; a reconstruction module 54 is connected to the first determination module 52 and is used to input the background area image into a reconstruction model to obtain a target code corresponding to the background area image, wherein the reconstruction model includes a multi-layer coding network trained by multiple groups of training data, each group of training data includes a background image and a code corresponding to the background image; a second determination module 56 is connected to the reconstruction module 54 and is used to determine whether the face area in the target image is a living area or a non-living area based on the target code and the background code of the face recognition device, wherein the background code is a background code corresponding to the background image of the area where the face recognition device is collected; an identification module 58 is connected to the second determination module 56 and is used to perform face recognition based on the face area when the face area in the target image is a living area.

[0084] Through the above-mentioned device, an acquisition module 50 is used to acquire a target image with a face to be identified that is collected by a face recognition device; a first determination module 52 determines a background area image in the target image; a reconstruction module 54 inputs the background area image into a reconstruction model to obtain a target code corresponding to the background area image, wherein the reconstruction model includes a multi-layer coding network trained by multiple groups of training data, and each group of training data includes a background image and a code corresponding to the background image; a second determination module 56 determines whether the face area in the target image is a living area or a non-living area based on the target code and the background code of the face recognition device, wherein the background code is a background code corresponding to the background image of the area where the face recognition device is collected; the recognition module 58 performs face recognition based on the face area when the face area in the target image is a living area. When comparing the background code of the background image with the target code of the background area of ​​the target image of the face to be identified, if the face to be identified in the target image is a living face, the background code and the target code should be consistent. If the face to be identified in the target image is a photo or video, the background around the face is inconsistent with the real background. The target code and background code of the target image are used to determine whether the face in the target image is alive, thereby achieving the technical effect of reducing the amount of data processing, simplifying the recognition process, and reducing the difficulty of comparison. This solves the technical problem of static detection of face recognition using photos or videos in related technologies, which requires accurate identification of subtle features of the face or background, and is complex and difficult to achieve.

[0085] According to another aspect of an embodiment of the present invention, a processor is further provided, and the processor is used to run a program, wherein the program executes any one of the above-mentioned face recognition methods when running.

[0086] According to another aspect of an embodiment of the present invention, a computer storage medium is further provided, the computer storage medium including a stored program, wherein when the program is running, the device where the computer storage medium is located is controlled to execute any one of the above-mentioned face recognition methods.

[0087] The serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages or disadvantages of the embodiments.

[0088] In the above embodiments of the present invention, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0089] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only schematic. For example, the division of the units can be a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.

[0090] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple units. Some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.

[0091] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.

[0092] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for a computer device (which can be a personal computer, a server or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), mobile hard disk, magnetic disk or optical disk and other media that can store program codes.

[0093] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principle of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.

Claims

1. A face recognition method, characterized in that: include: Obtaining a target image with a face to be identified collected by a face recognition device; Determine a background area image in the target image; Inputting the background area image into a reconstruction model to obtain a target code corresponding to the background area image, wherein the reconstruction model includes a multi-layer coding network trained by multiple sets of training data, each set of training data includes a background image and a code corresponding to the background image; Determining whether the face area in the target image is a living area or a non-living area by comparing the target code and the background code of the face recognition device, wherein the background code is a background code corresponding to the background image of the area where the face recognition device is collected; When the face region in the target image is a living body region, performing face recognition according to the face region; Among them, determining the background area picture in the target picture also includes: identifying the face area in the target picture, deleting the face area in the target picture, and obtaining a non-face area; determining a third motion detection area in the non-face area, wherein the third motion detection area is the motion detection area in the target picture; deleting the third motion detection area to obtain a selected background area; determining the image in the selected background area within a predetermined range from the face area as the background area picture.

2. The method according to claim 1, characterized in that Before determining whether the face area in the target image is a living area or a non-living area according to the target code and the background code of the face recognition device, the method further includes: using a photo of the background of the area collected when the face recognition device is in an idle state as the background picture; The background image is input into the reconstruction model to obtain a background code corresponding to the background image as the background code of the face recognition device.

3. The method according to claim 2, characterized in that Before inputting the background image into the reconstruction model to obtain the background code corresponding to the background image, the method further includes: Acquire a plurality of background training images without a first motion detection area, wherein the first motion detection area is an area in the background training image where motion is detected; Training the reconstruction model according to the background training picture; Input the background training picture into the reconstruction model, and the reconstruction model outputs the corresponding training code; Decoding is performed according to the training code to obtain a reconstructed picture, and an error between the background training picture and the reconstructed picture is determined; When the error satisfies an error threshold, or the number of training times reaches a preset number, it is determined that the training of the reconstruction model is completed.

4. The method according to claim 2, characterized in that: Before inputting the background image into the reconstruction model to obtain the background code corresponding to the background image, the method further includes: Deleting a second motion detection area in the background image, wherein the second motion detection area is an area in the background image where motion is detected; Preprocessing the deleted background image, wherein the preprocessing includes denoising and normalization; Inputting the background image into the reconstruction model to obtain the background code corresponding to the background image includes: The processed background image is input into the reconstruction model to obtain the code corresponding to the background image as the background code.

5. The method according to claim 1, characterized in that The third motion detection area is deleted, and the background area to be selected includes: When the position of the third motion detection area coincides with the position of the second motion detection area of ​​the background image, the third motion detection area is deleted to obtain a background area to be selected in the target image; When the position of the third motion detection area coincides with the position of the non-second motion detection area in the background image, the image of the third motion detection area is replaced with the corresponding background image in the background image to obtain the background area to be selected in the target image.

6. The method according to claim 1, characterized in that Inputting the background area picture into the reconstruction model to obtain the target code corresponding to the background area picture includes: Dividing the background area picture into a plurality of unit areas; Taking each of the plurality of unit regions as input to the reconstruction model, and outputting a reconstruction code of each unit region by the reconstruction model; The target code is generated according to the reconstructed code of each unit area.

7. The method according to any one of claims 1 to 6, characterized in that Determining whether a face region in the target image is a living region or a non-living region according to the target code and the background code of the face recognition device includes: Determine background reconstruction codes corresponding to a plurality of unit areas of the background area picture in the background picture, wherein the background picture is divided into a plurality of unit areas in the same division manner when generating corresponding background codes according to the reconstruction model, and determine background reconstruction codes corresponding to the plurality of unit areas; Calculating the reconstructed codes of the plurality of unit regions in the background region picture and the error parameters of the corresponding background reconstructed codes, wherein the error parameters include a likelihood ratio and a cumulative error; When the error parameters all meet the preset error conditions, determining that the face area in the target image is a living area; When none of the error parameters satisfy the preset error condition, it is determined that the face area in the target image is a non-living area.

8. A face recognition device, characterized in that: include: An acquisition module, used to acquire a target image with a face to be recognized collected by a face recognition device; A first determining module, used to determine a background area image in the target image; A reconstruction module, used for inputting the background area picture into a reconstruction model to obtain a target code corresponding to the background area picture, wherein the reconstruction model includes a multi-layer coding network trained by multiple sets of training data, each set of training data includes a background picture and a code corresponding to the background picture; A second determination module is used to determine whether the face area in the target image is a living area or a non-living area according to the target code and the background code of the face recognition device, wherein the background code is the background code corresponding to the background image of the area where the face recognition device is collected; A recognition module, configured to perform face recognition based on the face area in the target image when the face area is a living area; Among them, the device is also used to identify the face area in the target image, delete the face area in the target image, and obtain a non-face area; determine a third motion detection area in the non-face area, wherein the third motion detection area is the motion detection area in the target image; delete the third motion detection area to obtain a selected background area; determine that the image in the selected background area is within a predetermined range from the face area as the background area image.

9. A processor, characterized in that: The processor is used to run a program, wherein the program, when running, executes the face recognition method described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Power transmission line foreign matter detection method and device based on deep learning, and medium

    CN110751630A

  • Unsupervised fan blade defect detection method and system based on auto-encoder

    CN113256602A