Living body recognition method and device, electronic equipment and storage medium
By obtaining reference image frames with non-repetitive image frames during the live body recognition process for repetitive verification, the shortcomings of the existing live body recognition methods in defense against injection attacks are solved, and more efficient and accurate live body recognition is achieved.
Patent Information
- Application Number
- CN202410133017.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-29
- Publication Date
- 2025-07-29
AI Technical Summary
The existing live body recognition methods have major flaws in defending injection attacks. The mobile live body takes a long time and is easily attacked. The colorful live body fails under strong light conditions and has poor user experience, so it is impossible to fully defend against injection attacks.
By obtaining the non-repetitive image frames sent by the receiving terminal, the reference image frames are obtained and repetitive verification is performed, and the verification results of the image frames to be identified are judged, so as to prevent the injection attacker from using historical materials for live recognition.
It improves the accuracy and security of live body recognition, prevents injection attacks, and improves the efficiency and robustness of live body recognition.
Smart Images

Figure CN120388425A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of face recognition technology, and in particular to a living body recognition method, device, electronic device and storage medium. Background Art
[0002] Liveness detection is a technology that identifies whether a user is a living person, not a machine. It is typically integrated into the login portal of a business application (APP). Injection attacks are a common attack method against liveness detection. Motion liveness and color liveness are commonly used methods to defend against injection attacks. However, motion liveness detection is time-consuming and, due to the limited number of motion combinations, can be easily exploited by injection attacks. Color liveness detection can only protect against injection attacks if it strictly requires "color" feedback.
[0003] It can be seen that the current liveness recognition method has major security flaws and cannot completely defend against injection attacks. Summary of the Invention
[0004] The embodiments of the present application provide a living body recognition method, device, electronic device, and storage medium, which can improve the accuracy of living body recognition.
[0005] The technical solution of the embodiment of the present application is implemented as follows:
[0006] An embodiment of the present application provides a liveness recognition method, which includes: receiving a liveness recognition request sent by a terminal; the liveness recognition request includes an image frame to be recognized, and the image frame to be recognized is a non-repetitive image frame; obtaining a reference image frame of the image to be recognized; performing a repeatability check on the image frame to be recognized based on the reference image frame to obtain a verification result of the image frame to be recognized; and determining a liveness recognition result of a target object in the image frame to be recognized based on the verification result.
[0007] An embodiment of the present application provides a liveness recognition device, comprising: a request receiving module, configured to receive a liveness recognition request sent by a terminal; the liveness recognition request includes an image frame to be recognized, wherein the image frame to be recognized is a non-repetitive image frame; a reference image acquisition module, configured to obtain a reference image frame of the image frame to be recognized; a repeatability verification module, configured to perform a repeatability verification on the image frame to be recognized based on the reference image frame to obtain a verification result of the image frame to be recognized; and a recognition result determination module, configured to determine a liveness recognition result of a target object in the image frame to be recognized based on the verification result.
[0008] An embodiment of the present application provides an electronic device, comprising: a memory for storing computer-executable instructions; and a processor for implementing the liveness recognition method provided in the embodiment of the present application when executing the computer-executable instructions stored in the memory.
[0009] An embodiment of the present application provides a computer-readable storage medium storing a computer program or computer-executable instructions, which are used to implement the living body recognition method provided by the embodiment of the present application when being executed by a processor.
[0010] An embodiment of the present application provides a computer program product, including a computer program or computer-executable instructions, and when the computer program or computer-executable instructions are executed by a processor, the living body recognition method provided by the embodiment of the present application is implemented.
[0011] The embodiment of the present application has the following beneficial effects:
[0012] During the process of living body recognition, a reference image frame corresponding to the image frame to be recognized is obtained, the image frame to be recognized is repetitively verified based on the reference image frame, a verification result of the image frame to be recognized is obtained, and a living body recognition result of the target object included in the image frame to be recognized is determined based on the verification result. If the verification result is that an image frame repeated with the image frame to be recognized appears in the reference image frame, then the image frame to be recognized may not be obtained by real-time collection of the target object, but is historical video, photo and other materials replaced by an injecting attacker. At this time, a living body recognition failure result can be obtained based on this verification result. Therefore, the embodiment of the present application can prevent an injecting attacker from using materials such as photos to replace the target object for living body recognition by performing repetitive verification on the image frame to be recognized, and improve the accuracy of living body recognition. Description of the Drawings
[0013] Figure 1 is a schematic diagram of an action sequence of action living body in the related art;
[0014] Figure 2 is a schematic diagram of an interaction interface of action living body in the related art;
[0015] Figure 3 is a schematic diagram of an interaction interface of colorful living body in the related art;
[0016] Figure 4 is a schematic structural diagram of a living body recognition system architecture provided by the embodiment of the present application;
[0017] Figure 5A is a schematic structural diagram of a living body recognition device provided by the embodiment of the present application;
[0018] Figure 5B is a schematic structural diagram of a living body recognition device provided by the embodiment of the present application;
[0019] Figure 6A is a schematic flow diagram of a living body recognition method provided by the embodiment of the present application;
[0020] Figure 6BIt is a schematic flowchart of the living body recognition method provided by the embodiments of the present application;
[0021] Figure 6C It is a schematic flowchart of the living body recognition method provided by the embodiments of the present application;
[0022] Figure 7 It is a schematic flowchart of the repetitive verification method provided by the embodiments of the present application;
[0023] Figure 8 It is a schematic flowchart of the key image area extraction method provided by the embodiments of the present application;
[0024] Figure 9 It is an example diagram of the key image area extraction provided by the embodiments of the present application;
[0025] Figure 10 It is a schematic flowchart of the key image area matrix search provided by the embodiments of the present application;
[0026] Figure 11 It is a schematic flowchart of the coding repetitive verification provided by the embodiments of the present application;
[0027] [[ID=2⑦]] Figure 12 It is a schematic flowchart of the living body recognition method provided by the embodiments of the present application;
[0028] Figure 13 It is a schematic flowchart of the living body recognition method provided by the embodiments of the present application;
[0029] Figure 14 It is a schematic diagram of the interaction interface of the living body recognition method provided by the embodiments of the present application. Specific embodiments
[0030] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be construed as limiting the present application. All other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present application.
[0031] In the following description, reference is made to "some embodiments" which describe a subset of all possible embodiments. However, it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments and can be combined with each other without conflict.
[0032] In the following description, the terms "first", "second", and "third" only distinguish similar objects and do not represent a specific order for the objects. It can be understood that "first", "second", and "third" can be interchanged in a specific order or sequence when allowed, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.
[0033] Unless otherwise defined, all technical and scientific terms used in the embodiments of the present application have the same meaning as commonly understood by those skilled in the art to which the present application belongs. The terms used in the embodiments of the present application are only for the purpose of describing the embodiments of the present application and are not intended to limit the present application.
[0034] Before further elaborating on the embodiments of the present application, the nouns and terms involved in the embodiments of the present application are described. The nouns and terms involved in the embodiments of the present application are subject to the following explanations.
[0035] 1) Live recognition: A technology for identifying whether the user is a living body rather than a machine. Generally, the system structure consists of a collection end and a control end. The methods of live recognition generally include silent live, digital live, colorful live, motion live, etc.
[0036] 2) Collection end: Also known as the front end or client, it is a program for directly interacting with users and collecting data. It is usually deployed on the terminal for task execution.
[0037] 3) Control end: Also known as the decision-making end, back end, or server, it is usually deployed on the cloud server for making decisions and judgments.
[0038] 4) Injection attack: It refers to an attack method in which a hacker replaces the face image to be uploaded when the terminal collects the face image through hacking technology. For example, an injection attacker can control the camera data link of the face recognition system through hacking software, directly replace the image collected by the camera with an attack image, and then input the replaced attack image into the face recognition system.
[0039] Live recognition is a live detection device, technology, and method deployed on a machine for implementing the Turing test of the machine against the operator. Live recognition is a common functional module for a machine to determine whether the opposite end is a human. It usually does not exist alone but is integrated into the login end of the business APP and is also often used as a means of surprise inspection. For example, in daily life, the "teenager anti-addiction mode" in games uses live detection (face or verification code) methods to limit teenagers' excessive addiction to games. In the management of operating vehicles, "driver on-duty identity confirmation" uses live detection (face) methods to confirm the situation of manual on-duty. In various knowledge websites and APPs, "anti-crawler" uses live detection (verification code) to intercept crawlers.
[0040] Motion biopsy is the most widely used type of biopsy device, technology and method. It challenges the user to perform motions (such as Figure 1 As shown, the basic action sequence consists of a combination of opening the mouth, shaking the head left and right, up and down, and blinking. The algorithm engine then identifies and verifies whether the user has completed the specified action challenge. This algorithm engine utilizes technologies such as facial landmark location, face tracking, and action recognition to verify the user's authenticity. The core of action liveness is to implement challenges based on the randomness of movements and assume that only living humans can complete the corresponding action commands. Figure 2 This is an example diagram of the motion liveness interaction interface. The motion liveness principle steps are as follows: the mobile application sends a random action request to the controller; the controller issues a random action command and starts a countdown on the action command's lifecycle; the mobile terminal prompts the user with an action prompt based on the random action command and collects user action data; the user performs the relevant action and responds; the mobile terminal transmits the media back to the controller; the controller feeds the media content into an algorithm for parsing and calculates whether it is consistent with the random action command. Motion recognition has the following shortcomings: the user experience is poor, typically taking more than 5 seconds to complete two rounds of actions; and due to the limited number of action combinations, it is easy for attackers to use injection attacks to exhaustively enumerate various attack image combinations and replace the user image collected by the mobile terminal, making liveness recognition successful. Therefore, injection attacks are not defended against.
[0041] Colorful Liveness, also known as Glare Liveness or Light Liveness, proposes a liveness detection algorithm based on light sequence recognition. It challenges the user with light and uses the algorithm to identify whether a corresponding light sequence appears on the user's face. The algorithm engine utilizes technologies such as facial landmark location, face tracking, and color recognition to verify the user's liveness. The core of Colorful Liveness is based on the color and frequency of light changes to achieve the challenge, assuming that only faces can specifically reflect these light signal information. Figure 3 This is an example diagram of the colorful liveness interaction interface. The colorful liveness principle steps are as follows: the mobile application sends a random colorful password request to the controller; the controller issues a random colorful password and starts the colorful password lifecycle countdown; the mobile terminal sends a colorful prompt to the user based on the random colorful password and collects user data; the user responds silently and waits for the colorful detection and collection cycle to end; the mobile terminal sends the media back to the controller; the controller sends the media content to the algorithm for parsing and calculates whether it is consistent with the random colorful password. Motion recognition has the following shortcomings: poor system robustness. Under strong daylight conditions, the "colorful" screen cannot be effectively reflected from the target, causing the module to fail; poor user experience. High-intensity "colorful" can cause discomfort to the human eye; and "injection attacks" can only be defended under conditions where "colorful" feedback is strictly required.
[0042] It can be seen that the existing in-living body recognition methods have significant drawbacks in terms of user experience and security. For example, the overall user cooperation in the action-based in-living body recognition takes a long time, and it cannot defend against injection attacks. The colorful in-living body recognition has the problem of dazzling, and it also cannot completely resist injection attacks.
[0043] The embodiment of the present application provides an in-living body recognition method. A library is built for the in-living body image frames obtained during the in-living body recognition process corresponding to each request identity (Identity document, ID), and multiple reference image frames corresponding to each request identity are obtained. When a new in-living body request is initiated, the request identity is obtained, and it is checked whether the key image area of each frame of the image to be recognized already exists in the reference image frames corresponding to the request identity. If a frame repetition phenomenon occurs in the key image area (such as the face area), it is determined that the user has launched an injection attack. If no repetition occurs, the in-living body recognition result is further combined to determine whether the in-living body recognition is successful, preventing an injection attacker from using materials such as photos to replace the target object for in-living body recognition and improving the accuracy of in-living body recognition.
[0044] Before elaborating on the in-living body recognition method of the embodiment of the present application, first, an exemplary application of the in-living body recognition device provided by the embodiment of the present application for implementing this in-living body recognition method is described. This in-living body recognition device is an electronic device for implementing the in-living body recognition method. In one implementation manner, the in-living body recognition device (i.e., the electronic device) provided by the embodiment of the present application can be implemented as a terminal or as a server. In one implementation manner, the in-living body recognition device provided by the embodiment of the present application can be implemented as various types of user terminals such as a laptop computer, a tablet computer, a desktop computer, a set-top box, a mobile device (e.g., a mobile phone, a portable music player, a personal digital assistant, a dedicated messaging device, a portable game device), a smart phone, a smart speaker, a smart watch, a smart TV, a vehicle-mounted terminal, etc.; in another implementation manner, the in-living body recognition device provided by the embodiment of the present application can also be implemented as a server. Among them, the server can be an independent physical server, or a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery network (CDN, Content Delivery Network), and big data and artificial intelligence platforms. The terminal and the server can be directly or indirectly connected through wired or wireless communication methods, which are not limited in the embodiment of the present application. Next, the exemplary application when the in-living body recognition device is implemented as a server will be described.
[0045] See Figure 4 , Figure 4FIG. 0 is an optional architecture diagram of the live body recognition system 100 provided by the embodiments of the present application. To support any live body recognition application, the live body recognition application is used to identify and verify the target object, and the live body recognition result of the target object is obtained to improve the accuracy of live body recognition. At least the live body recognition application (client 400-1) is installed on the terminal of the embodiments of the present application. The live body recognition system 100 at least includes a server 200, a network 300, a terminal 400, and a database 500, where the server 200 is the server of the live body recognition application. The server 200 may constitute the live body recognition device of the embodiments of the present application, that is, the live body recognition method of the embodiments of the present application is implemented through the server 200. The terminal 400 is connected to the server 200 through the network 300. The network 300 may be a wide area network, a local area network, or a combination of the two.
[0046] When performing live body recognition, the user can input a live body recognition operation through the live body recognition application (client) 400-1 running on the terminal 400. The terminal 400 generates a live body recognition request in response to the live body recognition operation, and sends the live body recognition request to the server 200 through the network 300. Among them, the live body recognition request includes a to-be-recognized image frame, and the to-be-recognized image frame is a non-repeating image frame. After receiving the live body recognition request, the server 200 obtains a reference image frame of the to-be-recognized image frame in response to the live body recognition request; performs a repeatability check on the to-be-recognized image frame based on the reference image frame to obtain a check result of the to-be-recognized image frame; determines the live body recognition result of the target object in the to-be-recognized image frame based on the check result. At the same time, after obtaining the live body recognition result, the server 200 can also send the live body recognition result to the terminal 400, and the terminal 400 displays the live body recognition result to the user through the client 400-1.
[0047] See Figure 5A , Figure 5A FIG. 10 is a structural diagram of an electronic device provided by the embodiments of the present application. Figure 5A The electronic device shown in FIG. 10 may be a live body recognition device. The live body recognition device includes: at least one processor 410, a memory 450, at least one network interface 420, and a user interface 430. Each component in the live body recognition device is coupled together through a bus system 440. It can be understood that the bus system 440 is used to realize the connection and communication between these components. In addition to the data bus, the bus system 440 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clear illustration, in Figure 5A FIG. 10, all kinds of buses are labeled as the bus system 440.
[0048] The processor 410 may be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor may be a microprocessor or any conventional processor, etc.
[0049] The user interface 430 includes one or more output devices 431 that enable the presentation of media content, including one or more speakers and / or one or more visual display screens. The user interface 430 also includes one or more input devices 432, including user interface components that facilitate user input, such as a keyboard, a mouse, a microphone, a touch screen display, a camera, and other input buttons and controls.
[0050] The memory 450 may be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state memory, hard disk drives, optical disc drives, etc. The memory 450 optionally includes one or more storage devices that are physically remote from the processor 410.
[0051] The memory 450 includes volatile memory or non-volatile memory, and may also include both volatile and non-volatile memory. The non-volatile memory may be a read-only memory (ROM), and the volatile memory may be a random access memory (RAM). The memory 450 described in the embodiments of the present application is intended to include any suitable type of memory.
[0052] In some embodiments, the memory 450 is capable of storing data to support various operations. Examples of such data include programs, modules, and data structures, or subsets or supersets thereof, which are exemplarily described below.
[0053] The operating system 451 includes system programs for processing various basic system services and performing hardware-related tasks, such as a framework layer, a core library layer, a driver layer, etc., for implementing various basic services and processing hardware-based tasks;
[0054] The network communication module 452 is used to reach other electronic devices via one or more (wired or wireless) network interfaces 420. Exemplary network interfaces 420 include: Bluetooth, Wi-Fi (Wireless Fidelity), and USB (Universal Serial Bus), etc.;
[0055] A presentation module 453 for enabling presentation of information (e.g., a user interface for operating a peripheral device and displaying content and information) via one or more output devices 431 associated with the user interface 430 (e.g., a display screen, a speaker, etc.);
[0056] An input processing module 454 for detecting one or more user inputs or interactions from one of one or more input devices 432 and translating the detected inputs or interactions.
[0057] In some embodiments, as Figure 5B shown, the device provided by the embodiments of the present application may be implemented in software. Figure 5B Shown is a living body recognition device 455 stored in the memory 450, which may be software in the form of a program, a plug-in, etc., including the following software modules: a request receiving module 4551, a reference image acquisition module 4552, a repeatability verification module 4553, and a recognition result determination module 4554. These modules are logical, and thus can be combined arbitrarily or further split according to the functions implemented. The functions of each module will be described below.
[0058] In other embodiments, the device provided by the embodiments of the present application may be implemented in hardware. As an example, the device provided by the embodiments of the present application may be a processor in the form of a hardware decoding processor, which is programmed to execute the living body recognition method provided by the embodiments of the present application. For example, a processor in the form of a hardware decoding processor may employ one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field programmable gate arrays (FPGAs), or other electronic components.
[0059] The living body recognition method provided by each embodiment of the present application may be executed by an electronic device, where the electronic device may be a server or a terminal. That is, the living body recognition method provided by each embodiment of the present application may be executed by the server, may be executed by the terminal, or may also be executed through interaction between the server and the terminal.
[0060] Figure 6A is a schematic flowchart of the living body recognition method provided by the embodiments of the present application. Below will be combined with Figure 6AThe steps shown are explained as Figure 6A As shown, the liveness recognition method is described as an example in which the execution subject is a server. The method includes the following steps S101 to S104:
[0061] Step S101: receiving a liveness recognition request sent by a terminal.
[0062] The living body recognition request includes image frames to be recognized, and the image frames to be recognized are non-repeated image frames.
[0063] Here, the liveness recognition request includes the image frame to be recognized, that is, the liveness recognition request carries the image frame to be recognized. The image frame to be recognized is an image frame containing the target object. A non-repeating image frame to be recognized means that during this liveness recognition process, the terminal does not have any image frames that are duplicates of the image frame to be recognized in the video of the target object captured by the terminal. After capturing the image frame to be recognized, the terminal performs a duplicate frame check on the image frame to be recognized. If the test result indicates that no image frames duplicate the image frame to be recognized, the image frame to be recognized is a non-repeating image frame, and the terminal sends a liveness recognition request carrying the image frame to be recognized to the server. Target objects include those requiring liveness recognition verification in various scenarios, such as users requiring facial recognition in financial scenarios. However, due to injection attacks, the subject initiating liveness recognition on the terminal may not be the user themselves.
[0064] For example, in the absence of an injection attack, the target user is the user who has enabled liveness recognition on the terminal, and the image frames to be identified are video frames obtained by capturing real-time video of the target user on the terminal, i.e., photos of the target user captured in real time by the terminal using an image capture device. Alternatively, in the presence of an injection attack, the injection attacker enables liveness recognition on the terminal, and the image frames to be identified are video frames obtained by capturing real-time images or videos provided by the injection attacker.
[0065] In an embodiment of the present application, the liveness recognition request sent by the receiving terminal can be implemented in the following manner: when the terminal performs duplicate frame detection on the image frame to be recognized, the detection result is the first detection result, the liveness recognition request sent by the terminal is received. The first detection result is that the terminal does not have an image frame that is a duplicate of the image frame to be recognized. The terminal performs duplicate frame detection on the recognition image frame in the following manner: first, obtain the image frame to be recognized at the current moment. Then, based on the historical image frames to be recognized stored by the terminal, the image frame to be recognized is subjected to duplicate frame detection to obtain a detection result. The historical image frames to be recognized are all the image frames to be recognized acquired before this liveness recognition.
[0066] In an embodiment of the present application, during a liveness recognition process, the number of image frames to be recognized that are collected by the terminal is multiple. For any image frame to be recognized, the terminal can perform duplicate frame detection on the image frame to be recognized based on the historical image frames to be recognized in the local cache. If the detection result of the terminal performing duplicate frame detection on the image frame to be recognized is the first detection result, that is, the terminal does not have an image frame that is repeated with the image frame to be recognized, the server can receive the liveness recognition request for the image frame to be recognized sent by the terminal. At the same time, the terminal stores the image frame to be recognized in the local cache of the terminal as the historical image frame to be recognized used in the duplicate frame detection process of the next image frame to be recognized. Alternatively, if the detection result of the terminal performing duplicate frame detection on the image frame to be recognized is that the terminal has an image frame that is repeated with the image frame to be recognized, the liveness recognition process can be terminated directly, and a warning message of liveness recognition failure is displayed to the object that starts the liveness recognition operation at the terminal.
[0067] Exemplarily, the terminal acquires the i-th image frame to be identified. At this time, the historical image frames to be identified in the local cache of the terminal include all the image frames to be identified from the 1st frame to the i-1th frame. The terminal can perform duplicate frame detection on the i-th image frame to be identified based on all the image frames to be identified from the 1st frame to the i-1th frame, and obtain the detection result of the i-th image frame to be identified. If there is no image frame that is repeated with the i-th image frame to be identified in all the image frames to be identified from the 1st frame to the i-1th frame, a first detection result is obtained, and a liveness recognition request including the i-th image frame to be identified is sent to the server. At the same time, the i-th image frame to be identified is stored in the local cache of the terminal. Wherein, i is an integer greater than 1. It should be noted that for the 1st image frame to be identified, there is no historical image frame to be identified in the local cache of the terminal. At this time, the detection result of the 1st image frame to be identified is defaulted to be the first detection result.
[0068] Here, the liveness recognition request may include a request identifier and the image frame to be recognized. The request identifier may include the target subject's identity information, which may be in various forms including, but not limited to, numbers and letters. For example, the target subject's identity information may include an ID number, an account number, or the like.
[0069] In some embodiments, the request identifier may further include terminal identification information. The terminal identification information may be a unique identifier for a terminal device, and its expression may include but is not limited to various forms such as numbers and letters.
[0070] In the embodiments of the present application, by using the locally cached historical frames to be recognized of the terminal to perform duplicate frame detection on the frames to be recognized, it is possible to prevent an injection attacker from using a static picture of a target object to replace the image frame actually captured by the terminal, thereby improving the accuracy and security of liveness recognition. Additionally, only when the detection result indicates that there is no image frame in the terminal that duplicates the frame to be recognized, after receiving the liveness recognition request sent by the terminal, will the terminal start to perform a repeatability verification on the frame to be recognized carried in the liveness recognition request. In this way, the terminal has already performed duplicate frame detection on the frame to be recognized in advance. If the detection result of the terminal is that there is a duplicate frame of the frame to be recognized, the subsequent repeatability verification process can be skipped, and the result of liveness recognition failure can be directly obtained, thereby improving the efficiency of liveness recognition. Moreover, since only one duplicate frame detection is performed on the frame to be recognized by the terminal, it may lead to missed detections. However, in the embodiments of the present application, not only is duplicate frame detection performed on the frame to be recognized by the terminal, but also repeatability verification is performed on the frame to be recognized by the server, so the accuracy of liveness recognition can be further improved.
[0071] Step S102: Obtain a reference image frame of the frame to be recognized.
[0072] In the embodiments of the present application, the frame to be recognized includes a target object, and the reference image frame is the frame to be recognized stored during each liveness recognition process within a preset historical time period. The reference image frame can be stored in a database. The reference image frame corresponding to the frame to be recognized can be obtained based on the request identifier carried in the liveness recognition request. If the liveness recognition request carries terminal identifier information, the reference image frame corresponding to the frame to be recognized is each frame to be recognized obtained by the terminal corresponding to the terminal identifier information for any object within a preset historical time period. If the liveness recognition request carries the identity identifier information of the target object, the reference image frame corresponding to the frame to be recognized is each frame to be recognized obtained by the target object on any device within a preset historical time period. It should be noted that the embodiments of the present application do not limit the preset historical time period. After obtaining the request identifier, all reference data frames corresponding to the request identifier can be retrieved from the database according to the request identifier.
[0073] Step S103: Perform repeatability verification on the frame to be recognized based on the reference image frame to obtain the verification result of the frame to be recognized.
[0074] In the embodiments of the present application, each liveness recognition request sent by the terminal carries only one frame to be recognized. Figure 7 is a schematic diagram of the repeatability verification process. Refer to Figure 7, after receiving each image frame to be recognized sent by the terminal, the repeatability verification can be performed on the image frame to be recognized based on all the reference image frames, and the verification result of the image frame to be recognized can be obtained. If the verification result of the image frame to be recognized indicates that there is no duplicate frame of the image frame to be recognized among all the reference image frames, the image frame to be recognized is stored in the local cache as a new reference image frame for the repeatability verification of the next image frame to be recognized. Alternatively, if the verification result of the image frame to be recognized indicates that there is a duplicate frame of the image frame to be recognized among all the reference image frames, a request to end the current liveness recognition process is directly sent to the terminal.
[0075] In some embodiments, refer to Figure 7 , for each image frame to be recognized, after the verification result of the image frame to be recognized indicates that there is no duplicate frame of the image frame to be recognized among all the reference image frames, it can also be detected whether the image frame to be recognized is the last image frame to be recognized sent by the terminal. If the image frame to be recognized is the last frame, all the reference image frames in the local cache are synchronized to the database, which can increase the amount of data in the database when detecting duplicate frames subsequently, so as to improve the accuracy of liveness recognition.
[0076] In the embodiments of the present application, each image frame to be recognized that passes the repeatability verification is stored in the database as a reference image frame for the next repeatability verification, that is, all the historical image frames to be recognized are stored, rather than only storing key frames. This makes the reference image frames have a huge advantage in terms of data volume during the repeatability verification process, and can completely cover the range of attack materials used by injected attackers. Therefore, the embodiments of the present application make full use of the characteristic that injected attackers recycle attack materials, ensuring that the recycled attack materials cannot pass the repeatability verification, thereby improving the accuracy of liveness recognition.
[0077] In some embodiments, refer to Figure 6B , Figure 6A The steps shown in S103 can be implemented by the following steps S1031A to S1034A, which will be specifically described below.
[0078] Step S1031A, perform semantic segmentation on the image frame to be recognized to obtain multiple image regions.
[0079] Figure 8 It is a schematic flow diagram of the process for extracting key image regions of the image. Refer to Figure 8, in the embodiments of the present application, for each image frame to be recognized, the image frame to be recognized can be semantically segmented, and based on the category of each pixel in the image frame to be recognized, the image frame to be recognized is segmented into M image regions, where M is a positive integer. It should be noted that the embodiments of the present application do not limit the method of semantic segmentation. For example, a segmentation method based on thresholds, a region-based image segmentation method, an edge detection-based segmentation method, an image segmentation method based on wavelet analysis and wavelet transform, an image segmentation based on genetic algorithms, an active contour model-based segmentation method, etc. can be used, and a deep learning-based segmentation method (such as an SOTA segmentation model) can also be used.
[0080] Exemplarily, after semantically segmenting the image frame to be recognized, multiple image regions such as a face image region, a body image region, a background image region, and an item image region can be obtained.
[0081] Step S1032A, obtain the score of each image region.
[0082] In the embodiments of the present application, refer to Figure 8 , after obtaining multiple image regions, the multiple image regions can be scored to obtain the score of each image region. For each image region, the image region can be converted into a picture segment of a specific size (such as 200*200), after grayscale processing, it is input into a pre-trained classification network, and through the softmax function (normalized exponential function) for normalization processing, the score of the image region within the interval [0,1] is output. It should be noted that the embodiments of the present application do not specifically limit the pre-trained classification network. For example, it can be a residual network (Resnet, Residual Network). The pre-trained classification network can be pre-trained based on historical data that has been labeled, such as pictures of faces, water cups, clothing, etc.
[0083] Step S1033A, determine the key image regions of the image frame to be recognized according to the scores of the multiple image regions.
[0084] In the embodiments of the present application, the size of the key image region is smaller than the size of the image frame to be recognized. The key image region is an image intercepted from the image frame to be recognized for repeatability verification. That is, each pixel that makes up the key image region exists in the image frame to be recognized corresponding to the key image region. The number of key image regions of the image frame to be recognized can be one or multiple. After determining the score of each image region, the image regions can be arranged in descending order of score to obtain an image region sequence. The first N image regions in the image region sequence are determined as N key image regions; N is a positive integer. Refer to Figure 8, after obtaining the sequence of image regions, the first N image regions arranged in the sequence of image regions can be determined as key image regions. N is a positive integer, and N is less than or equal to M. It should be noted that the embodiments of the present application do not limit the value of N, and it can be set according to the actual situation.
[0085] Exemplarily, if N is 1, then the 1 image region with the highest score is selected from the sequence of image regions as the key image region. Figure 9 For the example diagram of key image region extraction, as Figure 9 shown, at this time, the 1 image region with the highest score extracted from the original image frame to be recognized is a face image region, that is, the face image region is determined as the key image region. Exemplarily, if N is 3, then the 3 image regions with the highest scores are selected from the sequence of image regions. If the 3 image regions with the highest scores are a face image region, a body image region, and a background image region, then the above face image region, body image region, and background image region are used as the 3 key image regions of the image frame to be recognized.
[0086] In some embodiments, after determining the score of each image region, the image regions can also be arranged in ascending order of score to obtain a sequence of image regions. The last N image regions in the sequence of image regions are determined as key image regions.
[0087] Step S1034A, determine the verification result of the image frame to be recognized according to the key image region and the reference image frame.
[0088] In an embodiment of the present application, it is possible to determine whether the key image area belongs to a reference image frame to obtain a verification result of the image frame to be identified. For each image frame to be identified, in the process of performing repeatability verification on the image frame to be identified based on all reference image frames, the key image area in the image frame to be identified can be used to replace the image frame to be identified. By determining whether the key image area in the image frame to be identified belongs to a reference image frame, the verification result of the image frame to be identified can be obtained. If the key image area in the image frame to be identified belongs to at least one reference image frame, a verification result is obtained that indicates that the image frame to be identified has a local repeated frame phenomenon, which is also a verification result that indicates that there are repeated frames of the image frame to be identified in all reference image frames. Alternatively, if the key image area in the image frame to be identified does not belong to any reference image frame, a verification result is obtained that indicates that the image frame to be identified does not have a local repeated frame phenomenon, which is also a verification result that indicates that there are no repeated frames of the image frame to be identified in all reference image frames. It should be noted that if there are multiple key image areas in the image frame to be identified, if these multiple key image areas belong to the same reference image frame, a verification result is obtained that indicates that there is a local repeated frame phenomenon in the image frame to be identified, which is also a verification result that indicates that there are repeated frames of the image frame to be identified in all reference image frames.
[0089] The embodiment of the present application improves the robustness of liveness recognition by extracting key image areas in the image frame to be identified and performing repeatability verification based on the reference image frame and the key image area. At the same time, the scope of repeatability verification is framed in the local key image area rather than the entire image frame to be identified, which greatly improves the recall speed and recognition accuracy of attack materials and realizes rapid and efficient detection of injection attacks.
[0090] In some embodiments, the verification result includes a first verification result and a second verification result. Figure 6B Step S1034A shown can be implemented by the following steps: obtaining image matrix data of the key image area and reference matrix data of the reference image frame. The data dimension of the image matrix data is less than the data dimension of the reference matrix data. If the image matrix data is a sub-matrix data of the reference matrix data, then determining the verification result of the image frame to be identified as a first verification result. If the image matrix data is not a sub-matrix data of the reference matrix data, then determining the verification result of the image frame to be identified as a second verification result.
[0091] In the embodiment of this application, Figure 10As shown, for each reference image frame, the reference image frame can be split into three RGB (Red, Green, Blue) channels to form a three-channel reference matrix data, and the representation form of the reference matrix data is a large cube. It should be noted that since common digital images are all in three RGB (Red, Green, Blue) channels, digital images can be directly converted into matrices through tools such as PIG and opencv. The specific implementation method of converting the reference image frame into the reference matrix data in the embodiments of the present application is not limited.
[0092] In the embodiments of the present application, as Figure 10 shown, for the key image region of each image frame to be recognized, the key image region can be split into three RGB (Red, Green, Blue) channels to form a three-channel image matrix data, and the representation form of the image matrix data is a small cube. If the image frame to be recognized includes multiple key image regions, each key image region can be split into three RGB channels respectively to obtain the image matrix data corresponding to each key image region, that is, multiple small cubes are obtained. The number of channels of the image matrix data is equal to the number of channels of the reference matrix data. The image matrix data is a three-channel matrix, and the data dimension of the image matrix data can include the number of rows and columns of the image matrix data; the data dimension of the reference matrix data can include the number of rows and columns of the reference matrix data. The data dimension of the image matrix data being less than the data dimension of the reference matrix data means that: the number of rows of the image matrix data is less than or equal to the number of rows of the reference matrix data, and the number of columns of the image matrix data is less than or equal to the number of columns of the reference matrix data.
[0093] In some embodiments, the key image region can be directly converted into the image matrix data through tools such as PIG and opencv.
[0094] In some embodiments, the step of obtaining the image matrix data corresponding to the key image region can be implemented in the following manner: Obtain the pixel values of the pixels in the key image region in the target color channels. Based on the pixel values, obtain the image matrix data of the key image region.
[0095] In the embodiments of the present application, the target color channels can include the R channel, the G channel, and the B channel. For the key image region of any image frame to be recognized, the pixel values of each pixel in the key image region in the R channel, the G channel, and the B channel can be obtained respectively. Here, the pixel value of a pixel is the RGB value of the corresponding color, and the value range is 0-255. The pixel values of each pixel under each target color channel form a two-dimensional matrix of each target color channel. Combine the two-dimensional matrices corresponding to each target color channel to obtain the image matrix data of the key image region.
[0096] The embodiment of the present application constructs image matrix data by obtaining the pixel values of pixels in the key image area under the target color channel. Compared with images, matrix data is more convenient for repeatable detection, thereby improving the speed of living body recognition.
[0097] In the embodiment of the present application, it is possible to determine whether the image matrix data is sub-matrix data in the reference matrix data, and obtain a verification result of the image frame to be identified.
[0098] For any reference matrix data, the reference matrix data can be divided into multiple sub-matrix data, where the data dimensions of the sub-matrix data are the same as those of the image matrix data. Each sub-matrix data can have overlapping rows and columns. For each image matrix data, all reference matrix data can be traversed according to a preset search order to detect whether the image matrix data is identical to the sub-matrix data in a reference matrix data.
[0099] In some embodiments, it should be noted that the present application embodiment does not limit the preset search order, for example, the search order may be up first, then down, left first, then right, or front, then back. If the reference matrix data and the image matrix data are represented in space by the X, Y, and Z axes, such as Figure 10 As shown in the figure, the above process is to search the image matrix data (small cube) corresponding to the key image area along the X, Y, and Z axes from small to large to see if it is the same as a small cube in a reference data matrix (large cube). The search process takes the search for a (2*2) two-dimensional matrix in a (3*3) two-dimensional matrix as an example. The (3*3) two-dimensional matrix is [1,2,3][4,5,6][7,8,9], and its sub-matrix data is: sub-matrix 1 [1,2][4,5], sub-matrix 2 [2,3][5,6], sub-matrix 3 [4,5][7,8], sub-matrix 4 [5,6][8,9]. The (2*2) two-dimensional matrix is searched to see if it is the same as the above four sub-matrix data.
[0100] In some embodiments, the number of key image regions in the image frame to be identified is one, and for any image matrix data, if the image matrix data is identical to at least one sub-matrix data in at least one reference matrix data, a first verification result is obtained indicating that the image frame to be identified corresponding to the image matrix data has a local repeated frame phenomenon. Alternatively, if the image matrix data is different from each sub-matrix data in each reference matrix data, a second verification result is obtained indicating that the image frame to be identified corresponding to the image matrix data does not have a local repeated frame phenomenon. Alternatively, the number of key image regions in the image frame to be identified is multiple, and if the image matrix data of each key image region in the image frame to be identified is identical to at least one sub-matrix data, and the identical sub-matrix data corresponding to each image matrix data belongs to the same reference matrix data, a first verification result is obtained indicating that the image frame to be identified has a local repeated frame phenomenon.
[0101] The embodiment of the present application converts the reference image frame into reference matrix data and the key image area into image matrix data for repeatability verification, which improves the processing speed and accuracy compared to directly using the image frame for repeatability verification.
[0102] In some embodiments, see Figure 6C , Figure 6A The illustrated step S103 can also be implemented by the following steps S1031B to S1032B, which are described in detail below.
[0103] Step S1031B: perform encoding processing on the image frame to be identified and the reference image frame respectively to obtain the image code to be identified of the image frame to be identified and the reference image code of the reference image frame.
[0104] In an embodiment of the present application, after obtaining a reference image frame, a coding engine can be used to compress the reference image frame into a corresponding reference image code. The reference image code is a unique code used to indicate that the total length of the reference image frame is multiple bits (bytes). After each image frame to be identified is received from a terminal, the coding engine can be used to compress the image frame to be identified into a corresponding image code to be identified. The image code to be identified is a unique code used to indicate that the total length of the image frame to be identified is multiple bits (bits). Therefore, repeatability verification can be achieved by detecting the repeatability of the unique code.
[0105] It should be noted that the embodiments of the present application do not specifically limit the encoding engine. For example, the encoding engine can be a hash algorithm such as Secure Hash Algorithm 1 (SHA1), Secure Hash Algorithm 256 (SHA256), Message Digest Algorithm 5 (MD5), Locality Sensitive Hash, etc., or a face recognition network (Facenet) or a face detection network (Retinaface) based on deep learning pre-training.
[0106] For example, taking the example of encoding the image frame to be identified and the reference image frame respectively using the MD5 encoding method, the obtained image code to be identified and the reference image code are both 128-bit (16-byte) codes in length, generally represented by 32-bit hexadecimal, that is, 1 byte can be represented as 2 consecutive hexadecimal digits, in the form of "902fbdd2b1df0c4f70b4a5d23525e932".
[0107] Step S1032B: Based on the reference image encoding, perform a repeatability check on the to-be-recognized image encoding to obtain the check result of the to-be-recognized image frame.
[0108] In the embodiments of the present application, Figure 11 is a schematic flowchart of the repeated frame detection. Refer to Figure 11 , the to-be-recognized image encoding corresponding to the latest received to-be-recognized image frame can be compared with all the reference image encodings to obtain the check result of this to-be-recognized image frame. It should be noted that the method for comparing the to-be-recognized image encoding with the reference image encoding in the embodiments of the present application is not specifically limited. For example, a cosine distance comparison algorithm can be used, etc.
[0109] In the embodiments of the present application, by converting the reference image frame into a reference image encoding and the to-be-recognized image frame into a to-be-recognized image encoding for repeatability check, compared with directly using the image frame for repeatability check, the processing speed and accuracy are improved.
[0110] In some embodiments, the to-be-recognized image encoding and the reference image encoding can be compared using a strong matching method, that is, it is required that each byte and bit are the same. If the encoding is represented in hexadecimal, every two hexadecimal numbers are a byte, and the byte in the to-be-recognized image encoding corresponds to the byte in the same position in the reference image encoding. For example, the third byte in the to-be-recognized image encoding corresponds to the third byte in the reference image encoding. Taking the example of encoding the to-be-recognized image frame and the reference image frame respectively through the MD5 encoding method, the obtained to-be-recognized image encoding is "902fbdd2b1df0c4f70b4a5d23525e932", and the reference image encoding is "902fbdd2b1df0c4f70b4a5d23525e932". At this time, each byte in the to-be-recognized image encoding is the same as that in the reference image encoding, and it is considered that the to-be-recognized image encoding and the reference image encoding are repeated.
[0111] In some embodiments, the byte of the to-be-recognized image encoding corresponds to the byte in the same position in the reference image encoding. Figure 6C The step S1032B shown can be implemented through the following steps: Detect whether each byte in the to-be-recognized image encoding is the same as the byte in the corresponding position in the reference image encoding. If each byte in the to-be-recognized image encoding is the same as the byte in the corresponding position in the reference image encoding, determine that the check result of the to-be-recognized image frame is the first type of check result. The first type of check result is used to indicate that there is repeatability between the to-be-recognized image frame and the reference image frame. Or, if there is at least one byte in the to-be-recognized image encoding that is different from the byte in the corresponding position in the reference image encoding, determine that the check result of the to-be-recognized image frame is the second type of check result. The second type of check result is used to indicate that there is no repeatability between the to-be-recognized image frame and the reference image frame.
[0112] In the embodiments of the present application, for any image to be recognized for encoding, the encoded image to be recognized can be byte-compared with each reference image encoding in sequence to detect whether each byte in the encoded image to be recognized is exactly the same as the corresponding byte in a certain reference image encoding, so as to obtain the verification result of the image frame to be recognized corresponding to the encoded image to be recognized. If each byte in the encoded image to be recognized is compared with at least one reference image encoding and the corresponding bytes are all the same, the verification result of the image frame to be recognized corresponding to the encoded image to be recognized is the first verification result. After obtaining the first verification result, a request to terminate the current liveness recognition and the result of the failure of the current liveness recognition are sent to the terminal. Alternatively, if each byte in the encoded image to be recognized is compared with each reference image encoding and there is at least one byte in each reference image encoding that is different from the corresponding byte in the encoded image to be recognized, a second verification result is obtained.
[0113] In the embodiments of the present application, by strongly matching the identity of each byte, the repeatability between the encoded image to be recognized and the reference image encoding is determined, avoiding the situation of misidentification injection attacks and improving the accuracy of liveness recognition.
[0114] In some embodiments, when performing repeatability verification, only the method of extracting the comparison matrix data of the key image region described above can be used, or only the method of comparing bytes after encoding described above can be used, or the method of extracting the comparison matrix data of the key image region described above and the method of comparing bytes after encoding described above can be used simultaneously.
[0115] In some embodiments, before Figure 6A the step S104 shown, the following steps can also be executed:
[0116] Based on the image frame to be recognized and the verification image frame, the target part of the target object is recognized to obtain the recognition result of the target part of the target object.
[0117] In an embodiment of the present application, a method for identifying a target part of a target object can adopt a silent living body. The verification image frame is verification information uploaded in advance by the target object for verifying the identity of the target object. The verification image frame is an image obtained by pre-capturing the target part. Taking the target object as a user as an example, the target part can be a face, palm, finger, eye and other parts. It should be noted that the embodiment of the present application does not limit the specific method for identifying the target part, which is similar to the face recognition method in the related art. Based on one or more image frames to be identified and all reference image frames in the database, the similarity between the image frame to be identified and the reference image frame can be determined, and then the target part recognition result of the target object can be obtained according to the similarity and a preset similarity threshold. When the similarity is greater than or equal to the similarity threshold, a target part recognition result that passes the recognition is obtained. Alternatively, when the similarity is less than the similarity threshold, a target part recognition result that fails the recognition is obtained.
[0118] Step S104: determining a liveness recognition result of the target object in the image frame to be recognized based on the verification result.
[0119] In the embodiment of the present application, a liveness recognition result of the target object can be obtained based on the recognition result of the target part of the target object and the verification result of the image frame to be recognized.
[0120] The embodiment of the present application obtains a reference image frame corresponding to the terminal based on the identification information carried in the liveness recognition request, performs a repeatability check on the image frame to be recognized based on the reference image frame, obtains a verification result of the image frame to be recognized, and determines the liveness recognition result of the target object based on the verification result. Therefore, using image frames with the same identification information as reference image frames has a huge advantage in terms of data volume. At the same time, performing a repeatability check on the image frame to be recognized based on the reference image frame fully utilizes the characteristic of the injection attacker recycling attack materials, ensuring that recycled attack materials cannot pass the repeatability check, thereby improving the accuracy of liveness recognition.
[0121] In some embodiments, Figure 6A The illustrated step S104 can also be implemented by the following steps: when the target part recognition result is recognition passed and the verification result is that there are no repeated frames of the image to be recognized, determining that the liveness recognition result of the target object is liveness recognition passed.
[0122] In an embodiment of the present application, if the verification result of an image frame to be identified indicates that there is a duplicate frame of the image frame to be identified, then regardless of whether the target part recognition result of the target object is passed or failed, a liveness recognition result for the target object indicating that liveness recognition failed is obtained. Alternatively, if the verification result of each image frame to be identified indicates that there are no duplicate frames of each image frame to be identified, and the target part recognition result of the target object is passed, then a liveness recognition result for the target object indicating that liveness recognition passed is obtained. The liveness recognition result for the target object is sent to the terminal, and the liveness recognition result is displayed to the target object.
[0123] The embodiment of the present application further determines the liveness recognition result of the target object by combining the repeatability verification result with the target part recognition result, thereby improving the accuracy of liveness recognition.
[0124] Figure 12 This is a flow chart of the liveness recognition method provided in the embodiment of the present application. Figure 12 The steps shown are explained as Figure 12 As shown, the liveness recognition method is described by taking the terminal as an example. The method includes the following steps S201 to S204:
[0125] Step S201 : In response to a liveness recognition operation initiated by a user, real-time acquisition is performed to obtain an image frame to be recognized.
[0126] In an embodiment of the present application, after a user performs a liveness recognition operation through a terminal, the terminal can respond to the liveness recognition operation initiated by the user by using an image acquisition device (e.g., a camera) to capture a real-time video of the target object at a fixed frequency of 10 frames per second (10fps, Frame Per Second), thereby obtaining multiple image frames to be recognized. The target object may be the same person as the user, or a different person. The real-time video of the target object may be injected with a video provided by an attacker, etc.
[0127] Step S202 : performing repeated frame detection on the image frame to be identified to obtain a detection result of the image frame to be identified.
[0128] In the embodiment of the present application, the terminal can perform duplicate frame detection on each image frame to be identified after capturing it, and obtain the detection result corresponding to the image frame to be identified. Specifically, local duplicate frame detection can be performed on the image frame to be identified based on the historical image frames to be identified in the local cache of the terminal. The process of performing local duplicate frame detection on the image frame to be identified is as follows: Figure 7 As shown, the process and method are the same as the method for the server to perform repeatability verification, and will not be repeated here.
[0129] Step S203, in the case where the detection result indicates that there is no duplicate frame of the image frame to be recognized, send a liveness recognition request to the server.
[0130] The liveness recognition request carries the image frame to be recognized, and the liveness recognition request is used to request the server to perform a repeatability check on the image frame to be recognized to obtain the check result of the image frame to be recognized.
[0131] In the embodiment of the present application, if the detection result corresponding to the image frame to be recognized indicates that there is no duplicate frame of the image frame to be recognized, the image frame to be recognized is constructed into a liveness recognition request and sent to the server. At the same time, the image frame to be recognized is stored in the local cache of the terminal to facilitate local duplicate frame detection for the next image frame to be recognized. Alternatively, if the detection result corresponding to the image frame to be recognized indicates that there is a duplicate frame of the image frame to be recognized, the local liveness recognition is directly terminated, and an alarm message indicating that a duplicate frame has been detected and there may be a risk of injection attack is sent to the target object.
[0132] Step S204, receive the check result returned by the server, and determine the liveness recognition result of the target object included in the image frame to be recognized based on the check result.
[0133] In the embodiment of the present application, the target part of the target object can be recognized first based on the image frame to be recognized and the verification image frame to obtain the target part recognition result of the target object. The verification image frame is the verification information uploaded in advance. After receiving the check result returned by the server, the liveness recognition result of the target object included in the image frame to be recognized is obtained based on the target part recognition result of the target object and the check result of the image frame to be recognized.
[0134] In the embodiment of the present application, by combining the repeatability check result with the target part recognition result, the liveness recognition result of the target object is further determined, improving the accuracy of liveness recognition.
[0135] Next, the liveness recognition method in the embodiment of the present application will be described in combination with the interaction between the terminal and the server in the liveness recognition system. It should be noted that the liveness recognition method here is a liveness recognition method implemented through the interaction between the terminal and the server, which is substantially the same as the liveness recognition method executed by the server or the terminal in the above embodiment. The only difference is that in the embodiment of the present application, the actions performed by the terminal during the execution of the liveness recognition method are also described. Moreover, some steps can be executed by either the terminal or the server. Therefore, for the steps that are the same as those in the above embodiment but have different execution entities in this embodiment, this embodiment is only an exemplary illustration, and in the implementation process, any one of the execution entities can execute them, and the embodiment of the present application does not make any limitations in this regard.
[0136] Figure 13Another optional process schematic diagram of the living body recognition method provided by the embodiments of the present application is as follows Figure 13 As shown, the method includes the following steps S301 to S312:
[0137] Step S301, the terminal responds to the living body recognition operation and collects multiple to-be-recognized image frames in real time.
[0138] In the embodiments of the present application, when the user initiates a living body recognition operation, the terminal responds to the living body recognition operation and collects multiple to-be-recognized image frames by performing video frame acquisition on the real-time video of the target object at a fixed frequency (for example, 10ps).
[0139] Step S302, for each to-be-recognized image frame, the terminal performs duplicate frame detection on the to-be-recognized image frame to obtain the detection result of the to-be-recognized image frame.
[0140] In the embodiments of the present application, after each to-be-recognized image frame is collected, the terminal performs a round of local duplicate frame detection on the to-be-recognized image frame to obtain the detection result of the to-be-recognized image frame. Alternatively, the terminal can also perform duplicate frame detection on the to-be-recognized image frame online.
[0141] Step S303, the terminal sends a living body recognition request to the server. Among them, the living body recognition request carries the to-be-recognized image frame.
[0142] In the embodiments of the present application, after each to-be-recognized image frame is collected, if the detection result of the to-be-recognized image frame indicates that there is no duplicate frame in the to-be-recognized image frame, the terminal sends a living body recognition request to the server.
[0143] Step S304, the terminal displays an alarm message to the user.
[0144] In the embodiments of the present application, after each to-be-recognized image frame is collected, if the detection result of the to-be-recognized image frame indicates that there is a duplicate frame in the to-be-recognized image frame, the terminal terminates the current living body recognition, displays that there is a duplicate frame to the user, and displays an alarm message indicating that there may be a risk of injection attack.
[0145] Step S305, the server obtains a reference image frame corresponding to the to-be-recognized image frame.
[0146] In the embodiments of the present application, after receiving the living body recognition request sent by the terminal, the server can respond to the living body recognition request, establish a connection with the terminal, and obtain a reference image frame corresponding to the to-be-recognized image frame from the database based on the request identifier in the living body recognition request. After the server establishes a connection with the terminal, when receiving a living body recognition request later, it is not necessary to repeat the above step of obtaining a reference image frame corresponding to the terminal from the database based on the request identifier in the living body recognition request.
[0147] Step S306: The server performs a repeatability check on the image frame to be recognized based on the reference image frame, and obtains the check result of the image frame to be recognized.
[0148] In the embodiment of the present application, after receiving each image frame to be recognized, the server performs a repeatability check on the image frame to be recognized, and obtains the check result of the image frame to be recognized.
[0149] Step S307: The server uploads the image frame to be recognized to the database.
[0150] In the embodiment of the present application, after performing a repeatability check on each image frame to be recognized, if the detection result of each image frame to be recognized indicates that there is no duplicate frame in the image frame to be recognized, then all the image frames to be recognized are uploaded to the database.
[0151] Step S308: The server sends an alarm message to the terminal.
[0152] In the embodiment of the present application, after receiving each image frame to be recognized, if the check result of the image frame to be recognized indicates that there is a duplicate frame in the image frame to be recognized, then the current liveness recognition is terminated, and an alarm message indicating that there is a duplicate frame and there may be a risk of injection attack is sent to the terminal.
[0153] Step S309: The server performs target part recognition on the target object based on the image frame to be recognized and the verification image frame, and obtains the target part recognition result of the target object.
[0154] In the embodiment of the present application, after performing a repeatability check on each image frame to be recognized, if the detection result of each image frame to be recognized indicates that there is no duplicate frame in the image frame to be recognized, then the server performs target part recognition on the target object based on the image frame to be recognized and the verification image frame, and obtains the target part recognition result of the target object.
[0155] Step S310: The server determines the liveness recognition result of the target object based on the check result and the target part recognition result.
[0156] In the embodiment of the present application, when the target part recognition result is passed and the check result is that there is no duplicate frame of the image frame to be recognized, it is determined that the liveness recognition result of the target object is that the liveness recognition is passed.
[0157] Step S311: The server sends the liveness recognition result of the target object to the terminal.
[0158] Step S312: The terminal displays the liveness recognition result to the user.
[0159] In another embodiment, step S309 and step S310 may also be executed by the terminal, and the specific actions executed are the same as those executed by the server, which will not be repeated here.
[0160] The following describes an exemplary application of the embodiments of the present application in a practical application scenario.
[0161] The solution of the embodiment of the present application can be applied to the "financial identity verification" commonly used in the financial industry, that is, the liveness detection method of the face category. It is usually used before applying for credit as the first step in regulatory identity verification requirements, and can effectively reduce the risks of some common attacks such as photocopying, recording, and injection.
[0162] The embodiments of the present application solve the problems of poor user experience such as “direct strong light” and “long time-consuming whole process” in related technologies. The embodiments of the present application also solve the problem of poor security of “inability to completely defend against injection attacks” in related technologies. The liveness recognition system provided by the embodiments of the present application can be composed of a client, a server and a database. The specific implementation steps are as follows: When the user initiates a liveness recognition operation, the client collects user image frames at a fixed frequency of 10fps, performs a round of local duplicate frame detection after each frame is collected, and stores the image frame to be identified in the local cache, and then transmits the image frame to be identified to the server for further judgment. The local duplicate frame detection of the client does not obtain historical data from the database. Taking performance factors into consideration, the initial screening results can be obtained as soon as possible, thereby improving the speed of identifying injection attacks. However, performing local duplicate frame detection only on the client may result in insufficient number of duplicate frame data collection, resulting in missed detections. When the number of collected frames is insufficient, the number of duplicate frame data collected may be insufficient. When the number of frames is greater than the attack data itself, there can be a 100% defense rate, so server detection is needed to expand the scale. However, both transmission and retrieval have delays (more than 500ms), so only performing repeatability checks on the server will result in an insufficient response speed. After the server establishes a link, it first pulls all the reference image frames of the ID in stock according to the request identifier (ID) and loads them into the local memory. After receiving each frame of the image frame to be identified from the client, it performs a round of repeatability checks and stores the image frame to be identified in the local cache. When the server receives the last frame of the image frame to be identified, it synchronizes all cached frames in the local cache to the database and returns the final liveness recognition result to the client. If a duplicate frame is detected in the middle, it ends directly and alerts the user. The detection process of duplicate frames by both the client and the server can be as follows Figure 7 As shown, the entire process is connected in series based on two core judgments: "whether the frame is repeated" and "whether it is the last frame".
[0163] During the duplicate frame check process for the image frame to be identified, a coding engine can be used to compress the image frame to be identified into a unique code, and duplicate frame detection can be achieved by detecting the repetitiveness of the unique code. For example, the coding engine can use MD5 to compress the image frame to be identified into a unique code with a total length of 128 bits. If SHA1 is used, the image frame to be identified can be compressed into a unique code with a total length of 160 bits. If SHA256 is used, the image frame to be identified can be compressed into a unique code with a total length of 256 bits.
[0164] To improve the robustness of the liveness detection system, a key image region extraction algorithm (such as Fast R-CNN and Mask R-CNN) can be introduced during repeated frame verification of the image frames to be identified. This algorithm includes a semantic segmentation engine and a key image region classifier (such as ResNet50, Yolov5, and RetinaFace). First, the semantic segmentation engine performs semantic segmentation on each image frame to be identified, dividing it into multiple regions based on the actual image content, such as face, body, background, and object regions. The key image region classifier then scores each region and selects the top M regions as key image regions. The reference image frame is split into three RGB channels, forming a three-channel matrix data A, which can be understood as a large cube. Each key image region is similarly split into a three-channel matrix data B, which can be understood as a small cube. Following a prescribed search sequence, the small cube is checked to see if it shares a pattern with the large cube. Finally, if the key image area (small cube) of the current image frame to be identified exists in a reference image frame (large cube), it is determined that there is a partial frame duplication phenomenon, indicating that the liveness detection is an injection attack. If no partial duplication is found, the final result is returned based on the silent liveness judgment. If the silent liveness is considered to be passed, it is passed; if the silent liveness is rejected, it is rejected.
[0165] The above-mentioned coding engine and key image area extraction algorithm are two different independent duplicate frame detection methods. You can use only one of them or use both together. In the embodiment of the present application, the client uses the coding engine to perform local duplicate frame detection, and the server uses a hybrid mode (coding engine and key image area extraction algorithm).
[0166] It should be noted that the key image region extraction algorithm and the key image region matrix search algorithm in the embodiments of the present application may be replaced by other algorithms with the same purpose but heterogeneous models.
[0167] The visualization interface of the living body recognition method provided in the embodiment of the present application can be as follows: Figure 14 As shown. Figure 14As shown, by using the living body recognition method provided in the embodiments of the present application, 18 repeated image frames can be detected in a total living body detection of ten seconds once.
[0168] Based on local repeated frame detection, the embodiments of the present application perform living body recognition, which can be combined with the excellent silent living body technology to return results within 1000 ms. Moreover, the silent living body technology itself does not require users to respond to challenges, so users do not need to make additional cooperation, greatly improving the user experience and thus increasing the overall living body passing rate. The embodiments of the present application make full use of the characteristic that attackers inject and recycle attack materials. All historical frames are stored instead of only key frames, which has a huge advantage in terms of data volume. In addition, the retrieval repetition range is framed in the local key image area instead of the whole frame, which greatly improves the recall speed and accuracy of attack materials, realizing the fast and efficient detection of injected attacks. From the user's perspective, the embodiments of the present application only require the user to face the camera and keep still for 500 ms, greatly reducing the user operation complexity and improving the user experience; from the engine perspective, each frame of the user's local key image area has never appeared before, which ensures the real-time nature of the image and realizes the resistance to injected attacks.
[0169] It can be understood that in the embodiments of the present application, data related to user information and the like are involved. When the embodiments of the present application are applied to specific products or technologies, user permission or consent needs to be obtained, and the collection, use, and processing of relevant data need to comply with relevant laws, regulations, and standards in relevant countries and regions.
[0170] Next, the implementation of the living body recognition device 455 provided in the embodiments of the present application as a software module will be further described. In some embodiments, as Figure 5B shown, the software module stored in the living body recognition device 455 in the memory 440 may include:
[0171] A request receiving module 4551, configured to receive a living body recognition request sent by a terminal; the living body recognition request includes an image frame to be recognized, and the image frame to be recognized is a non-repeated image frame; a reference image obtaining module 4552, configured to obtain a reference image frame of the image frame to be recognized; a repeatability verification module 4553, configured to perform repeatability verification on the image frame to be recognized based on the reference image frame to obtain a verification result of the image frame to be recognized; and a recognition result determination module 4554, configured to determine a living body recognition result of a target object in the image frame to be recognized based on the verification result.
[0172] In some embodiments, the request receiving module 4551 is further configured to receive a living body recognition request sent by the terminal when the detection result of the terminal for detecting repeated frames of the image frame to be recognized is a first detection result; the first detection result is that there is no image frame in the terminal that is repeated with the image frame to be recognized.
[0173] In some embodiments, the repeatability verification module 4553 is further configured to perform semantic segmentation on the image frame to be recognized, obtaining multiple image regions; acquiring the score of each image region; determining the key image region of the image frame to be recognized according to the scores of the multiple image regions; and determining the verification result of the image frame to be recognized according to the key image region and the reference image frame.
[0174] In some embodiments, the repeatability verification module 4553 is further configured to acquire the image matrix data of the key image region and the reference matrix data of the reference image frame; the data dimension of the image matrix data is less than that of the reference matrix data; if the image matrix data belongs to the sub-matrix data of the reference matrix data, determining that the verification result of the image frame to be recognized is the first verification result; if the image matrix data does not belong to the sub-matrix data of the reference matrix data, determining that the verification result of the image frame to be recognized is the second verification result.
[0175] In some embodiments, the repeatability verification module 4553 is further configured to acquire the pixel values of the pixels in the key image region in the target color channel; and obtaining the image matrix data of the key image region based on the pixel values.
[0176] In some embodiments, the repeatability verification module 4553 is further configured to perform encoding processing on the image frame to be recognized and the reference image frame respectively, obtaining the image encoding to be recognized of the image frame to be recognized and the reference image encoding of the reference image frame; and performing repeatability verification on the image encoding to be recognized based on the reference image encoding, obtaining the verification result of the image frame to be recognized.
[0177] In some embodiments, the bytes in the image encoding to be recognized correspond to the bytes at the same positions in the reference image encoding. The repeatability verification module 4553 is further configured to detect whether each byte in the image encoding to be recognized is the same as the byte at the corresponding position in the reference image encoding; if each byte in the image encoding to be recognized is the same as the byte at the corresponding position in the reference image encoding, determining that the verification result of the image frame to be recognized is the first type of verification result; the first type of verification result is used to characterize that there is repeatability between the image frame to be recognized and the reference image frame; if there is at least one byte in the image encoding to be recognized that is different from the byte at the corresponding position in the reference image encoding, determining that the verification result of the image frame to be recognized is the second type of verification result; the second type of verification result is used to characterize that there is no repeatability between the image frame to be recognized and the reference image frame.
[0178] In some embodiments, the liveness recognition device 455 also includes a target part recognition module, which is used to identify the target part of the target object based on the image frame to be identified and the verification image frame to obtain the target part recognition result of the target object; the recognition result determination module is also used to determine that the liveness recognition result of the target object is liveness recognition passed when the target part recognition result is recognition passed and the verification result is that there are no repeated frames of the image to be identified.
[0179] The present invention provides a computer program product comprising a computer program or computer-executable instructions stored in a computer-readable storage medium. A processor of an electronic device reads the computer-executable instructions from the computer-readable storage medium and executes the computer-executable instructions, causing the electronic device to perform the liveness detection method described in the present invention.
[0180] The embodiment of the present application provides a computer-readable storage medium storing computer-executable instructions, wherein the computer-executable instructions or computer programs are stored. When the computer-executable instructions or computer programs are executed by a processor, the processor will execute the living body recognition method provided by the embodiment of the present application, for example, Figure 6A The living body recognition method is shown.
[0181] In some embodiments, the computer-readable storage medium may be a memory such as RAM, ROM, flash memory, magnetic surface memory, optical disk, or CD-ROM; or may be various devices including one or any combination of the above memories.
[0182] In some embodiments, computer-executable instructions may be in the form of a program, software, software module, script, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment.
[0183] As an example, computer-executable instructions may, but need not, correspond to a file in a file system, may be stored as part of a file that stores other programs or data, such as, for example, in one or more scripts in a HyperText Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple coordinating files (e.g., files storing one or more modules, subroutines, or code portions).
[0184] By way of example, computer-executable instructions may be deployed to be executed on one electronic device, or on multiple electronic devices located at one site, or on multiple electronic devices distributed across multiple sites and interconnected by a communication network.
[0185] In summary, the liveness recognition method provided in the embodiment of the present application discloses a liveness recognition scheme based on local repeated frame detection. Related liveness recognition methods have major defects in user experience and security. For example, the overall user cooperation of motion liveness takes a long time and cannot defend against injection attacks. The colorful liveness has a dazzling problem and cannot completely resist injection attacks. Therefore, the embodiment of the present application builds a library of liveness frames for each request ID. When a new liveness request is initiated, it checks whether the key image area of each frame already exists in the past liveness frame. If frame duplication occurs in the key image area, it is determined that the user has been subjected to an injection attack. If no duplication occurs, it is further determined whether to release it in combination with the silent liveness recognition result.
[0186] The above description is merely an embodiment of the present application and is not intended to limit the scope of protection of the present application. Any modifications, equivalent replacements, and improvements made within the spirit and scope of the present application are included in the scope of protection of the present application.
Claims
1. A method for live body recognition, characterized in that, The method includes: Receiving a live recognition request sent by a terminal; the live recognition request includes an image frame to be recognized, and the image frame to be recognized is a non-repeating image frame; Obtaining a reference image frame of the image frame to be recognized; Performing a repeatability check on the image frame to be recognized based on the reference image frame to obtain a check result of the image frame to be recognized; Determining a live recognition result of a target object in the image frame to be recognized based on the check result.
2. The method according to claim 1, characterized in that, The receiving of the live recognition request sent by the terminal includes: Receiving the live recognition request sent by the terminal when a detection result of detecting duplicate frames for the image frame to be recognized at the terminal is a first detection result; the first detection result is that there is no image frame in the terminal that duplicates the image frame to be recognized; Wherein, the terminal performing a repeat frame detection on the image frame to be recognized includes: Obtaining an image frame to be recognized at the current moment; Performing a repeat frame detection on the image frame to be recognized based on historical image frames to be recognized stored in the terminal to obtain a detection result.
3. The method according to claim 1, characterized in that The performing a repeatability check on the image frame to be recognized based on the reference image frame to obtain a check result of the image frame to be recognized includes: Performing semantic segmentation on the image frame to be recognized to obtain a plurality of image regions; Obtaining a score for each image region; Determining a key image region of the image frame to be recognized according to the scores of the plurality of image regions; Determining the check result of the image frame to be recognized according to the key image region and the reference image frame.
4. The method according to claim 3, wherein The check result includes a first check result and a second check result, and the determining the check result of the image frame to be recognized according to the key image region and the reference image frame includes: Obtaining image matrix data of the key image region and reference matrix data of the reference image frame; the data dimension of the image matrix data is less than the data dimension of the reference matrix data; If the image matrix data belongs to sub-matrix data of the reference matrix data, determining that the check result of the image frame to be recognized is the first check result; If the image matrix data does not belong to sub-matrix data of the reference matrix data, determining that the check result of the image frame to be recognized is the second check result.
5. The method according to claim 4, wherein The obtaining the image matrix data of the key image region includes: Obtaining pixel values of pixels in the key image region in a target color channel; Based on the pixel values, obtaining the image matrix data of the key image region.
6. The method according to claim 1, wherein The performing a repeatability check on the image frame to be recognized based on the reference image frame to obtain a check result of the image frame to be recognized includes: Performing encoding processing on the image frame to be recognized and the reference image frame respectively to obtain a to-be-recognized image encoding of the image frame to be recognized and a reference image encoding of the reference image frame; Performing a repeatability check on the to-be-recognized image encoding based on the reference image encoding to obtain a check result of the image frame to be recognized.
7. The method according to claim 6, characterized in that, The bytes of the to-be-recognized image encoding correspond to the bytes at the same positions in the reference image encoding; the obtaining of the verification result of the to-be-recognized image frame by performing a repeatability verification on the to-be-recognized image encoding based on the reference image encoding includes: Detecting whether each byte in the to-be-recognized image encoding is the same as the byte at the corresponding position in the reference image encoding; If each byte in the to-be-recognized image encoding is the same as the byte at the corresponding position in the reference image encoding, determining that the verification result of the to-be-recognized image frame is a first type of verification result; the first type of verification result is used to characterize that there is repeatability between the to-be-recognized image frame and the reference image frame; If there is at least one byte in the to-be-recognized image encoding that is different from the byte at the corresponding position in the reference image encoding, determining that the verification result of the to-be-recognized image frame is a second type of verification result; the second type of verification result is used to characterize that there is no repeatability between the to-be-recognized image frame and the reference image frame.
8. The method according to any one of claims 1 to 7, characterized in that Before determining the liveness recognition result of the target object in the to-be-recognized image frame based on the verification result, the method further includes: Performing target part recognition on the target object based on the to-be-recognized image frame and the verification image frame to obtain the target part recognition result of the target object; The determining of the liveness recognition result of the target object in the to-be-recognized image frame based on the verification result includes: When the target part recognition result is recognition passed and the verification result is that there is no repeated frame of the to-be-recognized image, determining that the liveness recognition result of the target object is liveness recognition passed.
9. A living body recognition device, characterized in that, The apparatus includes: A request receiving module, configured to receive a liveness recognition request sent by a terminal; the liveness recognition request includes a to-be-recognized image frame, and the to-be-recognized image frame is a non-repeated image frame; A reference image obtaining module, configured to obtain a reference image frame of the to-be-recognized image frame; A repeatability verification module, configured to perform a repeatability verification on the to-be-recognized image frame based on the reference image frame to obtain the verification result of the to-be-recognized image frame; A recognition result determining module, configured to determine the liveness recognition result of the target object in the to-be-recognized image frame based on the verification result.
10. An electronic device, characterized in that, The electronic device includes: A memory, configured to store computer-executable instructions; A processor, configured to implement the liveness recognition method according to any one of claims 1 to 8 when executing the computer-executable instructions stored in the memory.
11. A computer-readable storage medium storing computer-executable instructions or a computer program, characterized in that, The computer-executable instructions or the computer program, when executed by the processor, implement the liveness recognition method according to any one of claims 1 to 8.
12. A computer program product, comprising computer-executable instructions or a computer program, characterized in that, The computer-executable instructions or the computer program, when executed by the processor, implement the liveness recognition method according to any one of claims 1 to 8.