Living body recognition method and device, electronic equipment and storage medium

By combining the dual verification of live body detection and environmental safety detection on the computer, the problem that live body recognition on the computer cannot defend against injection attacks is solved, and high security and high accuracy live body recognition is achieved, improving the user experience.

CN120340141APending Publication Date: 2025-07-18MASHANG CONSUMER FINANCE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410072666.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-16
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

In the prior art, the computer's live body recognition method cannot effectively defend against injection attacks, and the user experience is poor, especially the mobile live body is time-consuming and easy to be attacked. The colorful live body fails under strong light conditions, and the mobile injection detection program cannot be transplanted to the PC side for use.

Method used

By obtaining operating environment parameters while collecting video frames at the terminal, using independent environmental protection process detection and injection attacks, and combining live detection results for aggregate processing, the comprehensive judgment of live recognition results is achieved.

Benefits of technology

Effectively defend against injected attacks on the computer, improve the accuracy and security of live body recognition, while reducing user operation complexity and improving user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120340141A_ABST
    Figure CN120340141A_ABST
Patent Text Reader

Abstract

The invention provides a living body recognition method and device, electronic equipment and a storage medium. The method comprises the steps of receiving a living body detection request sent by a terminal; the living body detection request comprises a living body detection time sequence, the living body detection time sequence comprises video frames of a target object collected by the terminal in real time in a first time interval, and the first time interval is a time interval for collecting the video frames; determining operation environment parameters of the terminal in the first time interval; determining an environment safety detection time sequence of the terminal based on the operation environment parameters of the terminal, wherein the environment safety detection time sequence comprises an environment detection frame; and carrying out aggregation processing on the video frame of the living body detection time sequence and the environment detection frame of the environment safety detection time sequence to obtain a living body identification result of the target object. Through the method and the device, injection attacks can be defended during living body recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of face recognition technology, and in particular, to a method and device for live body recognition, an electronic device, and a storage medium. Background Art

[0002] Live body recognition is a technology for identifying whether a user is a live body rather than a machine, and is generally integrated into the login end of a business application (APP, Application). Injection attack is a common attack method against live body recognition. Action live body and colorful live body are common live body recognition methods for resisting injection attacks. However, action live body takes a long time, and due to the limited number of action combinations, it is easily exhausted and broken by injection attacks. Colorful live body can only defend against injection attacks under the condition of strictly requiring "colorful" feedback.

[0003] In related technologies, an injection detection program is used on the mobile side to detect injection attacks, but the injection detection program cannot be used on the computer side (PC, Personal Computer). Therefore, the current live body recognition method on the computer side cannot fully defend against injection attacks and has major defects in terms of security. Summary of the Invention

[0004] Embodiments of this application provide a method and device for live body recognition, an electronic device, and a storage medium, which can defend against injection attacks when performing live body recognition.

[0005] The technical solution of the embodiments of this application is implemented as follows:

[0006] Embodiments of this application provide a method for live body recognition, and the method includes:

[0007] Receiving a live body detection request sent by a terminal; the live body detection request includes a live body detection time series, and the live body detection time series includes: video frames of a target object collected in real time by the terminal within a first time interval, and the first time interval is a time interval for collecting video frames; determining operating environment parameters of the terminal within the first time interval; determining an environment security detection time series of the terminal based on the operating environment parameters of the terminal, and the environment security detection time series includes environment detection frames; performing aggregation processing on the video frames of the live body detection time series and the environment detection frames of the environment security detection time series to obtain a live body recognition result of the target object.

[0008] An embodiment of the present application provides a living body recognition device, including: a request receiving module, configured to receive a living body detection request sent by a terminal; the living body detection request includes a living body detection time series, and the living body detection time series includes: video frames of a target object collected in real time by the terminal within a first time interval, and the first time interval is a time interval for collecting video frames; an environmental parameter determination module, configured to determine the operating environmental parameters of the terminal within the first time interval; an environmental series determination module, configured to determine an environmental security detection time series of the terminal based on the operating environmental parameters of the terminal, and the environmental security detection time series includes environmental detection frames; an aggregation processing module, configured to perform aggregation processing on the video frames of the living body detection time series and the environmental detection frames of the environmental security detection time series to obtain a living body recognition result of the target object.

[0009] An embodiment of the present application provides an electronic device, where the electronic device includes: a memory, configured to store computer-executable instructions; a processor, configured to implement the living body recognition method provided by the embodiment of the present application when executing the computer-executable instructions stored in the memory.

[0010] An embodiment of the present application provides a computer-readable storage medium, storing a computer program or computer-executable instructions, which are configured to implement the living body recognition method provided by the embodiment of the present application when being executed by a processor.

[0011] The embodiment of the present application has the following beneficial effects:

[0012] The embodiment of the present application obtains the operating environmental parameters of the terminal within the first time interval, determines the environmental security detection time series of the terminal based on the operating environmental parameters, and performs aggregation processing on the video frames of the living body detection time series and the environmental detection frames of the environmental security detection time series to obtain a living body recognition result of the target object. Wherein, the first time interval is a time interval for collecting video frames, and the living body detection time series includes video frames of the target object collected in real time by the terminal within the first time interval. Since when the terminal is under an injection attack, the operating environmental parameters of the terminal will definitely change, and the collection of operating environmental parameters is carried out simultaneously with the collection of video frames, the obtained environmental security detection time series can be used to detect the injection attack risk during the video frame collection process. The embodiment of the present application can comprehensively judge whether the target object is a living body behavior through both living body detection and environmental security detection, realizes the defense against injection attacks during the living body recognition process, and improves the accuracy of living body recognition. Description of the Drawings

[0013] Figure 1 is a schematic diagram of the action sequence of an action living body in the related art;

[0014] Figure 2 is a schematic diagram of the interaction interface of an action living body in the related art;

[0015] Figure 3 It is a schematic diagram of the interaction interface of the colorful living body in the related art;

[0016] Figure 4 It is a schematic structural diagram of the living body recognition system architecture provided by the embodiment of the present application;

[0017] Figure 5 It is a schematic structural diagram of the living body recognition device provided by the embodiment of the present application;

[0018] Figure 6A It is a schematic flowchart of the living body recognition method provided by the embodiment of the present application;

[0019] Figure 6B It is a schematic flowchart of the living body recognition method provided by the embodiment of the present application;

[0020] Figure 6C It is a schematic flowchart of the living body recognition method provided by the embodiment of the present application;

[0021] Figure 6D It is a schematic flowchart of the living body recognition method provided by the embodiment of the present application;

[0022] Figure 6E It is a schematic flowchart of the living body recognition method provided by the embodiment of the present application;

[0023] Figure 7 It is a schematic diagram of the video frame provided by the embodiment of the present application;

[0024] Figure 8 It is a schematic structural diagram of the living body recognition system provided by the embodiment of the present application;

[0025] Figure 9 It is a schematic diagram of the environment detection frame provided by the embodiment of the present application;

[0026] Figure 10 It is a schematic diagram of the process information provided by the embodiment of the present application;

[0027] Figure 11 It is a schematic diagram of the file directory information provided by the embodiment of the present application;

[0028] Figure 12 It is a schematic diagram of the semantic analysis engine provided by the embodiment of the present application;

[0029] Figure 13 It is a schematic flowchart of the aggregation operation provided by the embodiment of the present application;

[0030] Figure 14 It is a schematic flowchart of the video frame processing provided by the embodiment of the present application;

[0031] Figure 15 It is a schematic flowchart of environmental detection frame processing provided by an embodiment of the present application;

[0032] Figure 16 It is a schematic flowchart of aggregation operation provided by an embodiment of the present application;

[0033] Figure 17 It is a schematic structural diagram of a living body recognition device provided by an embodiment of the present application;

[0034] Figure 18 It is a schematic flowchart of a living body recognition method provided by an embodiment of the present application;

[0035] Figure 19 It is a schematic diagram of an interaction interface of a living body recognition method provided by an embodiment of the present application. Detailed implementation manners

[0036] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be construed as limitations on the present application. All other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present application.

[0037] In the following description, "some embodiments" are involved, which describe a subset of all possible embodiments. However, it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict.

[0038] In the following description, the terms "first / second / third" involved are only used to distinguish similar objects and do not represent a specific order for the objects. It can be understood that "first / second / third" can be interchanged with a specific order or sequence when allowed, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.

[0039] Unless otherwise defined, all technical and scientific terms used in the embodiments of the present application have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present application belongs. The terms used in the embodiments of the present application are only for the purpose of describing the embodiments of the present application and are not intended to limit the present application.

[0040] Before further elaborating on the embodiments of the present application, the nouns and terms involved in the embodiments of the present application are described. The nouns and terms involved in the embodiments of the present application are applicable to the following explanations.

[0041] 1) Living body recognition: A technology for identifying whether a user is a living body rather than a machine. The system structure usually consists of a collection end and a control end. The methods of living body recognition generally include silent living body, digital living body, colorful living body, motion living body, etc.

[0042] 2) Collection end: also known as the front end or client, a program used to interact directly with the user and collect data, usually deployed on the terminal for task execution.

[0043] 3) Control end: also known as the decision-making end or back-end or server, usually deployed on a cloud server for making decisions and judgments.

[0044] 4) Injection attack: refers to an attack method that uses hacking technology to replace the face image to be uploaded when the terminal collects the face image. For example, the injection attacker can control the camera data link of the face recognition system through hacking software, directly replace the camera collected image with the attack image, and then input the replaced attack image into the face recognition system.

[0045] Liveness recognition is a liveness detection device, technology and method deployed on a machine to implement a Turing test on the operator. Liveness recognition is a common functional module for a machine to determine whether the other party is a human. It usually does not exist independently, but is integrated into the login end of the business APP. It is also often used as a means of surprise inspection. For example, in daily life, the "youth anti-addiction mode" in the game uses liveness detection (face or verification code) to limit teenagers from being overly addicted to games. In the management of operating vehicles, the "driver's identity confirmation" uses liveness detection (face) to confirm the status of manual work. In various knowledge websites and APPs, "anti-crawler" uses liveness detection (verification code) to intercept crawlers.

[0046] Motion liveness is the most widely used liveness detection device, technology and method. It challenges the user to perform actions (such as Figure 1 As shown in the figure, the basic action sequence is a combination of opening the mouth, shaking the head left and right, shaking the head up and down, blinking, etc.), and the algorithm engine is used to identify and verify whether the user has completed the specified action challenge. The algorithm engine uses technologies such as facial key point positioning, face tracking, and action recognition to verify whether the user is a real living person. The core of action liveness is to implement challenges based on the randomness of actions, and assume that only living humans can complete the relevant action commands. Figure 2This is an example diagram of the action liveness interactive interface. The action liveness principle steps are as follows: the mobile application sends a random action request to the controller; the controller gives a random action command and starts the countdown of the action command life cycle; the mobile terminal sends an action prompt to the user according to the random action command and collects user action data; the user performs relevant actions and responds; the mobile terminal returns the media to the controller; the controller sends the media content to the algorithm for parsing and calculates whether it is consistent with the random action command. Action recognition has the following disadvantages: poor user experience, usually it takes more than 5 seconds to complete two rounds of actions; due to the limited combination of actions, it is easy for attackers to use injection attacks to exhaust various combinations of attack images to replace the user image collected by the mobile terminal, so that liveness recognition is successful, so it is impossible to defend against injection attacks.

[0047] Colorful liveness, also known as glare liveness or light liveness, proposes a liveness recognition algorithm based on light sequence recognition. It challenges the user with light and uses the algorithm to identify whether the corresponding light sequence appears on the user's face. The algorithm engine uses technologies such as facial key point positioning, face tracking, and color recognition to verify whether the user is truly alive. The core of Colorful Liveness is to implement challenges based on the color and frequency of light changes, and assumes that only faces can specifically feedback these light signal information. Figure 3 This is an example diagram of the colorful live interactive interface. The principle steps of colorful live are: the mobile application sends a random colorful password request to the controller; the controller gives a random colorful password and starts the countdown of the colorful password life cycle; the mobile terminal sends a colorful prompt to the user according to the random colorful password and collects user data; the user responds silently and waits for the colorful detection collection cycle to end; the mobile terminal sends the media back to the controller; the controller sends the media content to the algorithm for parsing and calculates whether it is consistent with the random colorful password. Action recognition has the following shortcomings: the system has poor robustness. Under strong light conditions during the day, the "colorful" screen cannot be effectively reflected from the target, so the module fails; the user experience is poor, and the high-intensity "colorful" will cause discomfort to the human eye; only under the condition of strict requirements for "colorful" feedback can "injection attacks" be defended.

[0048] It can be seen that the liveness recognition methods in the related technologies have major defects in user experience and security. For example, the overall user cooperation of the action liveness takes a long time and cannot defend against injection attacks. The colorful liveness has the problem of dazzling and cannot completely resist injection attacks. In addition, the injection detection program in the related technologies generally runs on the mobile terminal and cannot be directly ported to the PC terminal for use. If it is directly ported to the PC terminal for use, the injection detection program will be at the browser web page permission level, and the PC terminal cannot effectively identify the injection detection.

[0049] Based on at least one of the above problems existing in the related art, the embodiments of the present application provide a living body recognition method, device, electronic device, storage medium and program product, which can defend against injection attacks when performing living body recognition on a computer side and improve the accuracy of living body recognition. The following describes an exemplary application of the living body recognition device provided by the embodiments of the present application. The living body recognition device is an electronic device for implementing the living body recognition method. In one implementation, the living body recognition device (i.e., the electronic device) provided by the embodiments of the present application can be implemented as a terminal or a server. In one implementation, the living body recognition device provided by the embodiments of the present application can be implemented as various types of user terminals such as a laptop computer, a tablet computer, a desktop computer, a set-top box, a mobile device (e.g., a mobile phone, a portable music player, a personal digital assistant, a dedicated messaging device, a portable game device), a smart phone, a smart speaker, a smart watch, a smart TV, a vehicle-mounted terminal, etc.; in another implementation, the living body recognition device provided by the embodiments of the present application can also be implemented as a server. Among them, the server can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs, Content Delivery Network), and big data and artificial intelligence platforms. The terminal and the server can be directly or indirectly connected through wired or wireless communication methods, which are not limited in the embodiments of the present application. The following will describe an exemplary application when the living body recognition device is implemented as a server.

[0050] See Figure 4 , Figure 4 FIG. is an optional architecture diagram of the living body recognition system 100 provided by the embodiments of the present application. To support any exemplary application, the living body recognition system 100 includes at least a server 200, a network 300, and a terminal 400. The server 200 is a server for the living body recognition application. The server 200 can constitute the living body recognition device of the embodiments of the present application, that is, the living body recognition method of the embodiments of the present application is implemented through the server 200. The terminal 400 is connected to the server 200 through the network 300. The network 300 can be a wide area network, a local area network, or a combination of the two.

[0051] When performing live body recognition, the user can input a live body recognition operation through the terminal 400. In response to the live body recognition operation, the terminal collects video frames of the target object in real time within a first time interval and sends a live body detection request to the server 200. The server 200 receives the live body detection request sent by the terminal. The live body detection request includes a live body detection time series, and the live body detection time series includes the video frames of the target object collected by the terminal in real time within the first time interval. The first time interval is the time interval for collecting video frames. The server 200 determines the operating environment parameters of the terminal within the first time interval. The server 200 determines an environmental security detection time series of the terminal based on the operating environment parameters of the terminal. The environmental security detection time series includes environmental detection frames. The server 200 performs an aggregation process on the video frames of the live body detection time series and the environmental detection frames of the environmental security detection time series to obtain a live body recognition result of the target object. At the same time, after obtaining the live body recognition result, the server 200 can also send the live body recognition result to the terminal 400, and the terminal 400 displays the live body recognition result to the user.

[0052] See Figure 5 , Figure 5 is a schematic structural diagram of an electronic device provided by an embodiment of the present application. Figure 5 The electronic device shown can be a live body recognition device, and the live body recognition device includes: at least one processor 410, a memory 450, at least one network interface 420, and a user interface 430. Each component in the live body recognition device is coupled together through a bus system 440. It can be understood that the bus system 440 is used to realize the connection and communication between these components. The bus system 440 includes, in addition to a data bus, a power bus, a control bus, and a status signal bus. However, for the sake of clear illustration, in Figure 5 all kinds of buses are labeled as the bus system 440.

[0053] The processor 410 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor can be a microprocessor or any conventional processor, etc.

[0054] The user interface 430 includes one or more output devices 431 that enable the presentation of media content, including one or more speakers and / or one or more visual display screens. The user interface 430 also includes one or more input devices 432, including user interface components that facilitate user input, such as a keyboard, a mouse, a microphone, a touch screen display, a camera, and other input buttons and controls.

[0055] The memory 450 can be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state memories, hard disk drives, optical disc drives, etc. The memory 450 optionally includes one or more storage devices that are physically remote from the processor 410.

[0056] The memory 450 includes volatile memory, non-volatile memory, or both volatile and non-volatile memory. The non-volatile memory can be a read-only memory (ROM), and the volatile memory can be a random access memory (RAM). The memory 450 described in the embodiments of the present application is intended to include any suitable type of memory.

[0057] In some embodiments, the memory 450 is capable of storing data to support various operations. Examples of such data include programs, modules, and data structures, or subsets or supersets thereof, which are illustrated below.

[0058] The operating system 451 includes system programs for handling various basic system services and performing hardware-related tasks, such as a framework layer, a core library layer, a driver layer, etc., for implementing various basic services and handling hardware-based tasks; the network communication module 452 is used to reach other electronic devices via one or more (wired or wireless) network interfaces 420. Exemplary network interfaces 420 include: Bluetooth, Wi-Fi (Wireless Fidelity), and USB (Universal Serial Bus), etc.; the presentation module 453 is used to enable the presentation of information (e.g., a user interface for operating peripheral devices and displaying content and information) via one or more output devices 431 associated with the user interface 430 (e.g., a display screen, a speaker, etc.); the input processing module 454 is used to detect and translate one or more user inputs or interactions from one of one or more input devices 432.

[0059] In some embodiments, the device provided in the embodiments of the present application can be implemented in software. Figure 5 Shown is a living body recognition device 455 stored in the memory 450, which can be software in the form of a program, a plugin, etc., and includes the following software modules: a request receiving module 4551, an environmental parameter determination module 4552, an environmental sequence determination module 4553, and an aggregation processing module 4554. These modules are logical, and thus can be combined arbitrarily or further split according to the functions implemented. The functions of each module will be described below.

[0060] In some other embodiments, the device provided by the embodiments of the present application may be implemented in a hardware manner. As an example, the device provided by the embodiments of the present application may be a processor in the form of a hardware decoding processor, which is programmed to execute the living body recognition method provided by the embodiments of the present application. For example, the processor in the form of a hardware decoding processor may employ one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field programmable gate arrays (FPGAs), or other electronic components.

[0061] The living body recognition methods provided by the embodiments of the present application may be executed by an electronic device, where the electronic device may be a server or a terminal. That is, the living body recognition methods provided by the embodiments of the present application may be executed by the server, or may be executed by the terminal, or may also be executed through the interaction between the server and the terminal.

[0062] Figure 6A is a schematic flowchart of the living body recognition method provided by the embodiments of the present application. The following will be described in conjunction with Figure 6A the steps shown. As Figure 6A shown, taking the execution entity of the living body recognition method as the server as an example for description, the method includes the following steps S101 to step S104:

[0063] Step S101, receiving a living body detection request sent by the terminal.

[0064] The living body detection request includes a living body detection time series. The living body detection time series includes: video frames of the target object collected in real time by the terminal within a first time interval. The first time interval is the time interval for collecting the video frames.

[0065] In the embodiments of the present application, a video frame is a frame of an image containing a target object. The target object includes objects that need to be verified for liveness recognition in various scenarios, that is, the target object can be an object that initiates a liveness recognition operation on a terminal. For example, a user who needs to perform face recognition in a financial scenario. However, due to the existence of injection attacks, the object that initiates a liveness recognition operation on a terminal may not be the user himself / herself. At this time, the target object can be the object contained in the pictures or videos provided by the injection attacker. Exemplarily, in the absence of injection attacks, the target object is the user who initiates a liveness recognition operation on the terminal, and the video frame is a frame of an image obtained by the terminal through image acquisition of the real-time video of the target object, that is, a photo of the target object collected in real time by the terminal through an image acquisition device. Or, in the presence of injection attacks, a non-genuine user initiates a liveness recognition operation on the terminal, the video frame is a video frame obtained by real-time acquisition of the pictures, videos, etc. provided by the injection attacker, and the target object is the person contained in the pictures or videos provided by the injection attacker.

[0066] In the embodiments of the present application, the first time interval is the time interval for collecting all video frames after the terminal starts the liveness recognition process. The terminal can send a liveness detection request to the server in response to the liveness recognition operation. After receiving the liveness detection request sent by the terminal, the server can obtain a liveness detection time series in response to the liveness detection request. The liveness detection time series can be used to determine the result of liveness detection within the first time interval. The result of liveness detection is used to characterize whether the target object contained in the video frames of the liveness detection time series is the same object as the verification object stored locally or on the server. The user can pre-store the verification object when registering an account. Exemplarily, when user A registers an account for application B, the verification object is pre-stored. When user A logs in to the account through face recognition in application B and inputs a liveness recognition operation to the terminal, the terminal starts the first liveness recognition process in response to the liveness recognition operation. The time for the terminal to collect all video frames of user A during the first liveness recognition process is 10s. Then, the first time interval for the first liveness recognition process is 10s, and the liveness detection time series composed of all video frames can be used to determine the result of liveness detection of user A within the first time interval.

[0067] The terminal can also collect video frames at a preset first frequency (for example, 10 frames per second). It should be noted that the embodiments of the present application do not specifically limit the first time interval. The first time interval can be negatively correlated with the first frequency, and the first time interval can also be positively correlated with the number of collected video frames. The first frequency, the number of collected video frames, and the first time interval can all be set according to the actual situation. See Figure 7, a video frame may include a timestamp and an image (data - pic) acquired at the timestamp. Sorting multiple video frames according to the time relationship of the corresponding timestamps yields a live detection time series. The time interval between two adjacent video frames in the live detection time series is negatively correlated with a preset first frequency.

[0068] In some embodiments, refer to Figure 7 , the video frame may further include a request identifier (request_id), and its manifestation forms include but are not limited to various ways such as numbers and letters. The request identifier is the unique identifier information used to represent the live detection initiated by the terminal. The request identifier may include the identity identifier information of the target object. Exemplarily, the identity identifier information of the target object may be an ID number, an account number, etc. The request identifier may also include the device identifier of the terminal, the start time of the live detection, etc. Video frames carrying the same request identifier are acquired in the same live recognition.

[0069] Step S102, determine the operating environment parameters of the terminal within the first time interval.

[0070] Here, the operating environment parameters are the operating environment information affected when the terminal performs video frame acquisition, such as process information and file directory information, etc. During the first time interval when the terminal performs video frame acquisition, the operating environment inside the terminal system can be synchronously acquired at a preset second frequency to obtain multiple operating environment parameters. It should be noted that in the embodiments of the present application, the first frequency at which the video frames are acquired and the second frequency at which the operating environment parameters are acquired are not specifically limited. The first frequency and the second frequency may be the same or different. Exemplarily, the first frequency may be less than the second frequency.

[0071] In the embodiments of the present application, determining the operating environment parameters of the terminal within the first time interval can be achieved based on the following method: First, obtain the first file directory information and the first process information of the terminal at the first environment frame sampling moment, where the first environment frame sampling moment is any moment within the first time interval; then, determine the first file directory information and the first process information as the operating environment parameters of the terminal at the first environment frame sampling moment.

[0072] Here, within the first time interval, the file directory information and process information of the terminal can be obtained based on a preset second frequency at multiple environmental frame sampling moments. For any environmental frame sampling moment, the file directory information and process information of the terminal at this environmental frame sampling moment are determined as the operating environment parameters of the terminal at this environmental frame sampling moment. One of the environmental frame sampling moments is taken as the first environmental frame sampling moment, that is, the file directory information and process information of the terminal obtained at this environmental frame sampling moment are the first file directory information and the first process information. The first environmental frame sampling moment is any environmental frame sampling moment within the first time interval.

[0073] In the embodiments of this application, the living body recognition method can be applied to the server in the living body recognition system. The client of the living body detection application and the environmental protection process independent of the client are deployed on the terminal at the same time. The environmental protection process is a process for collecting the operating environment parameters of the terminal. The operating environment of the terminal can be collected at a preset second frequency within the first time interval by calling the environmental protection process, and the operating environment parameters corresponding to multiple timestamps are obtained. The timestamp is the environmental frame sampling moment when the operating environment parameters are collected. The environmental protection process can be written in the Python language and finally packaged into an exe executable file. After calling the environmental protection process, the environmental protection process can obtain all the directory files on the entire disk within the first time interval through the system-level entry os.listdir(), and the environmental protection process can obtain all the process information of the currently running processes through the system-level entry psutil. Exemplarily, all the process information can be obtained through psutil.process_iter(), such as process.pid() process number, process.name() process name, process.status() process status, process.create_time() process creation time, process.cpu_percent() process CPU occupancy rate, process.memory_info() process memory occupancy, etc. After obtaining multiple directory files at each environmental frame sampling moment, for each environmental frame sampling moment, the multiple directory files at this environmental frame sampling moment can be constructed into file directory information with a tree-like form through the join() method.

[0074] Exemplarily, Figure 8 is a schematic diagram of the architecture of a living body recognition system. As Figure 8As shown in the figure, the terminal is a PC terminal, and the live body recognition system includes a PC terminal and a server. Among them, the server is a system running on a remote server, providing application programming interface (API) services, including scheduling and detection services. The environment protection process can be an executable application running on the PC terminal, which is an independent Windows program, can be directly run in the Windows system, is a system-level application, is an application at the same level as the browser, and is an application that can be privileged by the user. The permissions of the environment protection process can be manually assigned by the administrator (ordinary user). Therefore, the permissions are higher than those of the web program inside the browser. When the administrator (ordinary user) does not assign high permissions to the environment protection process, the environment protection process cannot run normally, that is, it cannot collect the running environment parameters, so it cannot release the user's live body recognition request, that is, it is determined that the live body result is not passed.

[0075] The client is a live body main application (live body detection application) running inside the PC terminal browser, which is a web program, and its permissions are granted by the browser. The client is used to interact with the user to perform live body detection on the user (such as face recognition). The web program can only obtain the usage permissions of external devices such as cameras and microphones for interacting with the user, and cannot obtain complex running environment parameters. The environment protection process is independent of the live body main application.

[0076] See Figure 8 , and the live body recognition system may further include an injection detection application. The injection detection application is at the browser web page permission level, cannot obtain running environment parameters, and cannot effectively identify injection attacks on the PC terminal. The injection detection application can only obtain the usage permission of the camera and detect injection attacks by obtaining the camera name.

[0077] Therefore, for the live body recognition system provided in the embodiment of the present application, the injection attack risk can be detected through an independent environment protection process. When the user himself does not manually assign high permissions to the environment protection process, the environment protection process cannot obtain the running environment parameters of the PC terminal, and naturally, a live body recognition result of failed live body recognition will be obtained; when the user himself manually assigns high permissions to the environment protection process, the environment protection process will detect and obtain the running environment parameters and send them to the server. The server processes the running environment parameters and the video frames sent by the live body main application to obtain the final live body recognition result. That is, even if the PC terminal is under an injection attack, it cannot bypass the environment protection process to pass face recognition, and even less can it perform operations such as stealing the financial accounts of logged-in users, greatly improving the security.

[0078] Step S103, determine the environmental security detection time series of the terminal based on the running environment parameters of the terminal.

[0079] The environmental security detection time series includes environmental detection frames.

[0080] Here, the environmental security detection time series includes environmental detection frames at each environmental frame sampling moment. For any environmental frame sampling moment, the environmental detection frame at this environmental frame sampling moment can be used to characterize the security of the terminal's operating environment at this environmental frame sampling moment, that is, it can be used to characterize whether there is an injection attack risk for the terminal at this environmental frame sampling moment. Based on the operating environment parameters of the terminal at each environmental frame sampling moment, the environmental detection frame at each environmental frame sampling moment is determined. The environmental detection frames at all environmental frame sampling moments within the first time interval constitute the environmental security detection time series. The environmental security detection time series can be used to determine the environmental security result within the first time interval corresponding to the video frame acquisition, so as to determine whether there is an injection attack during the entire live body recognition process.

[0081] In some embodiments, refer to Figure 6B , Figure 6A The steps shown in S103 can be implemented through the following steps S1031 to S1032, which will be specifically described below.

[0082] Step S1031: Determine the environmental detection frame at the first environmental frame sampling moment according to the first file directory information and the first process information of the terminal at the first environmental frame sampling moment.

[0083] In the embodiments of the present application, the latest environmental detection policy can be obtained, and based on the environmental detection policy, the operating environment parameters at the first environmental frame sampling moment are processed to obtain the environmental detection frame at the first environmental frame sampling moment. The environmental detection policy may include a detection identifier, and the detection identifier is used to characterize the detection method for detecting whether there is injection attack information in the operating environment parameters. Exemplarily, if the obtained environmental detection policy includes a detection identifier A, and the detection identifier A is used to characterize the detection method of regular matching, at this time, the environmental detection policy may further include the regular expression "process.name (process name) = virtual camera (virtual camera)". Based on the regular expression "process.name = virtual camera", it can be detected whether "virtual cam era" exists in the first file directory information and the first process information of the terminal at the first environmental frame sampling moment. For example, if the first file directory information or the first process information contains "virtualcamera", it is considered that there is software hijacking the camera function in the system, and an environmental detection frame indicating that there is an injection attack risk in the operating environment of the terminal at this time is obtained.

[0084] It should be noted that semantic matching algorithms such as natural language processing (NLP) or text semantic encoding algorithms can also be used to replace the regular matching detection method.

[0085] In the embodiments of the present application, by obtaining the latest environment detection policy and processing the running environment parameters at each environment frame sampling moment based on the environment detection policy, an environment detection frame at each environment frame sampling moment is obtained, which facilitates detecting the security of the running environment according to the latest injection attack software and a more effective detection method, and improves the security and accuracy of liveness recognition.

[0086] In some embodiments, when obtaining the environment detection policy, a request identifier is obtained, and based on the request identifier, the requests of the client and the environment protection process are associated, and the liveness detection time series and the environment security detection time series having the same request identifier as the liveness detection time series are associated, which facilitates subsequent aggregation processing.

[0087] See Figure 9 , the environment detection frame may include a time stamp and an environment status (data-status) obtained by processing the running environment parameters collected at the time stamp. The time stamp in the environment detection frame is the environment frame sampling moment. The environment status may include two cases: "1" and "0". When the environment status is "1", it indicates that there may be an injection attack risk in the running environment of the terminal at the environment frame sampling moment. When the environment status is "0", it indicates that there is no injection attack risk in the running environment of the terminal at the environment frame sampling moment.

[0088] In some embodiments, see Figure 9 , the environment detection frame may further include a request identifier (request_id), and its manifestation forms include but are not limited to various ways such as numbers and letters. The request identifier is the unique identifier information used to represent the liveness detection initiated by the terminal. The environment detection frames carrying the same request identifier are the environment states collected for the recognition environment in the same liveness recognition initiated by the target object. That is, if two environment detection frames have the same request identifier but different time stamps, it indicates that these two environment detection frames are the environment states collected for the running environment of the terminal at different environment frame sampling moments in the same liveness recognition.

[0089] The embodiment of the present application separates the environment protection process from the main live application (client) and directly places the separated environment protection process in the operating system for execution. The environment protection process is implemented at the system level. Since the environment protection process can be manually elevated to "run as administrator" by the user, while the main live application in the browser cannot be elevated by the user, the operating environment parameters obtained by the environment protection process and the live detection time series obtained by the client jointly determine the live recognition result of the target object, thereby increasing the authority of the main live application in disguise, thereby realizing the identification and defense of high-authority injection attacks, thereby greatly improving the security and accuracy of live recognition.

[0090] In some embodiments, Figure 6B Step S1031 shown can be implemented by the following steps: Searching for target injection attack information in the first process information and the first file directory information. If there is target injection attack information, determining the first environment classification value as the environment detection frame at the first environment frame sampling time. If there is no target injection attack information, determining the second environment classification value as the environment detection frame at the first environment frame sampling time. The first environment classification value and the second environment classification value have different values.

[0091] In the present application examples, see Figure 10 , process information can include process ID (Identity document), process name, process status, process creation time, process central processing unit (CPU) usage, and process memory usage. Figure 11 The file directory information is a tree-structured file directory, including the complete structure of all directory files starting from the root node of the terminal.

[0092] In some embodiments, the latest environment detection strategy can be obtained in response to a liveness detection request sent by a terminal. If the detection identifier carried by the environment detection strategy is used for characterization, the detection method used to detect whether there is injection attack information in the environment parameters is: directly detect whether there is target injection attack information in the operating environment parameters; then first obtain the target injection attack information from the environment detection strategy. Among them, the target injection attack information can be the process name or file name of the hijacking software for the injection attack. Then, the target injection attack information is searched in the first process information and the first file directory information at the first environment frame sampling moment, and it is determined whether the target injection attack information exists in the operating environment parameters at the first environment frame sampling moment.

[0093] In the embodiments of the present application, the environmental state includes a first environmental classification value and a second environmental classification value, and the values of the first environmental classification value and the second environmental classification value are different. The first environmental classification value is used to represent that there may be a risk of injection attack in the operating environment of the terminal at the sampling moment of the first environmental frame. For the process information and file directory information at any environmental frame sampling moment, if there is target injection attack information in the process information at the sampling moment of the environmental frame, and / or there is target injection attack information in the file directory information, it means that there is a risk of injection attack in the operating environment of the terminal at the sampling moment of the environmental frame, and the first environmental classification value can be determined as the environmental detection frame at the sampling moment of the environmental frame. At the same time, the environmental detection frame can also store the sampling moment of the environmental frame as a time stamp.

[0094] Exemplarily, the first environmental classification value can be "1". Taking the terminal operating system as the windows system as an example, the next level of the root node is the "C\D\E\F disk drives", and the next level is such as "Program Files". If the injection attack method is to install a camera hijacking software (for example, the file name is WeChat.exe) in the windows system, the target injection attack information is "WeChat.exe". Search for the target injection attack information in the file directory information. If there is a directory of "D:\Program Files(x86)\Tencent\WeChat\WeChat.exe" in the file directory information, that is, there is target injection attack information in the file directory information, which means that the injection attacker has installed the camera hijacking software in the windows system. At this time, there is a risk of injection attack, and the first environmental classification value "1" is determined as the environmental detection frame.

[0095] In the embodiments of the present application, the second environmental classification value is used to represent that there is no risk of injection attack in the operating environment of the terminal at the sampling moment of the first environmental frame. If there is no target injection attack information in the process information at the sampling moment of the first environmental frame, and there is also no target injection attack information in the file directory information, it means that there is no risk of injection attack in the operating environment of the terminal at the sampling moment of the first environmental frame, and the second environmental classification value is determined as the environmental detection frame at the sampling moment of the first environmental frame. At the same time, the environmental detection frame can also store the sampling moment of the first environmental frame as a time stamp. Exemplarily, the second environmental classification value can be "0".

[0096] In the embodiments of the present application, by analyzing whether there is target injection attack information in the process information and the file directory information, it can be obtained whether there is a risk software or process with the function of hijacking the camera in the operating environment of the terminal. If there is no risk software or process, the operating environment is considered safe. If there is a risk software or process, the operating environment is considered dangerous, which improves the security and accuracy of liveness recognition.

[0097] In some embodiments, referring to Figure 6C , Figure 6B the steps S1031 shown can be implemented by the following steps S10311 to S10313, which are specifically described below.

[0098] Step S10311, encode the first process information and the first file directory information to obtain an environmental vector.

[0099] In some embodiments, in response to a liveness detection request sent by a terminal, the latest environmental detection policy can be obtained. If the detection identifier carried by the environmental detection policy indicates that the detection method for whether there is injection attack information in the detected environmental parameters is to use a text semantic encoding algorithm, referring to Figure 12 , and both the process information and the file directory information are text information, the pre-trained semantic model can be first used to perform semantic extraction on the first process information and the first file directory information at the first environmental frame sampling moment, and encode to obtain the environmental vector at the first environmental frame sampling moment. It should be noted that the pre-trained semantic model in the embodiments of the present application is not specifically limited. For example, it can be a BERT-BASE model, or other text semantic encoding algorithms with different structures from the BERT-BASE model but capable of achieving the same purpose. Taking the BERT-BASE model as an example, the obtained environmental vector is a 768-bit vector.

[0100] Step S10312, classify the environmental vector to obtain the classification result of the environmental vector.

[0101] In the embodiments of the present application, the pre-trained semantic model can be used to classify the environmental vector at each environmental frame sampling moment to obtain the classification result of the environmental vector at each first environmental frame sampling moment. The classification result can represent the security of the running environment of the terminal, also known as the environmental state. The classification result can include a first type of classification result and a second type of classification result. The first type of classification result can be represented by the number "1". When the classification result is the first type of classification result "1", it indicates that there may be a risk of injection attack in the running environment at this environmental frame sampling moment. The second type of classification result can be represented by the number "0". When the classification result is the second type of classification result "0", it indicates that there is no risk of injection attack in the running environment at this environmental frame sampling moment.

[0102] Step S10313, determine the environmental detection frame at the first environmental frame sampling moment according to the classification result of the environmental vector.

[0103] In the embodiments of the present application, after obtaining the classification result at the first environmental frame sampling moment, the classification result at the first environmental frame sampling moment can be determined as the environmental detection frame at the first environmental frame sampling moment. At the same time, the environmental detection frame can also store the first environmental frame sampling moment as a timestamp.

[0104] In the embodiments of the present application, by encoding and classifying process information and file directory information, the security of the running environment is determined based on the obtained classification results, improving the efficiency and effectiveness of environment security detection, and further enhancing the security and accuracy of liveness recognition.

[0105] Step S1032: Determine the environmental security detection time series of the terminal according to the multiple environmental detection frames.

[0106] Here, the multiple environmental detection frames can be sorted according to the time relationship of the corresponding environmental frame sampling moments to obtain the environmental security detection time series of the terminal.

[0107] Step S104: Aggregate the video frames of the liveness detection time series and the environmental detection frames of the environmental security detection time series to obtain the liveness recognition result of the target object.

[0108] In the embodiments of the present application, referring to Figure 13 , after aggregating the liveness detection time series and the environmental security detection time series, a liveness recognition matrix can be obtained, and the liveness recognition result of the target object can be obtained after processing the liveness recognition matrix through a liveness recognition model.

[0109] In some embodiments, referring to Figure 6D , Figure 6A The steps shown in step S104 can be implemented through the following steps S1041 to S1044, which are specifically described below.

[0110] Step S1041: Determine the liveness detection score of each video frame in the liveness detection time series.

[0111] In the embodiments of the present application, referring to Figure 14 , obtain the image in each video frame of the liveness detection time series. A silent liveness detector can be used to determine the liveness detection score corresponding to the image in each video frame. It should be noted that the implementation method of the silent liveness detector in the embodiments of the present application is not limited, and it is similar to the face recognition method in the related art. Exemplarily, based on the image in the video frame and the pre-uploaded verification image, the similarity between the image in the video frame and the verification image can be determined as the liveness detection score corresponding to the image in the video frame. The liveness detection score still carries the corresponding timestamp information, that is, the video frame acquisition moment information.

[0112] Step S1042: Aggregate the multiple liveness detection scores and the multiple environmental detection frames to obtain a liveness recognition matrix.

[0113] In the embodiments of the present application, the mapping relationship between the multiple liveness detection scores and the multiple environmental detection frames can be determined, and a liveness recognition matrix can be constructed based on the mapping relationship.

[0114] In some embodiments, referring to Figure 6E , Figure 6D the steps shown in S1042 can be implemented by the following steps S10421 to S10424, which are specifically described below.

[0115] Step S10421, determine the first duration of each live detection score.

[0116] The first duration is the time difference between the first video frame acquisition time and the second video frame acquisition time. The first video frame acquisition time is the acquisition time of the first video frame of the live detection score, and the second video frame acquisition time is the acquisition time of the next video frame adjacent to the first video frame.

[0117] In the embodiments of the present application, referring to Figure 14 , for any live detection score, the time difference between the first video frame acquisition time of the first video frame corresponding to the live detection score and the second video frame acquisition time of the next video frame adjacent to the first video frame in the live detection time series can be calculated as the first duration of the live detection score.

[0118] Step S10422, determine the second duration of each environmental detection frame.

[0119] The second duration is the time difference between the second environmental frame sampling time and the third environmental frame sampling time. The second environmental frame sampling time is the sampling time of the first environmental detection frame, and the third environmental frame sampling time is the sampling time of the next environmental detection frame adjacent to the first environmental detection frame.

[0120] In the embodiments of the present application, referring to Figure 15 , for any current environmental detection frame, the time difference between the second environmental frame sampling time of the environmental detection frame and the third environmental frame sampling time of the next environmental detection frame adjacent to the environmental detection frame in the environmental safety detection time series can be calculated as the second duration of the environmental detection frame.

[0121] Step S10423, determine the mapping relationship between each live detection score and at least one environmental detection frame according to the first duration and the second duration.

[0122] The first duration of each live detection score is greater than or equal to the sum of the second durations of at least one environmental detection frame having a mapping relationship with the live detection score.

[0123] In the embodiments of the present application, the first frequency of video frame acquisition is less than the second frequency of environmental detection frame acquisition. The mapping relationship between each live detection score and at least one environmental detection frame can be determined based on the first duration of each live detection score and the second duration of each environmental detection frame. For any live detection score, the first duration of the live detection score is greater than or equal to the sum of the second durations of the environmental detection frames having a mapping relationship with the live detection score, and the video frame acquisition time corresponding to the live detection score is before or the same as the environmental frame sampling time corresponding to each environmental detection frame, and the video frame acquisition time of the next video frame adjacent to the live detection score is after or the same as the environmental frame sampling time corresponding to each environmental detection frame.

[0124] Exemplarily, the video frame acquisition time of the live detection score a is t0, the video frame acquisition time of the next live detection score b adjacent to the live detection score a in the live detection time series is t0 + 10 ms, and the first duration of the live detection score a is 10 ms; the environmental frame sampling time of the environmental detection frame 1 is t0, the environmental frame sampling time of the next environmental detection frame 2 adjacent to the environmental detection frame 1 in the environmental safety detection time series is t0 + 3 ms, and the second duration of the environmental detection frame 1 is 3 ms, and so on. If the environmental frame sampling time of the environmental detection frame 3 is t0 + 6 ms, the environmental frame sampling time of the environmental detection frame 4 is t0 + 9 ms, and the environmental frame sampling time of the environmental detection frame 5 is t0 + 12 ms. Then the sum of the second durations of the environmental detection frames 1, 2, and 3 is 9 ms, which is less than the first duration of the live detection score a of 10 ms, and the video frame acquisition time t0 of the live detection score a is before the environmental frame sampling times corresponding to the environmental detection frames 2 and 3, the video frame acquisition time t0 is the same as the environmental frame sampling time corresponding to the environmental detection frame 1, and the video frame acquisition time t0 + 10 ms of the live detection score b is after the environmental frame sampling times corresponding to the environmental detection frames 1, 2, and 3. The environmental detection frames 1, 2, and 3 are determined to be the environmental detection frames having a mapping relationship with the live detection score a.

[0125] Step S10423, based on the mapping relationship, construct a live recognition matrix with multiple live detection scores and multiple environmental detection frames.

[0126] Here, after determining the mapping relationship between each live detection score and at least one environmental detection frame, for any live detection score, the number of environmental detection frames having a mapping relationship with this live detection score can be obtained. The live detection score is replicated and increased to obtain the first number of this live detection score. The first number is the same as the number of the above-mentioned environmental detection frames having a mapping relationship. In this way, it can be achieved that each live detection score corresponds to an environmental detection frame. All the replicated and increased live detection scores are used as the elements of the first row of the live recognition matrix, and the environmental detection frames corresponding to each live detection score are filled into the corresponding positions of the second row of the live recognition matrix as the elements of the second row. The live detection scores and environmental detection frames in the same column of the live recognition matrix have a mapping relationship.

[0127] Alternatively, based on the number of live detection scores, the number of environmental detection frames, and the mapping relationship within the first time interval, a plurality of live detection scores and a plurality of environmental detection frames collected within the first time interval can be constructed into a live recognition matrix. First, determine the least common multiple of the number of live detection scores and the number of environmental detection frames. Then, replicate and increase the number of live detection scores and the number of environmental detection frames to the least common multiple respectively, and a live recognition matrix can be constructed based on the obtained least common multiple of live detection scores and environmental detection frames, as well as the mapping relationship.

[0128] Exemplarily, 20 live detection scores and 30 environmental detection frames are collected within the first time interval, then the least common multiple is 60. For a live detection score, this live detection score is replicated and increased to 3 live detection scores; for an environmental detection frame, this environmental detection frame is replicated and increased to 2 environmental detection frames. Finally, 60 live detection scores and 60 environmental detection frames are obtained. A live recognition matrix can be constructed based on the mapping relationship with 60 live detection scores and 60 environmental detection frames.

[0129] In the embodiment of the present application, by calculating the first duration of the live detection score and the second duration of the environmental detection frame, the alignment of the mapping relationship between the live detection score and the environmental detection frame can be realized through the first duration and the second duration, constructing a live recognition matrix, which is convenient for subsequent model processing to obtain the total live recognition score and improves the processing efficiency of live recognition.

[0130] Step S1043, input the live recognition matrix into a preset live recognition model to obtain the total live recognition score.

[0131] In the embodiment of the present application, refer to Figure 16, after obtaining the live body recognition matrix, the live body recognition matrix can be input into a preset live body recognition model. The preset live body recognition model may include a convolutional neural network (CNN). During the processing of the live body recognition model, the convolutional neural network performs a series of operations such as convolution and pooling on the live body recognition matrix, and then uses a normalization function (softmax function) to normalize the output of the convolutional neural network to obtain the total live body recognition score. The total live body recognition score is within the range of 0-1. It should be noted that the structure and internal operation process of the convolutional neural network are not specifically limited in the embodiments of the present application.

[0132] Exemplarily, the convolutional neural network can be a small-scale convolutional neural network resnet18.

[0133] Step S1044, when the total live body recognition score is greater than a preset score threshold, determine that the live body recognition result of the target object is that the live body recognition passes.

[0134] In the embodiments of the present application, the preset score threshold is not specifically limited and can be set according to actual situations, as long as it satisfies the range of 0-1. When the total live body recognition score is greater than the preset score threshold, it indicates that both the live body detection and the environmental security detection pass, and it is determined that the live body recognition result of the target object is that the live body recognition passes. Or, when the total live body recognition score is less than the preset score threshold, it indicates that one or both of the live body detection and the environmental security detection fail, and it is determined that the live body recognition result of the target object is that the live body recognition is rejected. Send and display the final live body recognition result to the terminal.

[0135] The embodiments of the present application perform aggregation processing on the video frames in the live body detection time series and the environmental detection frames in the environmental security detection time series to obtain the live body recognition result of the target object, and can comprehensively judge whether the target object is a live body behavior through both live body detection and environmental security detection, realizing the defense against injection attacks in the live body recognition process on the computer side and improving the accuracy of live body recognition.

[0136] In some embodiments, before executing Figure 6A the steps shown in S101, the following steps may also be executed: Receive a request identifier sent by the terminal. If the request identifier is successfully obtained, return a synchronization success result to the terminal, start the live body detection process, and the terminal performs video frame acquisition; if the request identifier is not successfully obtained, terminate the live body detection process and return a synchronization failure result to the terminal.

[0137] Next, an exemplary application of the embodiments of the present application in an actual application scenario will be described.

[0138] The solution of the embodiments of the present application can be applied to the "financial identity verification" commonly used in the financial industry, that is, the liveness detection method of the face category. It is usually used before applying for credit as the first step in regulatory identity verification requirements, and can effectively reduce some common attack risks such as photocopying, recording, and injection. The embodiments of the present application solve the problems of poor user experience in related technologies such as "direct strong light" and "long process". The embodiments of the present application also solve the poor security problems of "inability to completely defend against injection attacks" and "inability to detect injection attacks on the PC due to low web application permissions" in related technologies. The liveness recognition method provided in the embodiments of the present application can be applied to the following: Figure 17 In the liveness recognition system shown. The liveness recognition system includes an environment assurance process, a liveness main application and a server. The environment assurance process includes a process scanning module for process acquisition and process analysis. The environment assurance process also includes a file scanning module for file directory acquisition and file directory analysis. The liveness main application includes a real-time transmission acquisition module, a silent liveness recognition module and an injection detection module. The real-time transmission acquisition module is used for front-end acquisition and front-end compression (video frame acquisition and compression). The silent liveness recognition module is used for front-end facial preprocessing. The injection detection module is used for camera detection. The server is used for real-time acquisition, quality inspection, anti-copying, anti-mask, printing and collaborative real-time sharing algorithm (aggregate operation). The specific execution process of the liveness recognition method is as follows Figure 18 As shown: the user initiates liveness recognition on the PC, and the liveness main application and the environment assurance process collect information during the liveness period and submit it to the server, and finally the server makes a comprehensive judgment on the liveness recognition result. The user initiates a liveness recognition request in the liveness main application; the liveness main application first communicates with the environment assurance process and synchronizes the request identifier with the environment assurance process; the environment assurance process pulls the latest environment detection strategy from the server and uploads the request identifier to the server; after the synchronization of the request identifier is completed, the liveness main application will prompt the user to start liveness detection, please look at the camera; if the synchronization request fails, the user will be prompted to start the environment assurance process. When the liveness detection starts, the liveness main application will transmit the video frames obtained from the real-time video acquisition to the server at a rate of 10 frames per second in real time. The environment assurance process will obtain the current process information and directory information, analyze the process information and directory information according to the current environment detection strategy, and transmit the latest analysis results (environmental status) to the server in real time; after 500ms of basic continuous acquisition and recognition, the server makes a final result judgment through a collaborative real-time sharing algorithm (aggregate operation), notifies the liveness main application of the current liveness recognition result, and displays it to the user. As shown Figure 19As shown, a simple liveness recognition interactive interface displays an echo box. Specifically, the above-mentioned environment assurance process can be an analysis engine, which can be composed of a semantic extraction encoder and a classifier, and semantically extracts and classifies the directory information and process information in text form, and the classification result is the current environment security of 0 and 1 values. The server will link the environment security detection sequence value returned by the environment assurance process and the sequence value of the silent liveness detection uploaded by the liveness main application according to the request identifier, perform aggregation operations, and give the final liveness recognition result, and then choose to pass or reject the liveness recognition request of the user. The environment assurance process can report the detection result in real time based on the environment detection strategy. If the environment detection strategy is "pattern = r'wechat\.exe'", the environment assurance process will search for the pattern "pattern = r'wechat\.exe'" in the process name process.name in the above-mentioned process information and the directory information. If it is hit, it will return 1 (dangerous), and if it is not hit, it will return 0 (safe). If the environment detection strategy can also be a model file for the detection network, it represents the latest environment detection strategy.

[0139] It should be noted that the silent liveness algorithm used by the liveness main application in the embodiment of the present application can be replaced by other algorithms with the same purpose but heterogeneous models, such as motion liveness, lip reading liveness, colorful liveness, etc.

[0140] The embodiment of the present application is completely based on the reverse exclusion mechanism, rather than the challenge response mechanism. By detecting common abnormal behaviors, the attack behavior can be identified, and then the liveness determination can be achieved. Compared with the challenge response mechanism, such a mechanism is not perceived by the user and does not require complex coordination actions. The total duration of a one-way liveness detection is less than 1000ms, which has an excellent user experience. The embodiment of the present application makes full use of the high-privilege characteristics of desktop applications. Through the collaborative real-time sharing algorithm mechanism, the original low-privilege live main application is disguised as a high-privilege application, thereby realizing the identification of high-privilege injection attack applications. Thereby greatly improving the security of liveness recognition. From the user's perspective, the embodiment of the present application only needs to look directly at the camera, which greatly reduces the complexity of user operations and improves the user experience; but from the engine's perspective, the user is in an environment where injection attacks cannot be performed, that is, the device has the ability to defend against "injection attacks".

[0141] It is understandable that in the embodiments of the present application, related data such as user information is involved. When the embodiments of the present application are applied to specific products or technologies, user permission or consent is required, and the collection, use and processing of relevant data need to comply with relevant laws, regulations and standards of relevant countries and regions.

[0142] The following is a description of an exemplary structure of a live body recognition device 455 provided in an embodiment of the present application implemented as a software module. In some embodiments,Figure 5 As shown, the software module stored in the living body recognition device 455 of the memory 450 may include: a request receiving module 4551 for receiving a living body detection request sent by the terminal; the living body detection request includes a living body detection time series, and the living body detection time series includes: video frames of the target object collected in real time by the terminal within a first time interval, and the first time interval is the time interval for collecting the video frames. An environmental parameter determination module 4552 for determining the operating environmental parameters of the terminal within the first time interval. An environmental sequence determination module 4553 for determining the environmental security detection time series of the terminal based on the operating environmental parameters of the terminal, and the environmental security detection time series includes environmental detection frames. An aggregation processing module 4554 for performing aggregation processing on the video frames of the living body detection time series and the environmental detection frames of the environmental security detection time series to obtain the living body recognition result of the target object.

[0143] In some embodiments, the environmental parameter determination module 4552 is further configured to obtain the first file directory information and the first process information of the terminal at the first environmental frame sampling moment, and the first environmental frame sampling moment is any moment within the first time interval; determine the first file directory information and the first process information as the operating environmental parameters of the terminal at the first environmental frame sampling moment.

[0144] In some embodiments, the environmental sequence determination module 4553 is further configured to determine the environmental detection frame at the first environmental frame sampling moment according to the first file directory information and the first process information of the terminal at the first environmental frame sampling moment; determine the environmental security detection time series of the terminal according to multiple environmental detection frames.

[0145] In some embodiments, the environmental sequence determination module 4553 is further configured to search for target injection attack information in the first process information and the first file directory information; if there is target injection attack information, determine the first environmental classification value as the environmental detection frame at the first environmental frame sampling moment; if there is no target injection attack information, determine the second environmental classification value as the environmental detection frame at the first environmental frame sampling moment; wherein, the first environmental classification value and the second environmental classification value have different values.

[0146] In some embodiments, the environmental sequence determination module 4553 is further configured to encode the first process information and the first file directory information to obtain an environmental vector; classify the environmental vector to obtain the classification result of the environmental vector; determine the environmental detection frame at the first environmental frame sampling moment according to the classification result of the environmental vector.

[0147] In some embodiments, the aggregation processing module 4554 is further configured to determine the live detection score of each video frame in the live detection time series; perform aggregation processing on the multiple live detection scores and the multiple environmental detection frames to obtain a live recognition matrix; input the live recognition matrix into a preset live recognition model to obtain a total live recognition score; and determine that the live recognition result of the target object is a pass when the total live recognition score is greater than a preset score threshold.

[0148] In some embodiments, the aggregation processing module 4553 is further configured to determine the first duration of each live detection score; the first duration is the time difference between the acquisition time of the first video frame and the acquisition time of the second video frame, the acquisition time of the first video frame is the acquisition time of the first video frame of the live detection score, and the acquisition time of the second video frame is the acquisition time of the next video frame adjacent to the first video frame; determine the second duration of each environmental detection frame; the second duration is the time difference between the sampling time of the second environmental frame and the sampling time of the third environmental frame, the sampling time of the second environmental frame is the sampling time of the first environmental detection frame, and the sampling time of the third environmental frame is the sampling time of the next environmental detection frame adjacent to the first environmental detection frame; determine the mapping relationship between each live detection score and at least one environmental detection frame according to each first duration and each second duration; the first duration of each live detection score is greater than or equal to the sum of the second durations of at least one environmental detection frame having a mapping relationship with the live detection score; and construct a live recognition matrix from the multiple live detection scores and the multiple environmental detection frames based on the mapping relationship.

[0149] An embodiment of the present application provides a computer program product, which includes a computer program or computer-executable instructions, and the computer program or computer-executable instructions are stored in a computer-readable storage medium. The processor of the electronic device reads the computer-executable instructions from the computer-readable storage medium, and the processor executes the computer-executable instructions, so that the electronic device executes the live recognition method described above in the embodiments of the present application.

[0150] An embodiment of the present application provides a computer-readable storage medium storing computer-executable instructions, in which computer-executable instructions or a computer program are stored. When the computer-executable instructions or the computer program are executed by a processor, the processor will be caused to execute the live recognition method provided in the embodiments of the present application. For example, Figure 6A the live recognition method shown.

[0151] In some embodiments, the computer-readable storage medium may be a memory such as RAM, ROM, flash memory, magnetic surface memory, optical disc, or CD-ROM; or may be various devices including one or any combination of the above memories.

[0152] In some embodiments, the computer-executable instructions may be in the form of a program, software, software module, script, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including being deployed as a stand-alone program or being deployed as a module, component, subroutine, or other unit suitable for use in a computing environment.

[0153] As an example, the computer-executable instructions may or may not correspond to a file in a file system, may be stored as part of a file that holds other programs or data, for example, in one or more scripts in a Hyper Text Markup Language (HTML) document, stored in a single file dedicated to the program being discussed, or, stored in multiple cooperating files (such as files that store one or more modules, subroutines, or portions of code).

[0154] As an example, the computer-executable instructions may be deployed to execute on one electronic device, or on multiple electronic devices located at one location, or, on multiple electronic devices distributed at multiple locations and interconnected by a communication network.

[0155] In summary, the embodiments of the present application disclose a silent liveness recognition method for dual-application cooperation, which collects a portrait through a liveness main application, detects the risk of injection attacks through an environmental guarantee process, and performs dual-application coupling binding through real-time communication. Based on the results of the dual modules, it comprehensively determines whether the user's behavior is a liveness behavior, improving the accuracy of liveness recognition.

[0156] The above is only the embodiments of the present application and is not intended to limit the protection scope of the present application. Any modifications, equivalent replacements, and improvements made within the spirit and scope of the present application are included in the protection scope of the present application.

Claims

1. A method for live body recognition, characterized in that, The method includes: Receiving a liveness detection request sent by a terminal; the liveness detection request includes a liveness detection time series, and the liveness detection time series includes video frames of a target object collected in real time by the terminal within a first time interval, where the first time interval is the time interval for collecting the video frames; Determining the operating environment parameters of the terminal within the first time interval; Determining an environmental security detection time series of the terminal based on the operating environment parameters of the terminal, where the environmental security detection time series includes environmental detection frames; Performing an aggregation process on the video frames of the liveness detection time series and the environmental detection frames of the environmental security detection time series to obtain a liveness recognition result of the target object.

2. The method according to claim 1, wherein The determining the operating environment parameters of the terminal within the first time interval includes: Obtaining first file directory information and first process information of the terminal at a first environmental frame sampling moment, where the first environmental frame sampling moment is any moment within the first time interval; Determining the first file directory information and the first process information as the operating environment parameters of the terminal at the first environmental frame sampling moment.

3. The method according to claim 2, wherein Determining the environmental security detection time series of the terminal based on the operating environment parameters of the terminal includes: Determining an environmental detection frame at the first environmental frame sampling moment according to the first file directory information and the first process information of the terminal at the first environmental frame sampling moment; Determining the environmental security detection time series of the terminal according to multiple environmental detection frames.

4. The method according to claim 3, wherein, The determining the environmental detection frame at the first environmental sampling moment according to the first file directory information and the first process information of the terminal at the first environmental sampling moment includes: Searching for target injection attack information in the first process information and the first file directory information; If the target injection attack information exists, determining a first environmental classification value as the environmental detection frame at the first environmental frame sampling moment; If the target injection attack information does not exist, determining a second environmental classification value as the environmental detection frame at the first environmental frame sampling moment; Wherein, the values of the first environmental classification value and the second environmental classification value are different.

5. The method according to claim 3, wherein The determining the environmental detection frame at the first environmental frame sampling moment according to the first file directory information and the first process information of the terminal at the first environmental frame sampling moment includes: Encoding the first process information and the first file directory information to obtain an environmental vector; Classifying the environmental vector to obtain a classification result of the environmental vector; Determining the environmental detection frame at the first environmental frame sampling moment according to the classification result of the environmental vector.

6. The method according to claim 1, wherein The performing an aggregation process on the video frames of the liveness detection time series and the environmental detection frames of the environmental security detection time series to obtain a liveness recognition result of the target object includes: Determining a liveness detection score for each video frame of the liveness detection time series; Performing an aggregation process on multiple liveness detection scores and multiple environmental detection frames to obtain a liveness recognition matrix; Input the living body recognition matrix into a preset living body recognition model to obtain the total living body recognition score; When the total living body recognition score is greater than a preset score threshold, determine that the living body recognition result of the target object is a passing living body recognition.

7. The method according to claim 6, characterized in that The aggregating the multiple living body detection scores and the multiple environment detection frames to obtain a living body recognition matrix includes: Determine a first duration for each of the living body detection scores; the first duration is the time difference between a first video frame acquisition time and a second video frame acquisition time, the first video frame acquisition time is the acquisition time of the first video frame of the living body detection score, and the second video frame acquisition time is the acquisition time of the next video frame adjacent to the first video frame; Determine a second duration for each environment detection frame; the second duration is the time difference between a second environment frame sampling time and a third environment frame sampling time, the second environment frame sampling time is the sampling time of the first environment detection frame, and the third environment frame sampling time is the sampling time of the next environment detection frame adjacent to the first environment detection frame; According to each of the first durations and each of the second durations, determine a mapping relationship between each of the living body detection scores and at least one environment detection frame; the first duration of each living body detection score is greater than or equal to the sum of the second durations of at least one environment detection frame having a mapping relationship with the living body detection score; Based on the mapping relationship, construct the multiple living body detection scores and the multiple environment detection frames into the living body recognition matrix.

8. A living body recognition device, characterized in that, The apparatus includes: A request receiving module, configured to receive a living body detection request sent by a terminal; the living body detection request includes a living body detection time series, and the living body detection time series includes: video frames of a target object collected in real time by the terminal within a first time interval, and the first time interval is a time interval for collecting video frames; An environment parameter determining module, configured to determine the operating environment parameters of the terminal within the first time interval; An environment sequence determining module, configured to determine an environment security detection time series of the terminal based on the operating environment parameters of the terminal, and the environment security detection time series includes environment detection frames; An aggregation processing module, configured to perform aggregation processing on the video frames of the living body detection time series and the environment detection frames of the environment security detection time series to obtain the living body recognition result of the target object.

9. An electronic device, characterized in that, The electronic device includes: A memory, configured to store computer-executable instructions; A processor, configured to implement the living body recognition method according to any one of claims 1 to 7 when executing the computer-executable instructions stored in the memory.

10. A computer-readable storage medium storing computer-executable instructions or a computer program, characterized in that, The computer-executable instructions or the computer program, when executed by the processor, implement the living body recognition method according to any one of claims 1 to 7.