Living body detection method and device, electronic equipment, storage medium and program product

By performing signal processing on the area occupied by the object in the image frame sequence, determining the peak and valley values, and selecting key image frames for liveness detection, the problems of flexibility and low efficiency in traditional methods are solved, and more efficient liveness detection is achieved.

CN120689941APending Publication Date: 2025-09-23MASHANG CONSUMER FINANCE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510377433.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-27
Publication Date
2025-09-23

AI Technical Summary

Technical Problem

Traditional liveness detection methods are vulnerable to forgery attacks when facing contactless biometric recognition, have low flexibility and efficiency, and have strict requirements on user behavior.

Method used

By performing signal processing on the area occupied by the object in the image frame sequence and determining the peak and valley values, the first, second and third image frames can be flexibly selected for liveness detection, reducing constraints on user behavior.

Benefits of technology

It improves the flexibility and efficiency of liveness detection, can more accurately capture image frames of objects at different distances, reduces user behavior constraints, and improves detection accuracy and efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120689941A_ABST
    Figure CN120689941A_ABST
Patent Text Reader

Abstract

The invention provides a living body detection method and device, electronic equipment, a storage medium and a program product. The method comprises the following steps: for each image frame in an image frame sequence, determining an area occupied by an object in the image frame; performing signal processing on the area occupied by the object to obtain a data signal, and determining a peak value and a valley value of the data signal; determining a first numerical value between the peak value and the valley value, and determining a first image frame corresponding to the peak value, a second image frame corresponding to the valley value and a third image frame corresponding to the first numerical value from the image frame sequence; and performing living body detection on the object based on the first image frame, the second image frame and the third image frame. According to the invention, the flexibility and efficiency of living body detection can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to intelligent recognition technology, and in particular to a liveness detection method, device, electronic device, storage medium and program product. Background Art

[0002] In biometric identification and security systems, liveness detection is a crucial step in ensuring the accuracy of identity verification. Traditional image processing methods are vulnerable to forgery attacks when it comes to contactless biometric recognition, such as using photos, videos, or 3D masks to spoof.

[0003] In related art, in order to improve the security of liveness detection, detection is performed by instructing the user to move the face acquisition device or the face of the object to be detected from far to near or from near to far, but this method has low flexibility and efficiency. Summary of the Invention

[0004] The embodiments of the present application provide a liveness detection method, apparatus, electronic device, storage medium, and program product, which can improve the flexibility and efficiency of liveness detection.

[0005] The technical solution of the embodiment of the present application is implemented as follows:

[0006] The present invention provides a method for detecting a living body, which includes:

[0007] For each image frame in the image frame sequence, determining an area occupied by an object in the image frame;

[0008] performing signal processing on the area occupied by the object to obtain a data signal, and determining peak values ​​and valley values ​​of the data signal;

[0009] Determine a first value between the peak value and the valley value, and determine a first image frame corresponding to the peak value, a second image frame corresponding to the valley value, and a third image frame corresponding to the first value from the image frame sequence;

[0010] Perform livingness detection on the object based on the first image frame, the second image frame, and the third image frame.

[0011] The present invention provides a living body detection device, which includes:

[0012] an area determination module, configured to determine, for each image frame in the image frame sequence, an area occupied by an object in the image frame;

[0013] a signal processing module, configured to perform signal processing on the area occupied by the object to obtain a data signal, and determine peak values ​​and valley values ​​of the data signal;

[0014] a value determination module, configured to determine a first value between the peak value and the valley value, and determine, from the image frame sequence, a first image frame corresponding to the peak value, a second image frame corresponding to the valley value, and a third image frame corresponding to the first value;

[0015] A living body detection module is used to perform living body detection on the object based on the first image frame, the second image frame and the third image frame.

[0016] An embodiment of the present application provides an electronic device, comprising:

[0017] a memory for storing computer-executable instructions or computer programs;

[0018] The processor is configured to implement the liveness detection method provided in the embodiment of the present application when executing the computer executable instructions or computer program stored in the memory.

[0019] An embodiment of the present application provides a computer-readable storage medium storing a computer program or computer-executable instructions for implementing the liveness detection method provided in the embodiment of the present application when executed by a processor.

[0020] An embodiment of the present application provides a computer program product, including a computer program or computer-executable instructions. When the computer program or computer-executable instructions are executed by a processor, the liveness detection method provided in the embodiment of the present application is implemented.

[0021] The embodiments of the present application have the following beneficial effects:

[0022] During the liveness detection process, a data signal is obtained by performing signal processing on the area occupied by the object in multiple image frames, and the peak and valley values ​​of the data signal are determined. Based on the first image frame corresponding to the peak value, the second image frame corresponding to the valley value, and the third image frame corresponding to the first value between the peak and valley values, liveness detection of the object is achieved. Compared with the method in the related art that strictly requires the user to move at a constant speed from far to near or from near to far to collect image frames, and relies on timestamps to select intermediate image frames, the embodiment of the present application analyzes the data signal obtained by area changes and flexibly selects the first image frame, the second image frame, and the third image frame. It can more accurately capture image frames of the object at different distances, reduce the constraints on the user's behavior, and improve the flexibility of liveness detection. At the same time, the embodiment of the present application can accurately identify the most representative image frames in a shorter time, thereby improving the efficiency of liveness detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] Figure 1 Schematic diagram of key frame selection in related art;

[0024] Figure 2Schematic diagram of the architecture of the liveness detection system provided in an embodiment of the present application;

[0025] Figure 3 is a structural diagram of an electronic device provided in an embodiment of the present application;

[0026] Figure 4 This is a flow chart of the liveness detection method provided in the embodiment of the present application. Figure 1 ;

[0027] Figure 5 This is a flow chart of the liveness detection method provided in the embodiment of the present application. Figure 2 ;

[0028] Figure 6 This is a flow chart of the liveness detection method provided in the embodiment of the present application. Figure 3 ;

[0029] Figure 7 This is a block diagram of a liveness detection method provided in an embodiment of the present application;

[0030] Figure 8 This is a schematic diagram of extracting each key frame provided in an embodiment of the present application. DETAILED DESCRIPTION

[0031] In order to make the purpose, technical solutions and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limiting this application. All other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.

[0032] In the following description, reference is made to “some embodiments”, which describes a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.

[0033] In the following description, the terms "first\second\third" involved are merely used to distinguish similar objects and do not represent a specific ordering of the objects. It can be understood that "first\second\third" can be interchanged with a specific order or sequence where permitted, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.

[0034] In the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program that has a predetermined function and works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as processing circuits or memories) or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be part of an overall module or unit that includes the function of the module or unit.

[0035] Unless otherwise defined, all technical and scientific terms used in the embodiments of the present application have the same meanings as those commonly understood by those skilled in the art. The terms used in the embodiments of the present application are only for the purpose of describing the embodiments of the present application and are not intended to limit the present application.

[0036] The relevant data collection and processing in the embodiments of this application should be strictly in accordance with the requirements of relevant laws and regulations when applied in examples, and the informed consent or separate consent of the personal information subject should be obtained. Subsequent data use and processing should be carried out within the scope of authorization of laws and regulations and the personal information subject.

[0037] Before further describing the embodiments of the present application in detail, the nouns and terms involved in the embodiments of the present application are explained. The nouns and terms involved in the embodiments of the present application are subject to the following interpretations.

[0038] 1) Liveness detection: It is a method of determining the true physiological characteristics of an object in an identity verification scenario. It refers to the use of technical means to determine whether a biometric feature (such as a face, fingerprint, etc.) comes from a real living person, rather than from forged or deceptive means (such as photos, video playback, etc.).

[0039] 2) Image frame sequence: This refers to a collection of continuously captured image frames. Each image frame represents a scene captured at a specific timestamp. In liveness detection, an image frame sequence includes image frames captured at different timestamps of the subject.

[0040] 3) Data Signal: This is a numerical representation obtained by processing the information in the image frames. In the embodiment of the present application, the data signal is a sequence of signal values ​​extracted from the changes in the occupied area of ​​the object in multiple image frames, which is used for subsequent peak and valley value analysis.

[0041] 4) Peak: refers to the maximum value reached by a data signal within a local range (a limited interval or specific time period of the data signal).

[0042] 5) Valley value: refers to the minimum value reached by the data signal within a local range.

[0043] 6) Key points: refers to points with significant features in the object contained in the image frame. Taking the object as an example, the key points can be the positions of facial features such as eyes, nose, mouth, eyebrows, etc.

[0044] 7) Classification Model: A machine learning or statistical model is used to classify input data into predefined categories. In liveness detection, a classification model can be used to determine whether an object is truly alive based on extracted key points. The classification model learns to distinguish between real live objects and forged samples through training data sets.

[0045] In the related art, the far and near liveness detection scheme is to display face detection frames of different sizes in the interface. The face detection frame instructs the user to move the face acquisition device or the face from far to near or from near to far. When the boundary conditions of the matching frame are met, the user reaches the target position indicated by the face detection frame. Far and near liveness detection is achieved by instructing the user to complete the detection of multiple target positions. Specifically, the face image is captured during the distance change process to obtain the face video to be processed. According to the timestamp sequence of the user's moving target position, the corresponding frames of each target position are extracted from the face video to be processed. For the intermediate frames between the target positions, the user is constrained to move at a constant speed between the target positions, and the intermediate frames are taken according to the timestamp instructions.

[0046] Figure 1 This is a schematic diagram of the selection of key frames in related technologies. Figure 1 In the above scheme, the distance difference between the final frame and the initial frame is ensured by the user strictly moving to the corresponding face frame. For example, in the initial frame, the user moves to the face detection frame at position 1 at timestamp 1, and in the final frame, the user moves to the face detection frame at position 2 at timestamp 2. This imposes many constraints on the user's behavior and requires sufficient instruction, making it slightly less user-friendly. Selecting the intermediate frame based on the timestamp is based on the user strictly following the distance prompts and moving at a relatively uniform speed. Otherwise, the intermediate frame with a truly effective distance difference cannot be effectively captured. If the user does not move at a uniform speed between position 1 and position 2, for example, the first half is fast and the second half is slow, then selecting according to the intermediate time will result in a video frame that is closer to position 2, which is not the optimal frame. Therefore, the flexibility and efficiency of this liveness detection scheme are poor.

[0047] Furthermore, near-far liveness detection schemes extract facial keypoints from each keyframe and calculate the Euclidean distance between each pair of keypoints to form a new feature matrix. This artificially calculated feature matrix may lose the richer information that distinguishes real and fake faces. Alternatively, extracting facial keypoints from each keyframe and then calculating the positional offset of corresponding keypoints in adjacent frames (called pixel velocity) also loses feature information, reducing liveness detection accuracy.

[0048] Based on the problems existing in the related art, the embodiments of the present application provide a liveness detection method, device, electronic device, computer-readable storage medium and computer program product, which can improve the flexibility, efficiency and accuracy of liveness detection. In the liveness detection method provided in the embodiments of the present application, first, for each image frame in the image frame sequence, the area occupied by the object in the image frame is determined; then, the area occupied by the object is signal processed to obtain a data signal, and the peak and valley values ​​of the data signal are determined; then, a first value is determined between the peak and valley values, and the first image frame corresponding to the peak value, the second image frame corresponding to the valley value and the third image frame corresponding to the first value are determined from the image frame sequence; finally, based on the first image frame, the second image frame and the third image frame, liveness detection is performed on the object.

[0049] The following describes an exemplary application of a liveness detection device provided in an embodiment of the present application, which is an electronic device for implementing a liveness detection method. The electronic device provided in an embodiment of the present application can be implemented as various types of terminals such as a laptop computer, a tablet computer, a desktop computer, a set-top box, a smart phone, a smart speaker, a smart watch, a smart TV, a car terminal, etc., and can also be implemented as a server. Among them, the server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (Content Delivery Network, CDN), and big data and artificial intelligence platforms. The terminal and the server can be directly or indirectly connected via wired or wireless communication, which is not limited in the embodiment of the present application. Below, an exemplary application of the liveness detection device when it is implemented as a terminal or a server will be described.

[0050] See also Figure 2 , Figure 2 : This is a schematic diagram of the architecture of the liveness detection system provided in an embodiment of the present application. In order to perform the liveness detection operation, a liveness detection application can be provided. For example, the liveness detection application can be an application dedicated to liveness detection, or it can be a functional module in other applications (such as a liveness detection module in a financial application, etc.). The liveness detection system 100 in the embodiment of the present application includes at least a terminal 400, a network 300 and a server 200, wherein the server 200 is a server for the liveness detection application. The server 200 can constitute the liveness detection device in the embodiment of the present application, that is, the liveness detection method in the embodiment of the present application is implemented through the server 200. The terminal 400 is connected to the server 200 via the network 300, and the network 300 can be a wide area network or a local area network, or a combination of the two.

[0051] See also Figure 2 The terminal 400 collects image frames of the user at different timestamps in real time on the client side of the liveness detection application. Multiple image frames form an image frame sequence. After collecting the image frame sequence, the client encapsulates the image frame sequence into a liveness detection request and sends the liveness detection request to the server 200 via the network 300. In response to the received liveness detection request, the server 200 determines the area occupied by the object in each image frame in the image frame sequence. The server 200 performs signal processing on the area occupied by the object to obtain a data signal and determines the peak and valley values ​​of the data signal. The server 200 determines a first value between the peak and valley values ​​and determines from the image frame sequence a first image frame corresponding to the peak value, a second image frame corresponding to the valley value, and a third image frame corresponding to the first value. The server 200 performs liveness detection on the object based on the first, second, and third image frames. After obtaining the liveness detection result, the server 200 can send the liveness detection result to the terminal 400. The terminal 400 displays the liveness detection result on the current interface.

[0052] In some embodiments, the terminal 400 may also perform the liveness detection method of the embodiment of the present application, that is, the terminal 400 collects image frames of the user at different timestamps in real time on the client of the liveness detection application, and multiple image frames form an image frame sequence. After the client collects the image frame sequence, the terminal 400 determines the area occupied by the object in each image frame in the image frame sequence; the terminal 400 performs signal processing on the area occupied by the object to obtain a data signal, and determines the peak and valley values ​​of the data signal; the terminal 400 determines a first value between the peak and valley values, and determines the first image frame corresponding to the peak value, the second image frame corresponding to the valley value, and the third image frame corresponding to the first value from the image frame sequence; the terminal 400 performs liveness detection on the object based on the first image frame, the second image frame, and the third image frame. After obtaining the liveness detection result, the terminal 400 displays the liveness detection result on the current interface.

[0053] In an online banking application, users need to verify their identity through liveness detection to ensure the security of their accounts. When a user opens the banking application and enters the identity verification interface, terminal 400 (e.g., the user's smartphone) uses its front-facing camera to capture real-time image frames of the user at different timestamps, forming a sequence of image frames. The client encapsulates the image frame sequence into a liveness detection request and sends it to server 200 via network 300. After receiving the liveness detection request, server 200 begins processing the image frame sequence. For each image frame, server 200 determines the area occupied by the object (i.e., the user's face) in the image frame. Server 200 performs signal processing on the area, extracts the data signal, and determines the peak and valley values ​​of the data signal. Server 200 determines a first value between the peak and valley values ​​and finds the corresponding image frames (the first, second, and third frames). Based on the three image frames, server 200 performs liveness detection to determine whether the user is truly alive. After the liveness detection is complete, server 200 sends the liveness detection result back to terminal 400. The terminal 400 displays the liveness detection result (such as "verification successful" or "verification failed") on the current interface, and decides whether to allow the user to continue the operation based on the liveness detection result.

[0054] At the entrance of an office building or residential complex, an intelligent access control system is installed. This intelligent access control system uses liveness detection technology to identify the identities of people entering and exiting, ensuring that only authorized personnel can enter. A user stands in front of the intelligent access control system's camera. Terminal 400 (e.g., the intelligent access control system's camera device) begins to collect real-time image frames of the user at different timestamps, forming an image frame sequence. The client encapsulates the image frame sequence into a liveness detection request and sends it to server 200 via network 300. After receiving the liveness detection request, server 200 begins processing the image frame sequence. For each image frame, server 200 determines the area occupied by the object (i.e., the user's face) in the image frame. Server 200 performs signal processing on the area, extracts the data signal, and determines the peak and valley values ​​of the data signal. Server 200 determines a first value between the peak and valley values ​​and finds the corresponding image frames (the first, second, and third frames). Based on the three image frames, server 200 performs liveness detection to determine whether the user is truly alive and compares the value with the authorized list. After the liveness detection is completed, the server 200 sends the liveness detection result back to the terminal 400. The terminal 400 displays the liveness detection result (such as "open door" or "deny entry") on the current interface and controls whether the smart access control system is unlocked according to the liveness detection.

[0055] See also Figure 3 , Figure 3 is a structural diagram of an electronic device provided in an embodiment of the present application, Figure 3The electronic device shown includes: at least one processor 410, a memory 450, at least one network interface 420 and a user interface 430. The various components in the electronic device are coupled together via a bus system 440. It is understood that the bus system 440 is used to achieve connection and communication between these components. In addition to including a data bus, the bus system 440 also includes a power bus, a control bus and a status signal bus. However, for the sake of clarity, the bus system 440 is not shown in FIG. Figure 3 Various buses are labeled as bus system 440 .

[0056] The processor 410 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc., where the general-purpose processor can be a microprocessor or any conventional processor, etc.

[0057] The user interface 430 includes one or more output devices 431 that enable presentation of media content, including one or more speakers and / or one or more visual display screens. The user interface 430 also includes one or more input devices 432, including user interface components that facilitate user input, such as a keyboard, mouse, microphone, touch screen display, camera, other input buttons and controls.

[0058] The memory 450 may be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state memory, hard drives, optical drives, etc. The memory 450 may optionally include one or more storage devices that are physically remote from the processor 410.

[0059] The memory 450 includes volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory may be a read-only memory (ROM), and the volatile memory may be a random access memory (RAM). The memory 450 described in the embodiments of the present application is intended to include any suitable type of memory.

[0060] In some embodiments, the memory 450 can store data to support various operations, examples of which include programs, modules, and data structures, or a subset or superset thereof, as exemplified below.

[0061] Operating system 451, including system programs for processing various basic system services and performing hardware-related tasks, such as the framework layer, core library layer, and driver layer, which are used to implement various basic services and process hardware-based tasks;

[0062] A network communication module 452 is used to reach other electronic devices via one or more (wired or wireless) network interfaces 420. Exemplary network interfaces 420 include Bluetooth, Wi-Fi, and Universal Serial Bus (USB);

[0063] a presentation module 453 for enabling presentation of information via one or more output devices 431 (e.g., a display screen, a speaker, etc.) associated with the user interface 430 (e.g., a user interface for operating peripheral devices and displaying content and information);

[0064] The input processing module 454 is configured to detect one or more user inputs or interactions from one of the one or more input devices 432 and to translate the detected inputs or interactions.

[0065] In some embodiments, the apparatus provided in the embodiments of the present application may be implemented in software. Figure 3 A liveness detection device 455 stored in memory 450 is shown. This device can be software in the form of a program or plug-in, and includes the following software modules: an area determination module 4551, a signal processing module 4552, a value determination module 4553, and a liveness detection module 4554. These modules are logical and can be arbitrarily combined or further separated according to the functions they implement. The functions of each module will be described below.

[0066] In other embodiments, the device provided in the embodiments of the present application can be implemented in hardware. As an example, the device provided in the embodiments of the present application can be a processor in the form of a hardware decoding processor, which is programmed to execute the liveness detection method provided in the embodiments of the present application. For example, the processor in the form of a hardware decoding processor can adopt one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs) or other electronic components.

[0067] See also Figure 4 , Figure 4 This is a flow chart of the liveness detection method provided in the embodiment of the present application. Figure 1, will combine Figure 4 The steps shown are explained as Figure 4 As shown, the liveness detection method is described as an example in which the execution subject is a server. The method includes the following steps 101 to 104:

[0068] In step 101 , for each image frame in an image frame sequence, the area occupied by an object in the image frame is determined.

[0069] Here, an image frame sequence is a collection of multiple image frames. Each image frame is an image containing an object captured by a terminal using an image acquisition device (e.g., a camera). Different image frames correspond to different timestamps, where the timestamp corresponding to each image frame is the acquisition time of the image frame. In other words, the image frames in an image frame sequence are arranged in chronological order. The object refers to the target or subject for liveness detection, such as a face, fingerprint, or palm print. Each image frame is acquired by capturing the object at a different acquisition distance, where the acquisition distance is the distance between the terminal's image acquisition device and the object. For example, if the object is a face, the user can freely move the terminal or the face to change the distance between the terminal's image acquisition device and the face, completing the acquisition of the image frame.

[0070] For each image frame, object recognition is performed on the image frame to obtain the object area in the image frame. The bounding box of the object area can be determined, and the area occupied by the object can be calculated based on the bounding box. Alternatively, a binary mask of the object area can be generated to mark the object area. The number of pixels with a value of 1 in the binary mask is counted, and the pixels with a value of 1 represent the object area. The number of pixels with a value of 1 is used as the area occupied by the object. The embodiment of the present application does not limit the computer vision algorithm used for object recognition, for example, it can be a convolutional neural network, a rectangular feature (Haar) cascade classifier, etc.

[0071] For example, when a user logs into a financial application on a mobile terminal, the liveness detection program is triggered. The client interface of the financial application does not need to display a face collection indicator box, but only needs to display the movement mode. For example, if the movement mode is "far-near-far", the user can move the terminal or move the face according to the movement mode so that the collection distance meets the "far-near-far" change, and a sequence of image frames is collected. The face area in each image frame is detected using a face detection algorithm, and the bounding box of the face area is obtained as (x, y, w, h), where x and y are the coordinates of the upper left corner of the bounding box, w is the width, and h is the height. The product of the width and height is determined as the area occupied by the face in the image frame.

[0072] In step 102, signal processing is performed on the area occupied by the object to obtain a data signal, and peak values ​​and valley values ​​of the data signal are determined.

[0073] Here, based on the position of each image frame in the image frame sequence, signal processing can be performed on the area occupied by the object in multiple image frames to obtain a data signal. The position of the image frame in the image frame sequence is used to represent the timestamp of the image frame in the image frame sequence. Signal processing refers to normalizing the area occupied by the object in multiple image frames and using the normalized area as the signal value. The multiple signal values ​​are arranged according to the timestamps of the image frames corresponding to the signal values ​​to obtain a data signal. The data signal is a sequence that reflects the change in the area occupied by the object over time. The data signal can be represented in the form of a waveform using the horizontal axis "time" and the vertical axis "signal value". The embodiments of the present application do not limit the specific method for determining the peak and valley values ​​of the data signal. For example, the differential method can be used to find the local extreme points by calculating the difference between adjacent signal values, and the local extreme points are determined as peak or valley values; the derivative method can also be used to calculate the first-order or second-order derivative of the data signal, and the point where the derivative is zero is used as the peak or valley value; the extreme value detection algorithm (such as the peak finding function find_peaks function) can also be used to determine the peak and valley values ​​from the data signal.

[0074] In some embodiments, see Figure 5 , Figure 5 It is shown that in step 102, the signal processing of the area occupied by the object is performed to obtain a data signal, which can be achieved by the following steps 1021 to 1023:

[0075] In step 1021 , the largest area is screened out from the areas occupied by objects in multiple image frames.

[0076] Here, after determining the area occupied by the object in each image frame, the maximum value is selected from the multiple areas and determined as the maximum area. For example, if there are three image frames, the area occupied by the object in image frame 1 is 10, the area occupied by the object in image frame 2 is 15, and the area occupied by the object in image frame 3 is 9, then the maximum area is 15.

[0077] In step 1022, the ratio of the area occupied by the object to the maximum area in each image frame is determined as a signal value.

[0078] Here, for each image frame, the ratio of the area occupied by the object in the image frame to the maximum area is determined as the signal value corresponding to the image frame. That is, the area occupied by the object in each image frame can be normalized based on the maximum area to obtain a signal value between (0, 1]. For example, there are three image frames in total. The area occupied by the object in image frame 1 is 10, the area occupied by the object in image frame 2 is 15, and the area occupied by the object in image frame 3 is 9. The maximum area is 15. The signal value corresponding to image frame 1 is 10 / 15=0.667, the signal value corresponding to image frame 2 is 15 / 15=1, and the signal value corresponding to image frame 3 is 9 / 15=0.6.

[0079] In step 1023 , based on the position of each image frame in the image frame sequence, the signal values ​​corresponding to the multiple image frames are arranged to obtain a data signal.

[0080] Here, arranging the signal values ​​corresponding to the multiple image frames based on the position of each image frame in the image frame sequence means arranging the signal values ​​of the multiple image frames in order according to the timestamp of each image frame to obtain the arranged multiple signal values. The signal value sequence composed of the arranged multiple signal values ​​is the data signal. For example, the image frame sequence includes L image frames, L is a positive integer, and the data signal is the signal value sequence A = [a1, a2, ..., a L ], where a1 is the first signal value after arrangement, that is, the signal value corresponding to the image frame collected by the terminal at the initial moment, a2 is the second signal value after arrangement, a L is the Lth signal value after arrangement, that is, the signal value corresponding to the image frame collected by the terminal at the last moment.

[0081] The embodiments of the present application perform normalization processing on the area occupied by the object and arrange the data in time sequence, thereby generating a data signal that can accurately reflect the dynamic changes in the object acquisition distance, thereby significantly improving the accuracy and reliability of liveness detection.

[0082] In some embodiments, determining the peak and valley values ​​of the data signal in step 102 can be achieved in the following manner: first, based on the arrangement position of each signal value in the data signal, three consecutive signal values ​​are arbitrarily selected; then, if among the three consecutive signal values, the signal value at the middle position is greater than the other two signal values, the signal value at the middle position is determined as the peak value of the data signal; if among the three consecutive signal values, the signal value at the middle position is less than the other two signal values, the signal value at the middle position is determined as the valley value of the data signal.

[0083] Here, the arrangement position of the signal value in the data signal is the arrangement position of the image frame corresponding to the signal value in the image frame sequence. Three signal values ​​with consecutive arrangement positions are selected from the data signal, and the signal value at the middle position is compared with the numerical value between the other two signal values. The number of peak values ​​in the data signal can be one or more, and the number of valley values ​​can be one or more. If the peak value or valley value cannot be determined from the data signal using the above method, the maximum signal value in the data signal is used as the peak value, and the minimum signal value in the data signal is used as the valley value. The number of peak values ​​and valley values ​​determined from the data signal can be determined based on a preset movement pattern during the image frame acquisition process.

[0084] For example, if the preset movement mode during the image frame acquisition process is "far-near-far", then the number of peaks in the data signal is two and the number of valleys is one. The data signal is a signal value sequence A = [a1, a2, ..., a L ], randomly select three consecutive signal values ​​from the data signal: a1, a2, a3, if the signal value a2 at the middle position is greater than the signal value a1 and the signal value a3, then the signal value a2 is the peak value. Randomly select three consecutive signal values ​​from the data signal: a4, a5, a6, if the signal value a5 at the middle position is less than the signal value a4 and the signal value a6, then the signal value a5 is the valley value. Randomly select three consecutive signal values ​​from the data signal: a8, a9, a10, a111, a121, a131, a141, a151 10 , if the signal value a9 at the middle position is greater than the signal value a8 and the signal value a 10 , then the signal value a9 is the peak value.

[0085] It should be noted that if the number of peaks or valleys determined from the data signal based on the above method does not match the preset movement mode during the image frame acquisition process, for example, if the number of peaks is less than the number of "far" values ​​set in the movement mode, the following situations can be considered:

[0086] Case 1: Get multiple signal values ​​that are consecutive to the first signal value. The number of these multiple signal values ​​can be set based on actual needs. If the first signal value is greater than the multiple signal values, the first signal value can be used as the peak value. Alternatively, get multiple signal values ​​that are consecutive to the last signal value. The number of these multiple signal values ​​can be set based on actual needs. If the last signal value is greater than the multiple signal values, the last signal value can be used as the peak value. For example, the preset movement mode in the image frame acquisition process is "far-near-far". When the user performs liveness detection, the acquisition distance is maximum at the first image frame acquired. Then the user gradually reduces the acquisition distance from the terminal, and when reaching a certain position, increases the acquisition distance from the terminal. In this scenario, the first image frame acquired is the image frame acquired in the "far" movement mode, and the first signal value corresponding to the first image frame can be used as a peak value.

[0087] Case 2: Obtain multiple consecutive, equal signal values ​​from the data signal. Only one of these equal signal values ​​is retained, and the remaining equal signal values ​​are deleted to obtain a new data signal. The retained signal value is then used as the middle signal value, and two signal values ​​that are consecutive to this signal value are selected from the data signal to form three consecutive signal values. If the middle signal value of the three consecutive signal values ​​is greater than the other two signal values, the middle signal value is determined to be the peak value of the data signal.

[0088] When the number of valley values ​​does not match the preset movement pattern during the image frame acquisition process, you can refer to the above-mentioned situations, which will not be explained here.

[0089] The embodiment of the present application selects the continuous signal values ​​of any three positions in the data signal, and determines the peak and valley values ​​of the data signal by comparing the signal value at the middle position with the size of the other two signal values. It can efficiently and accurately identify the key moments when the object area changes, and then reflect the key moments when the acquisition distance changes, thereby accurately reflecting the user's movement pattern during liveness detection or dynamic acquisition. For example, under the preset movement mode of "far-near-far", this method can accurately locate two peaks (corresponding to the two moments when the user is farthest from the terminal) and one valley value (corresponding to the moment when the user is closest to the terminal), providing a reliable basis for subsequent analysis. In addition, for situations where the number of peak or valley values ​​does not conform to the preset pattern, a flexible adjustment strategy is introduced, such as considering boundary signal values ​​or processing continuous equal signal values, which further improves the robustness and adaptability of the algorithm. The embodiments of the present application not only simplify the feature extraction process of complex signals, but also significantly improve the accuracy of user behavior pattern recognition. Especially in the liveness detection scenario, it is convenient to extract key frames for liveness detection. There is no need to display a face indicator box on the interface to determine the image frame corresponding to the face indicator box as a key frame, which significantly improves the flexibility and efficiency of liveness detection.

[0090] In step 103 , a first value is determined between the peak value and the valley value, and a first image frame corresponding to the peak value, a second image frame corresponding to the valley value, and a third image frame corresponding to the first value are determined from the image frame sequence.

[0091] Here, if the data signal includes only one peak value and one valley value, one or more first values ​​can be directly determined between the numerical range of the peak value and the valley value. If the data signal includes multiple peak values ​​and / or multiple valley values, the peak values ​​and valley values ​​can be sorted in the order of the timestamps of the image frames to obtain a sorted peak-valley value sequence. For any adjacent peak values ​​and valley values ​​in the peak-valley value sequence, at least one first value can be determined between the peak values ​​and the valley values. For each first value, the target signal value with the smallest difference from the first value is determined from the data signal, and the image frame corresponding to the target signal value is determined as the third image frame. The image frame corresponding to the peak value is determined as the first image frame, and the image frame corresponding to the valley value is determined as the second image frame.

[0092] In an embodiment of the present application, if the data signal includes only one peak value and one valley value, in step 103, determining the first numerical value between the peak value and the valley value can be achieved in the following manner: first, determining the difference between the peak value and the valley value; then, obtaining a preset quantity threshold, where j is a positive integer starting from 1 and increasing sequentially to the quantity threshold; then, performing a preset operation on the difference, the quantity threshold and j to obtain the operation result; finally, determining the sum of the operation result and the valley value as the jth first numerical value.

[0093] Here, performing a preset operation on the difference, the quantity threshold, and j can be achieved by determining the sum of the quantity threshold and the preset value "1", and multiplying the ratio of j to the sum by the difference to obtain the operation result. The sum of the result of the preset operation and the valley value is determined as the jth first value. The jth first value satisfies the following formula (1).

[0094] v j =p l +(p h -p l )×j / (m+1) formula (1);

[0095] Among them, v j represents the jth first value, p l Indicates the valley value, p h represents the peak value, and m represents the quantity threshold.

[0096] The embodiment of the present application uses a method for determining a first numerical value based on the operation of a difference and a preset quantity threshold, which can accurately capture the key transition state of the object area change, ensure that the first numerical value is representative, and thus more comprehensively reflect the dynamic characteristics, significantly improving the accuracy and robustness of liveness detection.

[0097] In some embodiments, if the data signal includes multiple peak values ​​and / or multiple valley values, determining the first value between the peak values ​​and the valley values ​​in step 103 can be accomplished by sorting the peak values ​​and valley values ​​in the data signal according to the timestamp order of the image frames to obtain a sorted sequence of peak and valley values. From the sequence of peak and valley values, a second value and a third value that is continuous with the second value are selected. The first value is determined based on the difference between the second value and the third value and a preset quantity threshold.

[0098] Here, the second value is any peak or valley value among multiple peaks and valley values. When the second value is a peak, the third value that is continuous with the second value is a valley value; when the second value is a valley value, the third value that is continuous with the second value is a peak. The image frame corresponding to the third value is located before the image frame corresponding to the second value in the image frame sequence. Taking the movement mode of "far-near-far" as an example, two peaks and one valley value are obtained. The sorted peak and valley value sequence is [p1, p2, p3], where p1 and p3 are peaks and p2 is a valley value. If the second value is filtered out and it can be p2, the third value that is continuous in position is p1. If the second value is filtered out and it can be p3, the third value that is continuous in position is p2.

[0099] The preset quantity threshold refers to the number of first values ​​determined between each second value and its continuous third value. For different sets of second values ​​and third values, the value of the quantity threshold can be the same or different, and the quantity threshold can be set based on actual needs. For example, the quantity threshold is 1, and the sorted peak-valley value sequence is [p1, p2, p3], p1 and p3 are peak values, and p2 is a valley value. Then, one first value can be determined between the second value p2 and the third value p1, and another first value can be determined between the second value p3 and the third value p2. For each second value, the difference between the second value and the third value in a continuous position can be determined, and the first value between the second value and the third value in a continuous position can be determined based on the difference and the preset quantity threshold. The specific process of determining the first value is similar to the specific process of determining the first value between the peak value and the valley value when the data signal includes only one peak value and one valley value in the above embodiment.

[0100] Determining the first value based on the difference between the second value and the third value and a preset quantity threshold can be achieved by performing a preset operation on the difference, the quantity threshold, and j to obtain an operation result, where j is a positive integer; then, summing the result of the preset operation and the third value to determine the jth first value, where the image frame corresponding to the third value is located before the image frame corresponding to the second value in the image frame sequence. The first value can satisfy the following formula (2).

[0101] v j =p i-1 +(p i -p i-1 )×j / (m+1) formula (2);

[0102] Among them, v j represents the jth first value, p i-1 Represents the third value, that is, the i-1th value in the peak-valley value sequence, p i represents the second value, that is, the i-th value in the peak-valley value sequence, and m represents the quantity threshold.

[0103] By filtering the second value between peaks and valleys and their consecutive third values, and determining the first value based on the difference and a preset threshold, this embodiment of the application accurately captures the key intermediate values ​​of the object's area change, thereby identifying key frames for near and far liveness detection and ensuring that the key frames are representative. Compared to methods that directly identify image frames at intermediate times as key frames, this embodiment of the application does not require the user to move at a constant speed, thereby improving the accuracy and robustness of liveness detection.

[0104] In step 104 , liveness detection is performed on the object based on the first image frame, the second image frame, and the third image frame.

[0105] Here, the first image frame, the second image frame, and the third image frame may be used as key frames, and liveness detection may be performed on the object based on the key frames to obtain liveness detection results.

[0106] In some embodiments, see Figure 6 , Figure 6 It is shown that in step 104, liveness detection of the object is performed based on the first image frame, the second image frame, and the third image frame, which can be achieved by the following steps 1041 to 1043:

[0107] In step 1041 , the first coordinates of each key point in the key frame are determined.

[0108] The key frame is an image frame among the first image frame, the second image frame and the third image frame.

[0109] Here, for each key frame, key point detection can be performed on the object in the key frame to obtain multiple key points. For example, taking the object as a face, the key points can include parts such as eyes, nose, and mouth. The coordinates of the key points can be directly determined as the first coordinates, and the first coordinates include the coordinate value x of the horizontal axis (X) and the coordinate value y of the vertical axis (Y). Alternatively, the key points can be horizontally rotated so that the key points in the eye area are in a horizontal position, and the coordinates of the key points after horizontal rotation are determined as the first coordinates.

[0110] In some embodiments, determining the first coordinate of each key point in a key frame can be achieved in the following manner: first, performing key point detection on the object in the key frame to obtain the key points in the key frame; then, from the key points in the key frame, screening out the first key point corresponding to the first part of the object and the second key point corresponding to the second part; then, based on the coordinates of the first key point and the coordinates of the second key point, determining the rotation angle of the key frame; finally, mapping the coordinates of each key point based on the rotation angle to obtain the first coordinate of each key point in the key frame.

[0111] Here, a computer vision algorithm (such as the cross-platform computer vision library OpenCV) can be used to detect key points of the object in the key frame to obtain multiple key points in the key frame. From the key points in the key frame, a first key point corresponding to the first part of the object and a second key point corresponding to the second part are screened out, wherein the vertical distance between the first part and the second part is less than a preset threshold, that is, the first part and the second part are theoretically in a horizontal position. The preset threshold can be set voluntarily, for example, to 0.05. After determining the coordinates of the first key point and the coordinates of the second key point, the vector between the first key point and the second key point can be determined based on the coordinates of the first key point and the coordinates of the second key point. The angle between the vector and the horizontal axis (x-axis) is determined as the rotation angle. A rotation matrix can be constructed based on the rotation angle. The rotation matrix is ​​used to rotate a point on a two-dimensional plane around the origin by the rotation angle, and the rotation can be clockwise or counterclockwise. The cosine value of the rotation angle can be determined by the cosine function, the sine value of the rotation angle can be determined by the sine function, and the rotation matrix can be constructed based on the cosine value and the sine value. Mapping the coordinates of each key point based on the rotation angle can be achieved by multiplying the coordinates of each key point by the rotation matrix to obtain the first coordinate of each key point in the key frame.

[0112] For example, the object is a human face, the first part is the left eye, and the second part is the right eye. The coordinates of the center of the left eye are determined as the coordinates of the first key point (x1, y1), and the coordinates of the center of the right eye are determined as the coordinates of the second key point (x2, y2). The vector between the first key point and the second key point is calculated. The angle θ between the vector and the horizontal axis is determined based on the inverse tangent function, that is, the rotation angle θ = arctan((x2-x1) / (y2-y1)). The rotation matrix is Multiply the coordinates (x, y) of each key point by the rotation matrix to obtain the first coordinate of each key point.

[0113] The embodiment of the present application calculates the rotation angle of the key frame through the coordinates of the key points of the first part and the second part of the object, and maps the coordinates of all key points based on the rotation angle, and finally obtains the adjusted first coordinate of each key point. This not only improves the accuracy of key point positioning, but also effectively solves the coordinate offset problem caused by changes in shooting angle or object posture, thereby providing more stable and reliable basic data support for subsequent liveness detection tasks.

[0114] In step 1042 , the origin of the key points in the key frame is determined, and the difference between the first coordinate of each key point and the coordinate of the origin is determined as the second coordinate of each key point.

[0115] The origin of the keypoints in the keyframe can be customized based on actual needs. For example, for a face, the keypoint at the tip of the nose can be used as the origin (x0, y0). Subtracting the origin's coordinates from each keypoint's first coordinate (x, y) yields the second coordinate (x' = x-x0, y' = y-y0). This second coordinate eliminates the horizontal displacement of the face relative to the camera, retaining only the near-far displacement, improving the accuracy of liveness detection.

[0116] In step 1043, based on the second coordinate of each key point, the objects in the key frame are classified to obtain a living body detection result.

[0117] Here, the second coordinates of the key points can be normalized to obtain the third coordinates of the key points. Normalization can be achieved by determining the maximum coordinate value of the horizontal axis and the maximum coordinate value of the vertical axis for each of the multiple key points. For each key point, the ratio of the horizontal coordinate value of the key point to the maximum coordinate value of the horizontal axis is determined as the normalized horizontal coordinate value, and the ratio of the vertical coordinate value of the key point to the maximum coordinate value of the vertical axis is determined as the normalized vertical coordinate value. The normalized horizontal coordinate value and the normalized vertical coordinate value constitute the third coordinate of the key point. The third coordinates of the key points of multiple key frames are spliced ​​to obtain key point data. The key point data is classified using a pre-trained classification model to obtain a classification value. The classification model is a binary classification model. Therefore, when the classification value is greater than or equal to a preset threshold, the liveness detection result is determined to be non-live; when the classification value is less than the preset threshold, the liveness detection result is determined to be live. Exemplarily, the preset threshold is 1, that is, when the classification value output by the classification model is 1, the liveness detection result is determined to be non-liveness; when the classification value output by the classification model is 0, the liveness detection result is determined to be liveness.

[0118] The embodiments of the present application effectively reduce the interference caused by environmental factors and posture changes through precise key point coordinate conversion and relative position analysis, thereby improving the accuracy and reliability of liveness detection.

[0119] In some embodiments, features in other dimensions may be determined based on the second coordinates of the key points, and objects in the keyframes may be classified based on the features in other dimensions to obtain liveness detection results. For example, the square of the coordinate value in the second coordinate may be used as a feature, or the coordinate offset of each key point (i.e., the difference between the coordinate values ​​of two key points) may be used as a feature.

[0120] During the liveness detection process, the embodiment of the present application performs signal processing on the area occupied by the object in multiple image frames to obtain a data signal, and determines the peak and valley values ​​of the data signal. Based on the first image frame corresponding to the peak value, the second image frame corresponding to the valley value, and the third image frame corresponding to the first value between the peak and valley values, the liveness detection of the object is achieved. Compared with the method in the related art that strictly requires the user to move at a constant speed from far to near or from near to far to collect image frames, and relies on timestamps to select intermediate image frames, the embodiment of the present application analyzes the data signal obtained by area changes and flexibly selects the first image frame, the second image frame, and the third image frame. This can more accurately capture image frames of the object at different distances, reduce the constraints on the user's behavior, and improve the flexibility of liveness detection. At the same time, the embodiment of the present application can accurately identify the most representative image frames in a shorter time, thereby improving the efficiency of liveness detection.

[0121] The following describes an exemplary application of the embodiments of the present application in a practical application scenario.

[0122] In near-far liveness detection technology, a face frame is displayed on the client interface. The user must follow the instructions to move the face into the designated face frame and ensure that the face size is close to the face frame size. This operation is complicated, resulting in reduced flexibility and efficiency of liveness detection. The present application provides a liveness detection method, which is an improved interactive method for near-far liveness detection. This method does not require the face frame to be pre-displayed on the liveness detection client interface. The user only needs to change the acquisition position, and the key frame is automatically identified through the subsequent algorithm processing to complete liveness detection.

[0123] Figure 7 This is a block diagram of the liveness detection method provided in an embodiment of the present application.

[0124] See also Figure 7 , step 201, collecting video by moving far and near.

[0125] Among them, when the user performs liveness detection through the terminal, there is no need to display a face detection frame on the terminal interface, and it is only necessary to provide text prompt information indicating the user's movement method. For example, the text prompt information can be "move from far to near", "move from near to far", "far-near-far" or "near-far-near", etc. The user moves the terminal or moves the face according to the text prompt information, so that the distance between the terminal's image acquisition device and the face changes according to the requirements of the text prompt information. The terminal can capture the user's face image (corresponding to the image frame in the above embodiment) through an image acquisition device (such as a camera), and multiple frames of face images constitute a video (corresponding to the image frame sequence in the above embodiment).

[0126] Step 202: normalization processing.

[0127] The following preprocessing can be performed on each image frame in the video: first, a pre-trained face detection model is used to detect the face in the image frame and calculate the face area (corresponding to the area occupied by the object in the above embodiment). Then, the maximum face area (corresponding to the maximum area in the above embodiment) is screened out from the face areas of multiple image frames, and the face area of ​​each image frame is normalized to the interval (0, 1) according to the maximum face area, so as to obtain a normalized area sequence A = [a1, a2, ..., a L ], the sequence length is L, and L is the number of video frames.

[0128] Step 203: locate the peak points and valley points.

[0129] Among them, a signal processing algorithm can be used to locate at least one peak point in the normalized area sequence A. The distance between adjacent peak points is greater than or equal to L / (k-1), where k is the number of preset target positions. For example, the text prompt information can be "moving from far to near", then the number of target positions is 2, one target position is "far", and one target position is "near"; the text prompt information can be or, "far-near-far", then the number of target positions is 3, one target position is "far", one target position is "near", and one target position is "far". Use the same signal processing algorithm to process the sequence A'=[1-a1,1-a2,…,1-a L ] is processed to obtain at least one trough point in the normalized area sequence A, and the distance between adjacent trough points is greater than or equal to L / (k-1). When moving from far to near to far, there is 1 peak point and 2 trough points. When moving from near to far to near, there are 2 peak points and 1 trough point. The normalized face areas of all the image frames corresponding to the peak points and trough points are sorted by timestamp. Taking k=3 as an example, the area sequence [p1, p2, p3] corresponding to the peak and trough points is obtained. These peak and trough points ideally correspond to the far and near target positions prompted by the interface.

[0130] Exemplarily, the signal processing algorithm can be implemented based on a function for finding signal peaks in the signal. The basic principle is to identify local maximum points. The characteristic of local maximum points is that they are higher than other adjacent points within a certain range. The working principle of the signal peak search function is as follows: Local maximum: traverse the data points in the normalized area sequence A and find points that meet the local maximum conditions. If a data point is higher than the points on its left and right, then it is considered to be a local maximum. Height threshold: A minimum height threshold (height parameter) can be set, and only peaks above this minimum height threshold will be detected. Distance threshold: Set the minimum horizontal distance between peaks (distance parameter) to avoid detecting multiple peaks that are very close.

[0131] Step 204: Select the middle point between adjacent key frames according to the area change amplitude.

[0132] Among them, the image frames corresponding to the peak point and the valley point are taken as key frames. The number of intermediate frames extracted between every two adjacent far and near target positions in the area sequence [p1, p2, p3] corresponding to the peak and valley points is preset to be m, where m is a positive integer. First, any two adjacent area values ​​p in the area sequence [p1, p2, p3] can be determined. i-1 and p i , i is an integer greater than 1 and less than or equal to k. Then, determine two area values ​​p i-1 and p i The ideal area value v between j =p i-1 +(p i -p i-1 )×j / (m+1), j=1,…,m. corresponds to v j The image frame with the closest value is the extracted intermediate frame (corresponding to the third image frame in the above embodiment), and the intermediate frame is also used as a key frame. Figure 8 This is a schematic diagram of extracting each key frame provided in the embodiment of this application. Figure 8 , the horizontal axis is the timestamp, the vertical axis is the normalized face area, when k = 3, m = 1, a total of 6 key frames are extracted ( Figure 8 The key points are represented by “x”).

[0133] Step 205: extract facial key points in the key frame.

[0134] In this process, facial key points are detected for each key frame (including image frames corresponding to peaks and valleys, as well as extracted intermediate frames). The image is then horizontally rotated so that both eyes are in a horizontal position, and the x and y coordinates of each key point in the image are obtained. The key point located at the center of the face is used as the origin (e.g., the key point corresponding to the tip of the nose) (x0, y0). The coordinates of all key points (corresponding to the first coordinate in the above embodiment) are subtracted from the origin coordinates to perform a coordinate transformation. The transformed coordinates (corresponding to the second coordinate in the above embodiment) are x' = x-x0, y' = y-y0, eliminating the horizontal displacement of the face relative to the camera and retaining only the displacement information in the near and far directions. The maximum absolute value of all key points along each coordinate axis is calculated (max(abs(x')max(abs(y')), where max is the maximum value and abs is the absolute value). The coordinates of all key points along each coordinate axis are then normalized by dividing them by the maximum value to obtain the key point data x' / max(abs(x'), y' / max(abs(y')).

[0135] Step 206: Perform classification processing based on the classification model.

[0136] The keypoint data from each keyframe in the video can be concatenated to form a TxNx2 matrix, where T is the number of keyframes, N is the number of keypoints in each keyframe, and 2 is the number of coordinate axes. This concatenated matrix is ​​used as input to the classification model, which then outputs the liveness detection result. For example, if the classification model outputs a value of "0," the liveness detection result is real, while if it outputs a value of "1," the liveness detection result is fake. Training data can be collected in advance and used to train a binary classification convolutional model to obtain the classification model.

[0137] Training Data Collection: During the training phase, data collectors have wide latitude in performing random movements from far to near or from near to far, without having to adhere to the face prompt box's range constraints. A preset number of target locations, k (e.g., k = 3), can be used to execute a far-near-far or near-far-near movement pattern, collecting a certain amount of video data, including both real and fake subjects. The video data is then processed through various steps, and the resulting key point data is used as training data.

[0138] In an embodiment of the present application, liveness detection is achieved by detecting peaks and valleys in a normalized facial area sequence. Users no longer need to strictly follow the facial frame instructions to reach the target location to complete liveness detection based on distance judgment; they only need to reach the vicinity of the indicated target location, which increases ease of use and interaction efficiency. By extracting intermediate frames between the near and far target locations based on the size of the facial area, the optimal keyframes that facilitate classification detection can be accurately selected. By splicing key point coordinate vectors, the convolutional network learns classification features, which are richer than manually calculated key point distance features and improve the accuracy of liveness detection.

[0139] The following continues to describe the exemplary structure of the liveness detection device 455 provided in the embodiment of the present application as a software module. In some embodiments, such as Figure 3 As shown, the software modules stored in the living body detection device 455 of the memory 450 may include:

[0140] The area determination module 4551 is configured to determine, for each image frame in the image frame sequence, an area occupied by an object in the image frame.

[0141] The signal processing module 4552 is used to perform signal processing on the area occupied by the object to obtain a data signal and determine the peak value and valley value of the data signal.

[0142] The value determination module 4553 is used to determine a first value between the peak value and the valley value, and determine the first image frame corresponding to the peak value, the second image frame corresponding to the valley value, and the third image frame corresponding to the first value from the image frame sequence.

[0143] The living body detection module 4554 is used to perform living body detection on the object based on the first image frame, the second image frame and the third image frame.

[0144] In some embodiments, the signal processing module 4552 is also used to filter out the maximum area from the areas occupied by objects in multiple image frames; determine the ratio of the area occupied by the object in each image frame to the maximum area as the signal value; and arrange the signal values ​​corresponding to multiple image frames based on the position of each image frame in the image frame sequence to obtain a data signal.

[0145] In some embodiments, the signal processing module 4552 is also used to arbitrarily select three consecutive signal values ​​based on the arrangement position of each signal value in the data signal; if among the three consecutive signal values, the signal value at the middle position is greater than the other two signal values, then the signal value at the middle position is determined as the peak value of the data signal; if among the three consecutive signal values, the signal value at the middle position is less than the other two signal values, then the signal value at the middle position is determined as the valley value of the data signal.

[0146] In some embodiments, the numerical determination module 4553 is also used to determine the difference between the peak value and the valley value; obtain a preset quantity threshold, where j is a positive integer starting from 1 and increasing sequentially to the quantity threshold; perform a preset operation on the difference, the quantity threshold and j to obtain the operation result; and determine the sum of the operation result and the valley value as the jth first numerical value.

[0147] In some embodiments, the liveness detection module 4554 is also used to determine the first coordinate of each key point in the key frame, where the key frame is an image frame among the first image frame, the second image frame and the third image frame; determine the origin of the key point in the key frame, and determine the difference between the first coordinate of each key point and the coordinate of the origin as the second coordinate of each key point; based on the second coordinate of each key point, classify the objects in the key frame to obtain the liveness detection result.

[0148] In some embodiments, the liveness detection module 4554 is also used to perform key point detection on the object in the key frame to obtain the key points in the key frame; from the key points in the key frame, filter out the first key point corresponding to the first part of the object and the second key point corresponding to the second part; determine the rotation angle of the key frame based on the coordinates of the first key point and the coordinates of the second key point; map the coordinates of each key point based on the rotation angle to obtain the first coordinate of each key point in the key frame.

[0149] An embodiment of the present application provides a computer program product, which includes a computer program or computer-executable instructions stored in a computer-readable storage medium. A processor of an electronic device reads the computer-executable instructions from the computer-readable storage medium and executes the computer-executable instructions, causing the electronic device to perform the liveness detection method described in the embodiment of the present application.

[0150] The embodiment of the present application provides a computer-readable storage medium in which computer-executable instructions or computer programs are stored. When the computer-executable instructions or computer programs are executed by a processor, the processor will execute the living body detection method provided by the embodiment of the present application, for example, Figure 4 The liveness detection method is shown.

[0151] In some embodiments, the computer-readable storage medium may be a memory such as RAM, ROM, flash memory, magnetic surface memory, optical disk, or CD-ROM; or may be various devices including one or any combination of the above memories.

[0152] In some embodiments, computer-executable instructions may be in the form of a program, software, software module, script, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment.

[0153] As an example, computer-executable instructions may, but need not, correspond to a file in a file system, may be stored as part of a file that stores other programs or data, such as in one or more scripts in a HyperText Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple coordinating files (e.g., files storing one or more modules, subroutines, or code portions).

[0154] By way of example, computer-executable instructions may be deployed to be executed on one electronic device, or on multiple electronic devices located at one site, or on multiple electronic devices distributed across multiple sites and interconnected by a communication network.

[0155] In summary, the embodiments of the present application can complete near and far liveness detection without displaying a face prompt box on the interface, thereby improving the flexibility and efficiency of detection, and using key point coordinates for binary classification to obtain liveness detection results, which can improve detection accuracy.

[0156] The above description is merely an embodiment of the present application and is not intended to limit the scope of protection of the present application. Any modifications, equivalent replacements, and improvements made within the spirit and scope of the present application are included in the scope of protection of the present application.

Claims

1. A method for detecting a living body, characterized in that: The method comprises: For each image frame in the image frame sequence, determining an area occupied by an object in the image frame; performing signal processing on the area occupied by the object to obtain a data signal, and determining peak values ​​and valley values ​​of the data signal; Determine a first value between the peak value and the valley value, and determine a first image frame corresponding to the peak value, a second image frame corresponding to the valley value, and a third image frame corresponding to the first value from the image frame sequence; Perform livingness detection on the object based on the first image frame, the second image frame, and the third image frame.

2. The method according to claim 1, characterized in that The performing signal processing on the area occupied by the object to obtain a data signal includes: Filtering out the largest area from the areas occupied by the object in the plurality of image frames; determining a ratio of an area occupied by the object in each image frame to the maximum area as a signal value; Based on the position of each image frame in the image frame sequence, the signal values ​​corresponding to a plurality of the image frames are arranged to obtain the data signal.

3. The method according to claim 2, characterized in that Determining the peak value and the valley value of the data signal includes: Based on the arrangement position of each signal value in the data signal, arbitrarily select three consecutive signal values; If, among the three consecutive signal values, the signal value at the middle position is greater than the other two signal values, the signal value at the middle position is determined as the peak value of the data signal; If, among the three consecutive signal values, the signal value at the middle position is smaller than the other two signal values, the signal value at the middle position is determined as the valley value of the data signal.

4. The method according to claim 1, wherein The determining a first value between the peak value and the valley value includes: determining a difference between the peak value and the valley value; Obtain a preset quantity threshold, where j is a positive integer starting from 1 and increasing sequentially to the quantity threshold; Performing a preset operation on the difference, the quantity threshold, and j to obtain an operation result; The sum of the operation result and the valley value is determined as the jth first value.

5. The method according to any one of claims 1 to 4, characterized in that The performing living body detection on the object based on the first image frame, the second image frame, and the third image frame includes: determining a first coordinate of each key point in a key frame, wherein the key frame is an image frame among the first image frame, the second image frame, and the third image frame; Determine the origin of the key points in the key frame, and determine the difference between the first coordinate of each key point and the coordinate of the origin as the second coordinate of each key point; Based on the second coordinate of each key point, the object in the key frame is classified to obtain a living body detection result.

6. The method according to claim 5, characterized in that Determining the first coordinate of each key point in the key frame includes: Performing key point detection on the object in the key frame to obtain key points in the key frame; Filtering out, from the key points in the key frame, a first key point corresponding to a first part of the object and a second key point corresponding to a second part of the object; determining a rotation angle of the key frame based on the coordinates of the first key point and the coordinates of the second key point; The coordinates of each key point are mapped based on the rotation angle to obtain the first coordinates of each key point in the key frame.

7. A living body detection device, characterized in that: The device comprises: an area determination module, configured to determine, for each image frame in the image frame sequence, an area occupied by an object in the image frame; a signal processing module, configured to perform signal processing on the area occupied by the object to obtain a data signal, and determine peak values ​​and valley values ​​of the data signal; a value determination module, configured to determine a first value between the peak value and the valley value, and determine, from the image frame sequence, a first image frame corresponding to the peak value, a second image frame corresponding to the valley value, and a third image frame corresponding to the first value; A living body detection module is used to perform living body detection on the object based on the first image frame, the second image frame and the third image frame.

8. An electronic device, characterized in that: The electronic device comprises: a memory for storing computer-executable instructions or computer programs; The processor is configured to implement the liveness detection method according to any one of claims 1 to 6 when executing the computer executable instructions or computer program stored in the memory.

9. A computer-readable storage medium storing computer-executable instructions or a computer program, characterized in that: When the computer executable instructions or computer program are executed by a processor, the living body detection method according to any one of claims 1 to 6 is implemented.

10. A computer program product comprising computer executable instructions or a computer program, characterized in that When the computer executable instructions or computer program are executed by a processor, the living body detection method according to any one of claims 1 to 6 is implemented.