Face living body detection method and device, and electronic equipment

By acquiring the high-frequency and low-frequency feature components of video image frame sequences, calculating the inter-frame correlation, and matching the inter-frame correlation of live faces, the problem of low accuracy in face liveness detection on front-end devices is solved, achieving more efficient liveness detection.

CN116935457BActive Publication Date: 2026-05-08HANVON CORP
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HANVON CORP
Filing Date
2022-04-01
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

Existing face liveness detection methods have low accuracy on front-end devices, and methods based on deep features are computationally complex and not suitable for front-end devices.

Method used

By acquiring high-frequency and low-frequency feature components in the video image frame sequence, calculating the inter-frame correlation, and matching it with the pre-acquired inter-frame correlation of live faces, it is possible to detect whether the face to be detected is a live face.

Benefits of technology

It improves the accuracy of face liveness detection and reduces computational complexity, making it more suitable for front-end devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116935457B_ABST
    Figure CN116935457B_ABST
Patent Text Reader

Abstract

The application discloses a face living body detection method, belongs to the pattern recognition technical field, and is helpful to improve the face living body detection accuracy. The method comprises the following steps: acquiring high-frequency feature components and low-frequency feature components of each video image in a video image frame sequence used for face living body detection; acquiring interframe correlation of the video image frame sequence according to the high-frequency feature components and the low-frequency feature components of the video image; and detecting whether a to-be-detected face collected by the video image is a living body face according to a matching relationship between the interframe correlation and pre-acquired interframe correlation of a living body face. The method is based on the interframe detail change features of the video images in the video image sequence, the invariability features of the general situation and the contour, and the matching result of the distribution law of the above features in the living body face image, so that whether the video image sequence is a living body face is judged, and the accuracy of the living body face detection can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of pattern recognition technology, and in particular to methods, apparatus, electronic devices and computer-readable storage media for face liveness detection. Background Technology

[0002] Current liveness detection methods are mostly based on two-dimensional image features. Traditional image-based liveness feature methods have low accuracy and cannot achieve ideal results against attacks such as photos and videos. While depth-based methods for liveness detection can improve the accuracy of face liveness detection, they require the use of binocular image acquisition equipment or depth image acquisition equipment to extract depth features. Furthermore, depth feature calculation is complex and not suitable for front-end devices.

[0003] It is evident that existing methods for face liveness detection applied to front-end devices still require improvement. Summary of the Invention

[0004] This application provides a method for face liveness detection, which helps to improve the accuracy of face liveness detection.

[0005] In a first aspect, embodiments of this application provide a method for face liveness detection, including:

[0006] Obtain the high-frequency and low-frequency feature components of each frame of the video image sequence used for face liveness detection;

[0007] Based on the high-frequency feature components and low-frequency feature components of the video image, the inter-frame correlation of the video image frame sequence is obtained, wherein the inter-frame correlation includes inter-frame consistency and inter-frame difference;

[0008] Based on the matching relationship between the inter-frame correlation and the pre-acquired inter-frame correlation of live faces, it is determined whether the face to be detected in the video image is a live face.

[0009] Secondly, embodiments of this application provide a device for face liveness detection, comprising:

[0010] The high and low frequency component acquisition module is used to acquire the high frequency feature components and low frequency feature components of each video image in the video image frame sequence used for face liveness detection.

[0011] The inter-frame correlation acquisition module acquires the inter-frame correlation of the video image frame sequence based on the high-frequency feature components and the low-frequency feature components of the video image, wherein the inter-frame correlation includes inter-frame consistency and inter-frame difference.

[0012] The liveness detection module is used to detect whether the face to be detected in the video image is a live face based on the inter-frame correlation and the matching relationship between the inter-frame correlation and the pre-acquired inter-frame correlation of live faces.

[0013] Thirdly, embodiments of this application also disclose an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the face liveness detection method described in embodiments of this application.

[0014] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, comprises the steps of the face liveness detection method disclosed in embodiments of this application.

[0015] The face liveness detection method disclosed in this application obtains the high-frequency and low-frequency feature components of each frame of a video image in a video image frame sequence used for face liveness detection; obtains the inter-frame correlation of the video image frame sequence based on the high-frequency and low-frequency feature components of the video image; and detects whether the face to be detected in the video image is a live face based on the matching relationship between the inter-frame correlation and the pre-obtained inter-frame correlation of live faces, which helps to improve the accuracy of face liveness detection.

[0016] The above description is only an overview of the technical solution of this application. In order to better understand the technical means of this application and to implement it in accordance with the contents of the specification, and to make the above and other objects, features and advantages of this application more obvious and understandable, the following are specific embodiments of this application. Attached Figure Description

[0017] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0018] Figure 1 This is a flowchart of the face liveness detection method in Embodiments 1 and 2 of this application;

[0019] Figure 2 This is another flowchart of the face liveness detection method in Embodiments 1 and 2 of this application;

[0020] Figure 3 This is one of the schematic diagrams of the face liveness detection device in embodiments three and four of this application;

[0021] Figure 4 This is the second schematic diagram of the face liveness detection device in Embodiment 3 of this application;

[0022] Figure 5 This is the second schematic diagram of the face liveness detection device in Embodiment 4 of this application;

[0023] Figure 6 A block diagram schematically illustrates an electronic device for performing the method according to this application; and

[0024] Figure 7 A storage unit for holding or carrying program code implementing the method according to this application is illustrated schematically. Detailed Implementation

[0025] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0026] Example 1

[0027] This application discloses a method for face liveness detection, such as... Figure 1 As shown, the method includes steps 110 to 130.

[0028] Step 110: Obtain the high-frequency feature components and low-frequency feature components of each video image in the video image frame sequence used for face liveness detection.

[0029] The video image frame sequence used for face liveness detection described in this application embodiment can be a continuously acquired video image frame sequence of the face to be detected. The video image frame sequence includes at least two video images. The video images can be visible light images.

[0030] Visible light images contain both high-frequency and low-frequency information. Low-frequency information refers to slowly changing grayscale values, while high-frequency information refers to rapidly changing image frequencies, such as significant differences in grayscale between adjacent regions. For example, the general outline and contours of an image, i.e., global and coarse-structure information, constitute low-frequency information; while the edges between targets and background in an image typically show a clear difference, i.e., rapidly changing grayscale values ​​along the edges, which constitute high-frequency information.

[0031] Visible light images can be decomposed into low-frequency signals representing the global and coarse structure, and high-frequency signals representing the fine structure. Low-frequency images, composed of low-frequency signals, can be obtained by smoothing with a Gaussian filter. The high-frequency image is obtained by subtracting the low- and high-frequency components from the original visible light image. Similarly, the output feature maps of convolutional layers in convolutional neural networks can also be decomposed into high-frequency and low-frequency components. Subtracting the low-frequency features from the original features yields the high-frequency features. For deep learning, high-frequency and low-frequency images present different learning difficulties; low-frequency images are easier to learn, while high-frequency images are more challenging.

[0032] In some embodiments of this application, a pre-defined convolutional neural network can be used to perform convolution operations on each single-frame video image included in the video image frame sequence to obtain feature images of each frame video image after feature mapping. Then, a Gaussian filter is used to filter the feature images to obtain the low-frequency features of a single-frame video image. The low-frequency features of a single-frame video image constitute the low-frequency feature components of that frame video image. Further, the low-frequency feature components are subtracted from the feature images of the single-frame video image, and the remaining part is the high-frequency feature component. For ease of description, the high-frequency feature component of a single-frame video image will be denoted by the symbol "X" below. H The low-frequency feature components of a single frame video image are denoted by the symbol "X". L ".

[0033] For a specific example, when performing a regular convolution operation on a single frame of video image, a feature image X can be output. Where h and w represent the dimensions of the feature image, and c represents the number of feature images or channels. Next, a Gaussian filter t=2, where t is the variance of the Gaussian filter, is used to smooth the feature image X, resulting in the low-frequency image Xb of feature image X. L That is, the feature map composed of the low-frequency feature components of the feature image. Further, according to formula X... H =XX L That is, subtracting the low-frequency feature components from the original feature image to obtain the high-frequency feature components X of the original feature image. H .

[0034] In some embodiments of this application, the low-frequency feature component X is further... L The resolution is reduced to half of the original, and the area is reduced to a quarter of the original, which can reduce redundancy and obtain compact features.

[0035] Thus, for each video image in the video image frame sequence, after convolution and filtering, the high-frequency feature components and low-frequency feature components of each video image can be obtained.

[0036] In some embodiments of this application, convolution operations can be further performed on the high-frequency feature components and low-frequency feature components of each frame of video image to further extract and map features, thereby enhancing the low-frequency and high-frequency features.

[0037] Step 120: Obtain the inter-frame correlation of the video image frame sequence based on the high-frequency feature components and the low-frequency feature components of the video image.

[0038] For video image sequences capturing live human faces, there will be certain similarities and differences between different image frames in the video image sequence. In some embodiments of this application, these similarities and differences between different image frames in the video image sequence are collectively referred to as inter-frame correlation. That is, the inter-frame correlation includes inter-frame consistency and inter-frame difference.

[0039] In some embodiments of this application, the inter-frame correlation of the video image frame sequence can be determined based on the changes in details and the consistency of outlines and backgrounds between adjacent video images in the video image sequence.

[0040] The video image frame sequence includes at least two video images. As mentioned earlier, high-frequency components in visible light images represent detail information, while low-frequency components represent overall appearance and contour information. For face images, distinguishing between live and fake faces mainly relies on changes in details, such as expressions and subtle facial movements. Therefore, the frame-to-frame changes in high-frequency feature components in video images are key features to be extracted, while changes in low-frequency overall appearance and contour include large changes in the angle of the face position and are features to be ignored. In some embodiments of this application, obtaining the inter-frame correlation of the video image frame sequence based on the high-frequency and low-frequency feature components of the video images includes sub-steps S11 and S12.

[0041] Sub-step S11: Based on the high-frequency feature components of at least two adjacent video images, obtain the difference between the high-frequency feature components corresponding to each corresponding image position in the at least two adjacent video images to form the high-frequency frame difference feature of the video image frame sequence; and based on the low-frequency feature components of the at least two adjacent video images, obtain the average value of the low-frequency feature components corresponding to each corresponding image position in the at least two adjacent video images to form the low-frequency equalization feature of the video image frame sequence.

[0042] For example, for any two adjacent video frames in a video image sequence (denoted as y1 and y2), the difference between the high-frequency feature components at their corresponding image positions reflects the subtle changes between these two video frames. Therefore, the difference is a difference feature, which can be used to express the inter-frame differences between these two video frames, as part of the information on inter-frame correlation. Taking a face video image sequence as an example, the subtle differences between two video frames are the changes in the facial expression details of the same person in different image frames, which are reflected in the high-frequency differences. In some embodiments of this application, this can be achieved through formula X. H =X y1 H –X y2 H The difference X is obtained by subtracting the high-frequency feature components at the corresponding image positions of adjacent video images y1 and y2. H It reflects the subtle changes between these two video images and can be used as a high-frequency frame difference feature of the video image frame sequence.

[0043] On the other hand, for video images y1 and y2, the low-frequency feature components at their corresponding image positions are averaged. The resulting average reflects the general outline and contour of these two video images. Therefore, the average is a consistency feature and can be used to express the inter-frame consistency between these two video images, serving as partial information about inter-frame correlation. Taking a face video image sequence as an example, the general outline and contour information of two video images, such as the facial contour of the same person in different image frames, are reflected in the low-frequency consistency. In some embodiments of this application, this can be achieved through formula X. L =(X y1 L +X y2 L The average value X is obtained by averaging the low-frequency feature components at the corresponding image positions of adjacent video images y1 and y2. L It reflects the unchanging general outline and contour in these two video images, and can be used as the low-frequency equalization feature of the video image frame sequence.

[0044] Sub-step S12 involves weighted concatenation of the high-frequency frame difference feature and the low-frequency equalization feature to obtain the inter-frame correlation of the video image frame sequence.

[0045] Furthermore, it can be done through formulas The high-frequency frame difference feature and the low-frequency equalization feature are weighted and concatenated to obtain the concatenated vector V. p splicing vector V p This is used to characterize the inter-frame correlation of the video image frame sequence. Where w1 is the high-frequency frame difference feature X. H The weights, w2 is the low-frequency frame difference feature X LThe weights w1 and w2 are used to adjust the ratio of high-frequency feature components to low-frequency feature components. The values ​​of w1 and w2 are between 0 and 1, and are determined based on the accuracy of the face liveness test results. In some embodiments of this application, because high-frequency information in face images can better highlight the subtle changes between frames—for example, a live face will produce subtle facial expressions with varying light, resulting in inter-frame differences between video images—while the subtle changes in the attack sample are smaller, in order to make the high-frequency features stand out more, the high-frequency features that highlight subtle changes need to be amplified, i.e., the value of w1 is greater than the value of w2. In some embodiments of this application, the value of w1 can be set to 0.75, and the value of w2 can be set to 0.25.

[0046] Step 130: Based on the matching relationship between the inter-frame correlation and the pre-acquired inter-frame correlation of live faces, detect whether the face to be detected in the video image is a live face.

[0047] After processing the high-frequency and low-frequency feature components of two or more adjacent video frames in the video image sequence according to a preset method to obtain the inter-frame correlation characterizing the video image sequence, the obtained inter-frame correlation is further matched with the pre-acquired inter-frame correlation of live faces, and the detection face collected in the video image frame sequence is determined as a live face based on the matching result.

[0048] In some embodiments of this application, the pre-acquired inter-frame correlation of the live face is: a surface that characterizes the distribution law of inter-frame correlation of high-frequency and low-frequency features in the live face video image sequence, obtained by fitting live face samples.

[0049] In some embodiments of this application, such as Figure 2 As shown, before detecting whether the face to be detected in the video image is a live face based on the matching relationship between the inter-frame correlation and the pre-acquired inter-frame correlation of live faces, the method further includes: step 100.

[0050] Step 100: Obtain the inter-frame correlation of a live face.

[0051] In some embodiments of this application, obtaining the inter-frame correlation of a live face includes: constructing a sample set comprising a plurality of positive samples using a video image sequence of a live face as positive samples; obtaining the high-frequency feature components and low-frequency feature components of each frame of video image in each positive sample in the sample set; for each positive sample, obtaining the inter-frame correlation characterizing the positive sample based on the high-frequency feature components and low-frequency feature components of at least two frames of video image in the positive sample; and performing surface fitting on the inter-frame correlation of the positive samples in the sample set to obtain a surface characterizing the distribution law of the inter-frame correlation of high-frequency and low-frequency features in the live face video image sequence.

[0052] Because the spoofed faces used in attacks are diverse, varied, and constantly evolving, and because obtaining spoofed face samples is difficult, costly, and involves limited data, traditional binary classification methods suffer from sample imbalance and significant errors. In the embodiments of this application, modeling is performed only on live face samples (i.e., positive samples) to obtain a high-dimensional feature surface showing the distribution of inter-frame correlations of live faces.

[0053] For example, we can first acquire several video image sequences of live faces, each video image sequence including at least two video images, and construct a positive sample based on each video image sequence, thus obtaining a positive sample set including several positive samples. Next, for each positive sample, we use the same method as in the previous step of acquiring the inter-frame correlation of the video image sequence to be detected to obtain the inter-frame correlation of each training sample. As can be seen from the previous steps, the inter-frame correlation of the positive sample is obtained by weighted concatenation of the high-frequency frame difference features and low-frequency equalization features of adjacent video images in the positive sample, and is a multi-dimensional vector. Next, based on the inter-frame correlation of each positive sample in the positive sample set, we fit a high-dimensional feature distribution surface of the inter-frame correlation, which expresses the distribution law of high-frequency feature components and low-frequency feature components in the live face sequence.

[0054] For a specific implementation of fitting a high-dimensional feature distribution surface of inter-frame correlation based on the inter-frame correlation of each positive sample in the positive sample set, please refer to the implementation of fitting a surface based on several multi-dimensional vectors in the prior art. It will not be repeated in the embodiments of this application.

[0055] In some embodiments of this application, after performing surface fitting on the inter-frame correlation of the positive samples in the sample set to obtain a surface characterizing the distribution law of inter-frame correlation of high-frequency and low-frequency features in the live face video image sequence, the method further includes: substituting the inter-frame correlation of the positive samples in the sample set into the surface to obtain the distance between the inter-frame correlation of each positive sample and the surface; and determining a distance threshold between the inter-frame correlation of the live face video image sequence and the surface based on the distribution of the distance between the inter-frame correlation of each positive sample and the surface. For example, after fitting the surface characterizing the distribution law of inter-frame correlation of high-frequency and low-frequency features in the live face video image sequence, the inter-frame correlation of each positive sample in the positive sample set is substituted into a preset surface model. The surface model is used to calculate the distance between the inter-frame correlation of each positive sample and the surface. For example, the surface model can be expressed as: |f(V t )-z|<ε, where ε is a very small constant. Then, by adjusting the value of ε and statistically analyzing the cases where the inter-frame correlation of each positive sample in the positive sample set satisfies the above surface model with respect to the distance between the surface and the positive sample, except for a few abnormal positive samples, all other positive samples in the positive sample set satisfy the surface model determined by the adjusted ε. The value of ε at this time is set as the distance threshold between the inter-frame correlation of the live face video image sequence and the fitted surface.

[0056] In some embodiments of this application, the distance threshold may also be determined empirically.

[0057] In some embodiments of this application, the step of detecting whether the face to be detected in the video image acquisition is a live face based on the matching relationship between the inter-frame correlation and the pre-acquired inter-frame correlation of a live face includes: substituting the acquired inter-frame correlation into the surface to determine whether the acquired inter-frame correlation conforms to the inter-frame correlation distribution law of high-frequency and low-frequency features in the live face video image sequence expressed by the surface; in response to the acquired inter-frame correlation conforming to the inter-frame correlation distribution law of high-frequency and low-frequency features in the live face video image sequence expressed by the surface, determining that the face to be detected in the video image acquisition is a live face; in response to the acquired inter-frame correlation not conforming to the inter-frame correlation distribution law of high-frequency and low-frequency features in the live face video image sequence expressed by the surface, determining that the face to be detected in the video image acquisition is a non-live face.

[0058] In the liveness detection stage, video image sequences that do not conform to the inter-frame correlation rules represented by the surface obtained from the modeling can be directly identified as non-live faces, thus improving the generalization ability of liveness detection.

[0059] In some embodiments of this application, the step of substituting the acquired inter-frame correlation into the surface to determine whether the acquired inter-frame correlation conforms to the inter-frame correlation distribution law of high-frequency and low-frequency features in the live face video image sequence expressed by the surface includes: substituting the acquired inter-frame correlation into the surface to obtain the distance between the inter-frame correlation and the surface; in response to the distance being less than a predetermined distance threshold, determining that the inter-frame correlation conforms to the inter-frame correlation distribution law of high-frequency and low-frequency features in the live face video image sequence expressed by the surface; in response to the distance being greater than or equal to the predetermined distance threshold, determining that the inter-frame correlation does not conform to the inter-frame correlation distribution law of high-frequency and low-frequency features in the live face video image sequence expressed by the surface.

[0060] Taking the pre-acquired inter-frame correlation of a live face as an example, represented by the surface f(V) = z, where z is a constant. During face liveness detection, the inter-frame correlation V, calculated in the aforementioned steps based on the high-frequency and low-frequency feature components of two adjacent video frames in the current video image frame sequence, is used. p Substitute the surface model |f(V) t If |f(V)| < ε, then |f(V) p If |f(V)| < ε holds, it indicates that the currently acquired inter-frame correlation conforms to the inter-frame correlation distribution law of high-frequency and low-frequency features in the live face video image sequence expressed by the surface; if |f(V)| < ε holds, it means that the inter-frame correlation obtained conforms to the inter-frame correlation distribution law of high-frequency and low-frequency features in the live face video image sequence expressed by the surface. p If -z|<ε is not true, it means that the currently obtained inter-frame correlation does not conform to the inter-frame correlation distribution law of high-frequency and low-frequency features in the live face video image sequence expressed by the surface; otherwise...

[0061] Furthermore, if, after substituting the acquired inter-frame correlation into the surface, it is determined that the inter-frame correlation conforms to the inter-frame correlation distribution pattern of high-frequency and low-frequency features in the live face video image sequence expressed by the surface, then the face to be detected in the video image frame sequence is considered a live face; if, after substituting the acquired inter-frame correlation into the surface, it is determined that the inter-frame correlation does not conform to the inter-frame correlation distribution pattern of high-frequency and low-frequency features in the live face video image sequence expressed by the surface, then the face to be detected in the video image frame sequence is considered a non-live face.

[0062] The face liveness detection method disclosed in this application obtains the high-frequency and low-frequency feature components of each frame of a video image in a video image frame sequence used for face liveness detection; obtains the inter-frame correlation of the video image frame sequence based on the high-frequency and low-frequency feature components of the video image; and detects whether the face to be detected in the video image is a live face based on the matching relationship between the inter-frame correlation and the pre-obtained inter-frame correlation of live faces, which helps to improve the accuracy of face liveness detection.

[0063] The face liveness detection method disclosed in this application determines whether a video image sequence contains live faces based on the inter-frame detail variation features of video images in a video image sequence, as well as the matching results of the distribution patterns of the invariant features of the overall shape and contour with the aforementioned features in live face images. Compared to simply comparing the distinguishing features between two frames, this method can improve the accuracy of live face detection. Furthermore, this method uses offline fitting of the inter-frame correlation of live faces as a comparison reference in the detection stage. In the detection stage, only simple convolution operations are needed to obtain the inter-frame correlation of the current video image. Compared to face liveness detection based on depth features, this method has a lower computational load and is more suitable for front-end devices.

[0064] Example 2

[0065] This application discloses a method for face liveness detection, such as... Figure 1 As shown, the method includes steps 110 to 130.

[0066] Step 110: Obtain the high-frequency feature components and low-frequency feature components of each video image in the video image frame sequence used for face liveness detection.

[0067] Step 120: Obtain the inter-frame correlation of the video image frame sequence based on the high-frequency feature components and the low-frequency feature components of the video image.

[0068] Step 130: Based on the matching relationship between the inter-frame correlation and the pre-acquired inter-frame correlation of live faces, detect whether the face to be detected in the video image is a live face.

[0069] The specific implementation method for obtaining the high-frequency feature components and low-frequency feature components of each video image frame in the video image frame sequence used for face liveness detection is described in Example 1, and will not be repeated in this example.

[0070] Unlike Embodiment 1, in some embodiments of this application, the video image frame sequence includes multiple video images, and obtaining the inter-frame correlation of the video image frame sequence based on the high-frequency feature components and the low-frequency feature components of the video images includes: sub-steps S21 to S23.

[0071] Sub-step S21: The high-frequency feature components of multiple specified video images in the video image frame sequence are concatenated into a multi-dimensional high-frequency feature component matrix in the form of one-dimensional high-frequency feature components corresponding to each frame of the video image; and the low-frequency feature components of the multiple specified video images are concatenated into a multi-dimensional low-frequency feature component matrix in the form of one-dimensional low-frequency feature components corresponding to each frame of the video image.

[0072] For example, firstly, for each video frame with a high-frequency feature component of size h*w, the high-frequency feature components corresponding to each column of pixels in each video frame can be sequentially concatenated end to end to obtain a one-dimensional high-frequency feature component of length h*w; then, using the one-dimensional high-frequency feature component corresponding to each video frame as row or column elements, the one-dimensional high-frequency feature components corresponding to N video frames in the video frame sequence are sequentially concatenated into a multi-dimensional high-frequency feature component matrix including N rows or N columns, according to the chronological order of the video frame timestamps.

[0073] Similarly, for a low-frequency feature component of size h*w for each video frame, the low-frequency feature components corresponding to each column of pixels in each video frame can be concatenated end to end to obtain a one-dimensional low-frequency feature component of length h*w. Then, using the one-dimensional low-frequency feature component corresponding to each video frame as row or column elements, the one-dimensional low-frequency feature components corresponding to N video frames in the video frame sequence are sequentially concatenated into a multi-dimensional low-frequency feature component matrix including N rows or N columns, according to the order of the video frame timestamps.

[0074] In this way, each frame of video image will obtain a (h*w)*N multidimensional high-frequency feature component matrix and a (h*w)*N multidimensional low-frequency feature component matrix.

[0075] Sub-step S22 involves performing low-rank sparse matrix decomposition on the multidimensional high-frequency feature component matrix and the multidimensional low-frequency feature component matrix to obtain a first matrix representing low-frequency background features, a second matrix representing low-frequency detail features, a third matrix representing high-frequency background features, and a fourth matrix representing high-frequency detail features.

[0076] Next, the multidimensional high-frequency feature component matrix (e.g., denoted as "A") is obtained based on the high-frequency and low-frequency feature components of the video images in the current video image frame sequence. H ") and multidimensional low-frequency feature component matrix (e.g., denoted as "A") ...L After that, low-rank sparse matrix decomposition was performed using the RPCA (Robust Principal Component Analysis) method to obtain the first matrix (e.g., denoted as "L") used to characterize the low-frequency background features. L The second matrix used to characterize low-frequency detail features (e.g., denoted as "S") L The third matrix used to characterize high-frequency background features (e.g., denoted as "L") H ), and a fourth matrix (e.g., denoted as "S") used to characterize high-frequency detail features. H ”).

[0077] For specific implementation methods of low-rank sparse matrix decomposition of multidimensional high-frequency feature component matrices and low-rank sparse matrix decomposition of multidimensional low-frequency feature component matrices, please refer to the low-rank sparse matrix decomposition methods in the prior art, which will not be repeated in the embodiments of this application.

[0078] Sub-step S23: The inter-frame consistency is represented by the nuclear norm of the first matrix, and the inter-frame differences are represented by the column and norm of the fourth matrix, thereby obtaining the inter-frame correlation of the video image frame sequence.

[0079] After low-rank sparse matrix decomposition, the low-rank part of the multidimensional low-frequency feature component matrix is ​​decomposed into a first matrix representing low-frequency background features, and the kernel norm of the matrix is ​​suitable for expressing the changes in the low-rank part. Therefore, the kernel norm of the first matrix can be used as a representation of background feature changes; the smaller the kernel norm, the smaller the background feature changes, i.e., the greater the consistency. After low-rank sparse matrix decomposition, the sparse part of the multidimensional high-frequency feature component matrix is ​​decomposed into a fourth matrix representing high-frequency detail features, and the column norm and norm of the matrix are suitable for expressing the changes in the sparse part. The column norm, i.e., the 1-norm, is the maximum value of the sum of the absolute values ​​of all matrix column vectors. Therefore, the column norm and norm of the fourth matrix can be used as a representation of detail feature changes; the larger the column norm and norm, the greater the detail feature changes, i.e., the greater the difference.

[0080] After determining the inter-frame correlation of the currently acquired video image sequence of the face to be detected, the next step is to match the acquired inter-frame correlation with the pre-acquired inter-frame correlation rules for live faces to obtain the face liveness detection results.

[0081] Unlike Embodiment 1, in this embodiment, the pre-acquired inter-frame correlation of the live face is: a column and norm threshold characterizing the distribution pattern of high-frequency feature inter-frame correlation in the live face video image sequence, and a kernel norm threshold characterizing the distribution pattern of low-frequency feature inter-frame correlation in the live face video image sequence.

[0082] Accordingly, the step of detecting whether the face to be detected in the video image is a live face based on the matching relationship between the inter-frame correlation and the pre-acquired inter-frame correlation of live faces includes: comparing the kernel norm of the first matrix with the kernel norm threshold, and comparing the column and norm of the fourth matrix with the column and norm threshold; in response to the kernel norm of the first matrix being less than the kernel norm threshold, and the column and norm of the fourth matrix being greater than or equal to the column and norm threshold, detecting the face to be detected in the video image as a live face; in response to the kernel norm of the first matrix being greater than or equal to the kernel norm threshold, or the column and norm of the fourth matrix being less than the column and norm threshold, detecting the face to be detected in the video image as a non-live face.

[0083] For example, the kernel norm of the first matrix, which is used as the inter-frame correlation of the currently acquired video image frame sequence, is compared with a pre-acquired kernel norm threshold. If the kernel norm of the first matrix is ​​less than the kernel norm threshold, it indicates that the rank of the low-rank matrix is ​​small, and the low-frequency features (such as general outline and contour information) of each video image in the video sequence that generates the low-rank matrix are highly correlated. It can be assumed that the background of each video image in the video sequence has not been switched. As another example, the column and norm of the fourth matrix, which is used as the inter-frame correlation of the currently acquired video image frame sequence, is compared with a pre-acquired column and norm threshold. If the column and norm of the fourth matrix is ​​greater than or equal to the column and norm threshold, it indicates that the sparsity of the sparse matrix is ​​high, and the high-frequency features (such as detail information) of each video image in the video sequence that generates the sparse matrix change significantly. It can be assumed that the faces in each video image in the video sequence are changing.

[0084] If the comparison results simultaneously satisfy the condition that the nuclear norm of the first matrix is ​​less than the nuclear norm threshold, and the column norm of the fourth matrix is ​​greater than or equal to the column norm threshold, then the face to be detected in the video image frame sequence can be considered a live face; otherwise, the face to be detected in the video image frame sequence can be considered a non-live face.

[0085] Determining whether changes in a video image sequence are detail changes or background changes based on the physical properties of low-rank sparse subspaces reveals several key aspects. In the case of a live face, details change but the overall background outline remains unchanged. In the case of an attack, however, the background or outline changes significantly, while details remain unchanged or change very little. Therefore, using the decomposition results of the low-rank sparse matrix to determine whether a video image frame sequence represents a live face or an attack face can reduce the difficulty of sample collection and improve generalization ability.

[0086] like Figure 2As shown, before detecting whether the face to be detected in the video image is a live face based on the matching relationship between the inter-frame correlation and the pre-acquired inter-frame correlation of live faces, the method further includes: step 100.

[0087] Step 100: Obtain the inter-frame correlation of a live face.

[0088] Unlike Embodiment 1, in some embodiments of this application, the column norm threshold and nuclear norm threshold can be determined in advance based on experience.

[0089] In other embodiments of this application, the step of obtaining the inter-frame correlation of a live face includes: constructing a sample set including a plurality of positive samples using a video image sequence of a live face as positive samples; obtaining the high-frequency feature components and low-frequency feature components of each frame of video image in each positive sample in the sample set; and for each positive sample, performing the following inter-frame correlation acquisition operation: concatenating the high-frequency feature components of multiple specified video images in the positive sample into a multi-dimensional high-frequency feature component matrix in the form of a one-dimensional high-frequency feature component corresponding to each frame of video image in the positive sample; and concatenating the high-frequency feature components of multiple specified video images in the positive sample into a multi-dimensional high-frequency feature component matrix in the form of a one-dimensional high-frequency feature component corresponding to each frame of video image in the positive sample; and concatenating the high-frequency feature components of multiple specified video images in the positive sample into a multi-dimensional high-frequency feature component matrix in the form of a one-dimensional low-frequency feature component corresponding to each frame of video image in the positive sample. In a quantitative form, the low-frequency feature components of the specified multiple frames of video images are concatenated into a multi-dimensional low-frequency feature component matrix; low-rank sparse matrix decomposition is performed on the multi-dimensional high-frequency feature component matrix and the multi-dimensional low-frequency feature component matrix to obtain a first matrix representing low-frequency background features and a fourth matrix representing high-frequency detail features; the kernel norm of the first matrix and the column norm of the fourth matrix are obtained as the inter-frame correlation of the corresponding positive samples; based on the inter-frame correlation of all positive samples in the sample set, the kernel norm threshold of the first matrix used to represent low-frequency background features and the column norm threshold of the fourth matrix used to represent high-frequency detail features are determined.

[0090] For a specific implementation method of constructing a sample set including several positive samples using video image sequences of live human faces as positive samples, please refer to Embodiment 1, which will not be repeated in this embodiment.

[0091] For a detailed implementation of obtaining the high-frequency and low-frequency feature components of each frame of video image in each positive sample in the sample set, please refer to Embodiment 1, which will not be repeated in this embodiment.

[0092] For a detailed implementation of obtaining the inter-frame correlation of each positive sample, please refer to the description of the step of obtaining the inter-frame correlation of the current video image frame sequence in this embodiment, which will not be repeated here.

[0093] After obtaining the inter-frame correlation of all positive samples in the positive sample set (i.e., the nuclear norm of the corresponding first matrix and the column norm of the fourth matrix), the nuclear norm of the first matrix and the column norm of the fourth matrix are further statistically analyzed, and the nuclear norm of the first matrix that is satisfied by all positive samples in the positive sample set is determined as the nuclear norm threshold of the first matrix. In addition, the column norm of the fourth matrix that is satisfied by all positive samples in the positive sample set is determined as the column norm threshold of the fourth matrix.

[0094] The face liveness detection method disclosed in this application obtains the high-frequency and low-frequency feature components of each frame of a video image in a video image frame sequence used for face liveness detection; obtains the inter-frame correlation of the video image frame sequence based on the high-frequency and low-frequency feature components of the video image; and detects whether the face to be detected in the video image is a live face based on the matching relationship between the inter-frame correlation and the pre-obtained inter-frame correlation of live faces, which helps to improve the accuracy of face liveness detection.

[0095] The face liveness detection method disclosed in this application determines whether a video image sequence contains live faces based on the inter-frame detail variation features of video images in a video image sequence, as well as the matching results of the distribution patterns of the invariant features of the overall shape and contour with the aforementioned features in live face images. Compared to simply comparing the distinguishing features between two frames, this method can improve the accuracy of live face detection. Furthermore, this method uses offline statistical analysis of the inter-frame correlation of live faces as a comparison reference in the detection stage. In the detection stage, only simple convolution operations are needed to obtain the inter-frame correlation of the current video image. Compared to face liveness detection based on depth features, this method has a lower computational load and is more suitable for front-end devices.

[0096] Example 3

[0097] This application discloses a face liveness detection device, such as... Figure 3 As shown, the device includes:

[0098] The high- and low-frequency component acquisition module 310 is used to acquire the high-frequency feature components and low-frequency feature components of each video image in the video image frame sequence used for face liveness detection.

[0099] The inter-frame correlation acquisition module 320 acquires the inter-frame correlation of the video image frame sequence based on the high-frequency feature components and the low-frequency feature components of the video image, wherein the inter-frame correlation includes inter-frame consistency and inter-frame difference.

[0100] The liveness detection module 330 is used to detect whether the face to be detected in the video image is a live face based on the inter-frame correlation and the matching relationship between the inter-frame correlation and the pre-acquired inter-frame correlation of live faces.

[0101] In some embodiments of this application, the video image frame sequence includes at least two video images, such as... Figure 4 As shown, the inter-frame correlation acquisition module 320 further includes: a first inter-frame correlation acquisition submodule 3201.

[0102] The first inter-frame correlation acquisition submodule 3201 is configured to: acquire, based on the high-frequency feature components of at least two adjacent video images, the difference between the high-frequency feature components corresponding to each corresponding image position in the at least two adjacent video images, thereby constituting a high-frequency frame difference feature of the video image frame sequence; acquire, based on the low-frequency feature components of the at least two adjacent video images, the average value of the low-frequency feature components corresponding to each corresponding image position in the at least two adjacent video images, thereby constituting a low-frequency equalization feature of the video image frame sequence; and perform weighted concatenation of the high-frequency frame difference feature and the low-frequency equalization feature to acquire the inter-frame correlation of the video image frame sequence.

[0103] In some embodiments of this application, the pre-acquired inter-frame correlation of the live face is: a surface that characterizes the distribution law of inter-frame correlation of high-frequency and low-frequency features in the live face video image sequence, obtained by fitting live face samples; such as Figure 4 As shown, the liveness detection module 330 further includes:

[0104] The first liveness detection submodule 3301 is used to substitute the acquired inter-frame correlation into the surface to determine whether the acquired inter-frame correlation conforms to the inter-frame correlation distribution law of high-frequency features and low-frequency features in the live face video image sequence expressed by the surface.

[0105] The first liveness detection submodule 3301 is further configured to, in response to the acquisition of the inter-frame correlation conforming to the inter-frame correlation distribution law of high-frequency and low-frequency features in the live face video image sequence expressed by the surface, determine that the face to be detected in the video image acquisition is a live face; or,

[0106] In response to the fact that the obtained inter-frame correlation does not conform to the inter-frame correlation distribution pattern of high-frequency and low-frequency features in the live face video image sequence expressed by the surface, the face to be detected in the video image frame is determined to be a non-live face.

[0107] In some embodiments of this application, before detecting whether the face to be detected in the video image is a live face based on the matching relationship between the inter-frame correlation and the pre-acquired inter-frame correlation of live faces, such as... Figure 4 As shown, the liveness detection module 330 further includes: a first inter-frame correlation distribution pattern acquisition submodule 3302.

[0108] The first inter-frame correlation distribution pattern acquisition submodule 3302 is used to construct a sample set including a plurality of said positive samples, using the video image sequence of a live human face as positive samples; and,

[0109] Obtain the high-frequency feature components and low-frequency feature components of each frame of video image in each positive sample in the sample set;

[0110] The first inter-frame correlation distribution pattern acquisition submodule 3302 is further configured to, for each positive sample, acquire the inter-frame correlation characterizing the positive sample based on the high-frequency feature components and the low-frequency feature components of at least two frames of the video images in the positive sample; and,

[0111] A surface is fitted to the inter-frame correlation of the positive samples in the sample set to obtain a surface that characterizes the distribution pattern of inter-frame correlation of high-frequency and low-frequency features in the live face video image sequence.

[0112] In some embodiments of this application, after performing surface fitting on the inter-frame correlation of the positive samples in the sample set to obtain a surface characterizing the distribution law of inter-frame correlation of high-frequency and low-frequency features in the live face video image sequence, as follows: Figure 4 As shown, the liveness detection module 330 further includes: a distance threshold determination submodule 3303.

[0113] The distance threshold determination submodule 3303 is used to substitute the inter-frame correlation of the positive samples in the sample set into the surface to obtain the distance between the inter-frame correlation of each positive sample and the surface; and

[0114] Based on the distribution of the distance between the inter-frame correlation and the surface of each positive sample, a distance threshold between the inter-frame correlation and the surface of the live face video image sequence is determined.

[0115] The step of substituting the acquired inter-frame correlation into the surface to determine whether the acquired inter-frame correlation conforms to the inter-frame correlation distribution pattern of high-frequency and low-frequency features in the live face video image sequence expressed by the surface includes:

[0116] Substitute the obtained inter-frame correlation into the surface to obtain the distance between the inter-frame correlation and the surface;

[0117] In response to the distance being less than a predetermined distance threshold, it is determined that the inter-frame correlation conforms to the inter-frame correlation distribution pattern of high-frequency and low-frequency features in the live face video image sequence expressed by the surface.

[0118] In response to the distance being greater than or equal to the predetermined distance threshold, it is determined that the inter-frame correlation does not conform to the inter-frame correlation distribution pattern of high-frequency and low-frequency features in the live face video image sequence expressed by the surface.

[0119] The face liveness detection apparatus disclosed in this application is used to implement the face liveness detection method described in Embodiment 1 of this application. The specific implementation methods of each module of the apparatus will not be repeated here, but can be found in the specific implementation methods of the corresponding steps in the method embodiment.

[0120] The face liveness detection apparatus disclosed in this application acquires high-frequency and low-frequency feature components of each frame of a video image in a video image frame sequence used for face liveness detection; obtains the inter-frame correlation of the video image frame sequence based on the high-frequency and low-frequency feature components of the video image; and detects whether the face to be detected in the video image is a live face based on the matching relationship between the inter-frame correlation and the pre-acquired inter-frame correlation of live faces, which helps to improve the accuracy of face liveness detection.

[0121] The face liveness detection apparatus disclosed in this application determines whether a video image sequence contains live faces based on the matching results of the inter-frame detail change features of video images in a video image sequence, as well as the invariance features of the overall shape and contour, and the distribution patterns of the aforementioned features in live face images. Compared to simply comparing the distinguishing features between two frames, this method can improve the accuracy of live face detection. Furthermore, this method uses offline fitting of the inter-frame correlation of live faces as a comparison reference in the detection stage. In the detection stage, only simple convolution operations are needed to obtain the inter-frame correlation of the current video image. Compared to face liveness detection based on depth features, this method has a lower computational load and is more suitable for front-end devices.

[0122] Example 4

[0123] This application discloses a face liveness detection device, such as... Figure 3 As shown, the device includes:

[0124] The high- and low-frequency component acquisition module 310 is used to acquire the high-frequency feature components and low-frequency feature components of each video image in the video image frame sequence used for face liveness detection.

[0125] The inter-frame correlation acquisition module 320 acquires the inter-frame correlation of the video image frame sequence based on the high-frequency feature components and the low-frequency feature components of the video image, wherein the inter-frame correlation includes inter-frame consistency and inter-frame difference.

[0126] The liveness detection module 330 is used to detect whether the face to be detected in the video image is a live face based on the inter-frame correlation and the matching relationship between the inter-frame correlation and the pre-acquired inter-frame correlation of live faces.

[0127] In some embodiments of this application, such as Figure 5 As shown, the inter-frame correlation acquisition module 320 further includes: a second inter-frame correlation acquisition submodule 3202;

[0128] The second inter-frame correlation acquisition submodule 3202 is used to concatenate the high-frequency feature components of multiple specified video images in the video image frame sequence into a multi-dimensional high-frequency feature component matrix in the form of one-dimensional high-frequency feature components corresponding to each frame of the video image; and to concatenate the low-frequency feature components of multiple specified video images into a multi-dimensional low-frequency feature component matrix in the form of one-dimensional low-frequency feature components corresponding to each frame of the video image.

[0129] The second inter-frame correlation acquisition submodule 3202 is further configured to perform low-rank sparse matrix decomposition on the multi-dimensional high-frequency feature component matrix and the multi-dimensional low-frequency feature component matrix respectively to obtain a first matrix representing low-frequency background features, a second matrix representing low-frequency detail features, a third matrix representing high-frequency background features, and a fourth matrix representing high-frequency detail features; and to obtain the inter-frame correlation of the video image frame sequence by representing inter-frame consistency through the kernel norm of the first matrix and representing inter-frame differences through the column and norm of the fourth matrix.

[0130] In some embodiments of this application, the pre-acquired inter-frame correlation of live faces comprises: column and norm thresholds characterizing the distribution pattern of high-frequency feature inter-frame correlation in the live face video image sequence, and kernel norm thresholds characterizing the distribution pattern of low-frequency feature inter-frame correlation in the live face video image sequence; the liveness detection module 330 further comprises: a second liveness detection submodule 3305.

[0131] The second liveness detection submodule 3305 is used to compare the nuclear norm of the first matrix with the nuclear norm threshold, and to compare the column norm of the fourth matrix with the column norm threshold;

[0132] The second liveness detection submodule 3305 is further configured to, in response to the fact that the nuclear norm of the first matrix is ​​less than the nuclear norm threshold, and the column and sum norms of the fourth matrix are greater than or equal to the column and sum norm thresholds, determine that the face to be detected in the video image frame sequence is a live face; or,

[0133] In response to the nuclear norm of the first matrix being greater than or equal to the nuclear norm threshold, or the column norm of the fourth matrix being less than the column norm threshold, the face to be detected in the video image frame sequence is determined to be a non-living face.

[0134] In some embodiments of this application, before detecting whether the face to be detected in the video image is a live face based on the matching relationship between the inter-frame correlation and the pre-acquired inter-frame correlation of live faces, such as... Figure 5 As shown, the liveness detection module 330 further includes: a nuclear norm threshold and column norm threshold determination submodule 3306;

[0135] The nuclear norm threshold and column norm threshold determination submodule 3306 is used to construct a sample set including a number of positive samples using a video image sequence of a live human face as a positive sample; and to obtain the high-frequency feature components and low-frequency feature components of each frame of video image in each positive sample in the sample set.

[0136] The nuclear norm threshold and column norm threshold determination submodule 3306 is further configured to perform the following inter-frame correlation acquisition operation for each positive sample:

[0137] The high-frequency feature components of multiple specified video images in the positive sample are concatenated into a multi-dimensional high-frequency feature component matrix in the form of one-dimensional high-frequency feature components corresponding to each frame of the video image in the positive sample; and the low-frequency feature components of multiple specified video images are concatenated into a multi-dimensional low-frequency feature component matrix in the form of one-dimensional low-frequency feature components corresponding to each frame of the video image.

[0138] The multidimensional high-frequency feature component matrix and the multidimensional low-frequency feature component matrix are respectively subjected to low-rank sparse matrix decomposition to obtain a first matrix representing low-frequency background features and a fourth matrix representing high-frequency detail features.

[0139] Obtain the nuclear norm of the first matrix and the column norm of the fourth matrix as the inter-frame correlation of the corresponding positive samples; and,

[0140] Based on the inter-frame correlation of all the positive samples in the sample set, the kernel norm threshold of the first matrix used to characterize low-frequency background features and the column and norm thresholds of the fourth matrix used to characterize high-frequency detail features are determined.

[0141] The face liveness detection apparatus disclosed in this application is used to implement the face liveness detection method described in Embodiment 2 of this application. The specific implementation methods of each module of the apparatus will not be repeated here, but can be found in the specific implementation methods of the corresponding steps in the method embodiment.

[0142] The face liveness detection apparatus disclosed in this application acquires high-frequency and low-frequency feature components of each frame of a video image in a video image frame sequence used for face liveness detection; obtains the inter-frame correlation of the video image frame sequence based on the high-frequency and low-frequency feature components of the video image; and detects whether the face to be detected in the video image is a live face based on the matching relationship between the inter-frame correlation and the pre-acquired inter-frame correlation of live faces, which helps to improve the accuracy of face liveness detection.

[0143] The face liveness detection apparatus disclosed in this application determines whether a video image sequence contains live faces based on the matching results of the inter-frame detail change features of video images in a video image sequence, as well as the invariance features of the overall shape and contour, and the distribution patterns of the aforementioned features in live face images. Compared to simply comparing the distinguishing features between two frames, this method can improve the accuracy of live face detection. Furthermore, this method uses offline statistical analysis of the inter-frame correlation of live faces as a comparison reference in the detection stage. In the detection stage, only simple convolution operations are needed to obtain the inter-frame correlation of the current video image. Compared to face liveness detection based on depth features, this method has a lower computational load and is more suitable for front-end devices.

[0144] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus embodiments, since they are fundamentally similar to the method embodiments, the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments.

[0145] The present application provides a detailed description of a method and apparatus for face liveness detection. Specific examples have been used to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only for the purpose of helping to understand the method and its core idea. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of the present application. Therefore, the content of this specification should not be construed as a limitation of the present application.

[0146] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0147] The various component embodiments of this application can be implemented in hardware, or as software modules running on one or more processors, or a combination thereof. Those skilled in the art will understand that microprocessors or digital signal processors (DSPs) can be used in practice to implement some or all of the functions of some or all of the components in the electronic device according to the embodiments of this application. This application can also be implemented as a device or apparatus program (e.g., a computer program and computer program product) for performing part or all of the methods described herein. Such a program implementing this application can be stored on a computer-readable medium, or can be in the form of one or more signals. Such signals can be downloaded from an Internet website, provided on a carrier signal, or provided in any other form.

[0148] For example, Figure 6 An electronic device is shown that can implement the methods according to this application. The electronic device may be a PC, mobile terminal, personal digital assistant, tablet computer, etc. The electronic device conventionally includes a processor 610 and a memory 620, and program code 630 stored on the memory 620 and executable on the processor 610, which, when executing the program code 630, implements the methods described in the above embodiments. The memory 620 may be a computer program product or a computer-readable medium. The memory 620 may be an electronic memory such as flash memory, EEPROM (Electrically Erasable Programmable Read-Only Memory), EPROM, hard disk, or ROM. The memory 620 has a storage space 6201 for the program code 630 of a computer program for performing any of the method steps described above. For example, the storage space 6201 for the program code 630 may include various computer programs for implementing the various steps in the methods described above. The program code 630 is computer-readable code. These computer programs can be read from or written to one or more computer program products. These computer program products include program code carriers such as hard disks, CDs, memory cards, or floppy disks. The computer program includes computer-readable code that, when executed on an electronic device, causes the electronic device to perform the method according to the above embodiments.

[0149] This application also discloses a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the steps of the face liveness detection method as described in Embodiment 1 of this application.

[0150] Such a computer program product can be a computer-readable storage medium, which can have the same characteristics as... Figure 6 The memory 620 in the illustrated electronic device is similarly arranged with storage segments, storage spaces, etc. Program code can be stored, for example, in a compressed form on the computer-readable storage medium. The computer-readable storage medium is typically as shown in the reference... Figure 7 The portable or fixed storage unit is described above. Typically, the storage unit includes computer-readable code 630', which is code read by a processor and, when executed by the processor, implements the various steps of the method described above.

[0151] The terms "an embodiment," "embodiment," or "one or more embodiments" as used herein mean that a particular feature, structure, or characteristic described in connection with an embodiment is included in at least one embodiment of this application. Furthermore, please note that the examples of the phrase "in one embodiment" do not necessarily all refer to the same embodiment.

[0152] Numerous specific details are set forth in the specification provided herein. However, it will be understood that embodiments of this application may be practiced without these specific details. In some instances, well-known methods, structures, and techniques have not been shown in detail so as not to obscure the understanding of this specification.

[0153] In the claims, any reference signs placed between parentheses should not be construed as limiting the claims. The word "comprising" does not exclude the presence of elements or steps not listed in the claims. The word "a" or "an" preceding an element does not exclude the presence of a plurality of such elements. This application can be implemented by means of hardware comprising several different elements and by means of a suitably programmed computer. In a unit claim enumerating several means, several of these means may be embodied by the same item of hardware. The use of the words first, second, and third, etc., does not indicate any order. These words can be interpreted as names.

[0154] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.

Claims

1. A method for face liveness detection, characterized in that, include: Obtain the high-frequency and low-frequency feature components of each frame of the video image sequence used for face liveness detection; Based on the high-frequency feature components and low-frequency feature components of the video image, the inter-frame correlation of the video image frame sequence is obtained, wherein the inter-frame correlation includes inter-frame consistency and inter-frame difference; Based on the matching relationship between the inter-frame correlation and the pre-acquired inter-frame correlation of live faces, it is determined whether the face to be detected in the video image is a live face; The video image frame sequence includes at least two video images. Obtaining the inter-frame correlation of the video image frame sequence based on the high-frequency feature components and the low-frequency feature components of the video images includes: Based on the high-frequency feature components of at least two adjacent video images, the difference between the high-frequency feature components corresponding to each corresponding image position in the at least two adjacent video images is obtained to form the high-frequency frame difference feature of the video image frame sequence; and based on the low-frequency feature components of the at least two adjacent video images, the average value of the low-frequency feature components corresponding to each corresponding image position in the at least two adjacent video images is obtained to form the low-frequency equalization feature of the video image frame sequence. The high-frequency frame difference feature and the low-frequency equalization feature are weighted and concatenated to obtain the inter-frame correlation of the video image frame sequence.

2. The method according to claim 1, characterized in that, The pre-acquired inter-frame correlation of the live face is a surface that characterizes the distribution law of inter-frame correlation of high-frequency and low-frequency features in the live face video image sequence, obtained by fitting live face samples. The step of detecting whether the face to be detected in the video image acquisition is a live face based on the matching relationship between the inter-frame correlation and the pre-acquired inter-frame correlation of the live face includes: Substitute the obtained inter-frame correlation into the surface to determine whether the obtained inter-frame correlation conforms to the inter-frame correlation distribution pattern of high-frequency and low-frequency features in the live face video image sequence expressed by the surface. In response to the fact that the obtained inter-frame correlation conforms to the inter-frame correlation distribution law of high-frequency and low-frequency features in the live face video image sequence expressed by the surface, the face to be detected in the video image acquisition is determined to be a live face. In response to the fact that the obtained inter-frame correlation does not conform to the inter-frame correlation distribution pattern of high-frequency and low-frequency features in the live face video image sequence expressed by the surface, the face to be detected in the video image frame is determined to be a non-live face.

3. The method according to claim 2, characterized in that, Before detecting whether the face to be detected in the video image is a live face based on the matching relationship between the inter-frame correlation and the pre-acquired inter-frame correlation of live faces, the method further includes: Using video image sequences of live human faces as positive samples, a sample set including several of the positive samples is constructed; Obtain the high-frequency feature components and low-frequency feature components of each frame of video image in each positive sample in the sample set; For each positive sample, the inter-frame correlation characterizing the positive sample is obtained based on the high-frequency feature components and the low-frequency feature components of at least two frames of the video images in the positive sample. A surface is fitted to the inter-frame correlation of the positive samples in the sample set to obtain a surface that characterizes the distribution pattern of inter-frame correlation of high-frequency and low-frequency features in the live face video image sequence.

4. The method according to claim 3, characterized in that, After performing surface fitting on the inter-frame correlation of the positive samples in the sample set to obtain a surface characterizing the distribution law of inter-frame correlation of high-frequency and low-frequency features in the live face video image sequence, the method further includes: Substitute the inter-frame correlation of the positive samples in the sample set into the surface to obtain the distance between the inter-frame correlation of each positive sample and the surface; Based on the distribution of the distance between the inter-frame correlation and the surface of each positive sample, a distance threshold between the inter-frame correlation and the surface of the live face video image sequence is determined. The step of substituting the acquired inter-frame correlation into the surface to determine whether the acquired inter-frame correlation conforms to the inter-frame correlation distribution pattern of high-frequency and low-frequency features in the live face video image sequence expressed by the surface includes: Substitute the obtained inter-frame correlation into the surface to obtain the distance between the inter-frame correlation and the surface; In response to the distance being less than a predetermined distance threshold, it is determined that the inter-frame correlation conforms to the inter-frame correlation distribution pattern of high-frequency and low-frequency features in the live face video image sequence expressed by the surface. In response to the distance being greater than or equal to the predetermined distance threshold, it is determined that the inter-frame correlation does not conform to the inter-frame correlation distribution pattern of high-frequency and low-frequency features in the live face video image sequence expressed by the surface.

5. The method according to claim 1, characterized in that, The video image frame sequence includes multiple video images. Obtaining the inter-frame correlation of the video image frame sequence based on the high-frequency feature components and the low-frequency feature components of the video images includes: The high-frequency feature components of multiple specified video images in the video image frame sequence are concatenated into a multi-dimensional high-frequency feature component matrix in the form of one-dimensional high-frequency feature components corresponding to each frame of the video image; and the low-frequency feature components of multiple specified video images are concatenated into a multi-dimensional low-frequency feature component matrix in the form of one-dimensional low-frequency feature components corresponding to each frame of the video image. The multidimensional high-frequency feature component matrix and the multidimensional low-frequency feature component matrix are respectively subjected to low-rank sparse matrix decomposition to obtain a first matrix representing low-frequency background features, a second matrix representing low-frequency detail features, a third matrix representing high-frequency background features, and a fourth matrix representing high-frequency detail features. The inter-frame correlation of the video image frame sequence is obtained by representing inter-frame consistency through the nuclear norm of the first matrix and inter-frame differences through the column and norm of the fourth matrix.

6. The method according to claim 5, characterized in that, The pre-acquired inter-frame correlation of the live face consists of: a column and norm threshold characterizing the distribution pattern of high-frequency feature inter-frame correlation in the live face video image sequence, and a kernel norm threshold characterizing the distribution pattern of low-frequency feature inter-frame correlation in the live face video image sequence; the step of detecting whether the face to be detected in the video image acquisition is a live face based on the matching relationship between the inter-frame correlation and the pre-acquired inter-frame correlation of the live face includes: The nuclear norm of the first matrix is ​​compared with the nuclear norm threshold, and the column norm of the fourth matrix is ​​compared with the column norm threshold. In response to the fact that the nuclear norm of the first matrix is ​​less than the nuclear norm threshold, and the column norm of the fourth matrix is ​​greater than or equal to the column norm threshold, the face to be detected in the video image frame sequence is determined to be a live face. In response to the nuclear norm of the first matrix being greater than or equal to the nuclear norm threshold, or the column norm of the fourth matrix being less than the column norm threshold, the face to be detected in the video image frame sequence is determined to be a non-living face.

7. The method according to claim 6, characterized in that, Before detecting whether the face to be detected in the video image is a live face based on the matching relationship between the inter-frame correlation and the pre-acquired inter-frame correlation of live faces, the method further includes: Using video image sequences of live human faces as positive samples, a sample set including several of the positive samples is constructed; Obtain the high-frequency feature components and low-frequency feature components of each frame of video image in each positive sample in the sample set; For each positive sample, perform the following inter-frame correlation acquisition operation: The high-frequency feature components of multiple specified video images in the positive sample are concatenated into a multi-dimensional high-frequency feature component matrix in the form of one-dimensional high-frequency feature components corresponding to each frame of the video image in the positive sample; and the low-frequency feature components of multiple specified video images are concatenated into a multi-dimensional low-frequency feature component matrix in the form of one-dimensional low-frequency feature components corresponding to each frame of the video image. The multidimensional high-frequency feature component matrix and the multidimensional low-frequency feature component matrix are respectively subjected to low-rank sparse matrix decomposition to obtain a first matrix representing low-frequency background features and a fourth matrix representing high-frequency detail features. Obtain the nuclear norm of the first matrix and the column norm of the fourth matrix as the inter-frame correlation of the corresponding positive samples; Based on the inter-frame correlation of all the positive samples in the sample set, the kernel norm threshold of the first matrix used to characterize low-frequency background features and the column and norm thresholds of the fourth matrix used to characterize high-frequency detail features are determined.

8. A device for face liveness detection, characterized in that, include: The high and low frequency component acquisition module is used to acquire the high frequency feature components and low frequency feature components of each video image in the video image frame sequence used for face liveness detection. The inter-frame correlation acquisition module acquires the inter-frame correlation of the video image frame sequence based on the high-frequency feature components and the low-frequency feature components of the video image, wherein the inter-frame correlation includes inter-frame consistency and inter-frame difference. The liveness detection module is used to detect whether the face to be detected in the video image is a live face based on the matching relationship between the inter-frame correlation and the pre-acquired inter-frame correlation of live faces; The video image frame sequence includes at least two video images, and the inter-frame correlation acquisition module includes: a first inter-frame correlation acquisition submodule; The first inter-frame correlation acquisition submodule is configured to: acquire, based on the high-frequency feature components of at least two adjacent video images, the difference between the high-frequency feature components corresponding to each corresponding image position in the at least two adjacent video images, thereby constituting a high-frequency frame difference feature of the video image frame sequence; acquire, based on the low-frequency feature components of the at least two adjacent video images, the average value of the low-frequency feature components corresponding to each corresponding image position in the at least two adjacent video images, thereby constituting a low-frequency equalization feature of the video image frame sequence; and perform weighted concatenation of the high-frequency frame difference feature and the low-frequency equalization feature to acquire the inter-frame correlation of the video image frame sequence.

9. An electronic device, comprising a memory, a processor, and program code stored in the memory and executable on the processor, characterized in that, When the processor executes the program code, it implements the face liveness detection method according to any one of claims 1 to 7.

10. A computer-readable storage medium having program code stored thereon, characterized in that, When executed by a processor, the program code implements the steps of the face liveness detection method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Automatic database creating video sectioning method for automatic recognition of micro-expressions

    CN103426005A

  • Human face in-vivo detection method and system

    CN103679118A