Face recognition methods and devices in video mode

By combining color-cycled heart rate signals and spectral signals for face liveness detection, and extracting feature signals from the forehead and spectral regions of interest, reliable facial features are generated. This solves the problem of insufficient reliability of face recognition methods in video mode under liveness attacks and lighting changes, and improves recognition accuracy.

CN116704618BActive Publication Date: 2025-10-31FOSHAN HONGSHI INTELLIGENT INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210186728.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-02-28
Publication Date
2025-10-31
Estimated Expiration
2042-02-28

AI Technical Summary

Technical Problem

Existing video-based face recognition methods are not reliable enough in handling liveness attacks and under different lighting conditions, and do not make full use of video information, especially periodic heart rate signals and frequency domain information related to facial color changes.

Method used

By combining color-cycled heart rate signals and spectral signals, facial liveness detection is performed. Multidimensional facial features are obtained through imaging control, and feature signals are extracted using the forehead region and the region of interest in the spectrum to generate reliable facial features for recognition.

Benefits of technology

It improves the robustness and reliability of face recognition, solves the problem of uneven facial imaging under different lighting conditions, and enhances the recognition accuracy of the algorithm.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116704618B_ABST
    Figure CN116704618B_ABST
Patent Text Reader

Abstract

This application relates to the field of face recognition and discloses a face recognition method and apparatus in a video mode. The video-based face recognition method includes the following steps: acquiring a face video frame; locating the forehead region based on the face video frame; generating a liveness region of interest (GRI) and a spectral region of interest (GRI) in the forehead region; extracting a color-periodic heart rate signal from the GRI and extracting a facial spectral signal from the GRI; performing a weighted combination of the color-periodic heart rate signal and the facial spectral signal for face liveness determination; and acquiring multi-dimensional facial features through imaging control for enhanced face recognition. Combining the color-periodic heart rate signal and the spectral signal for face liveness determination makes the liveness determination more robust and reliable. Simultaneously, obtaining more reliable and discriminative multi-dimensional facial features for face recognition through imaging control can solve the problem of uneven facial imaging under different lighting conditions, improving the reliability of the algorithm.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of facial recognition, and in particular to facial recognition technology in video mode. Background Technology

[0002] Currently, deep learning is being used more and more widely in the field of facial recognition. However, various problems arise in real-world applications, leading to a decrease in the reliability of facial recognition algorithms. These include various liveness detection attacks and handling different lighting conditions.

[0003] For these practical applications, some solutions employ multi-view cameras or other similar structured light methods. However, such solutions significantly increase product costs and subsequent maintenance costs.

[0004] Furthermore, existing face recognition methods still have shortcomings in utilizing video information. Many commonly used methods are based on image processing techniques and do not fully utilize spatial and frequency domain information. For example, they use frequency domain information for image correction but do not consider the extraction of periodic signals. Other methods consider the relationship model between face anchor points but do not account for various interferences caused by different lighting conditions.

[0005] The inventors of this application have discovered that existing facial recognition methods do not consider the following points regarding the use of video information:

[0006] (1) Utilization of the periodic heart rate signal of facial color change;

[0007] (2) Facial feature extraction under different lighting conditions;

[0008] (3) There is uneven brightness on the face.

[0009] Therefore, there is an urgent need for a robust face recognition method that can fully utilize video information in a pure video mode. Summary of the Invention

[0010] The purpose of this application is to provide a face recognition method and apparatus in video mode, which combines color periodic heart rate signals and spectral signals for face liveness detection, making the liveness detection more robust and reliable; at the same time, face recognition based on multi-dimensional face features obtained by imaging control can solve the recognition problem under different lighting conditions and improve the reliability of the algorithm.

[0011] This application discloses a face recognition method in video mode, including the following steps:

[0012] Acquire face video frames;

[0013] Based on the facial video frames, locate the forehead area;

[0014] A live region of interest and a spectral region of interest are generated in the forehead region;

[0015] The color periodic heart rate signal is extracted from the region of interest of the living organism, and the facial spectrum signal is extracted from the region of interest of the spectrum.

[0016] The color periodic heart rate signal and the face spectrum signal are weighted and combined to determine face liveness.

[0017] Enhanced facial recognition is achieved by acquiring multi-dimensional facial features through imaging control.

[0018] In a preferred embodiment, the step of "acquiring multi-dimensional facial features through imaging control to enhance facial recognition" includes the following sub-steps:

[0019] Image groups are generated through imaging control;

[0020] Based on the image set, facial features under different brightness levels are obtained to generate a facial feature set;

[0021] By segmenting facial regions, brightness feature groups of facial blocks are extracted.

[0022] The facial feature group and the segmented brightness feature group are combined to synthesize the facial feature;

[0023] Facial recognition is performed based on the facial features.

[0024] In a preferred embodiment, the step of "generating an image group via imaging control" includes the following sub-steps:

[0025] The camera shutter speed, gain, and white balance are locked, and the brightness of the LED fill light is controlled to obtain the image group.

[0026] In a preferred embodiment, the step of "combining the face feature group and the segmented brightness feature group to synthesize the face feature" includes the following sub-steps:

[0027] Reliable face features are obtained by combining the face feature group and the block brightness feature group using a feature concatenation method.

[0028] The reliable facial features are then dimensionality-reduced to obtain lower-dimensional facial features.

[0029] In a preferred embodiment, the step of "generating a live region of interest and a spectral region of interest in the forehead region" includes the following sub-steps:

[0030] The forehead center position is obtained using a forehead localization method, and the living body region of interest is generated at this forehead center position. At the same time, the spectral region of interest is also generated at this forehead center position.

[0031] In a preferred embodiment, the step of "extracting the color periodic heart rate signal from the live body region of interest" includes the following sub-steps:

[0032] The video sequence is decomposed into pyramid multi-resolution to obtain images at different scales;

[0033] Temporal bandpass filtering is performed on the image at each scale to obtain several frequency bands of interest;

[0034] For each frequency band, the signal is differentially approximated using Taylor series, and the approximation result is linearly amplified.

[0035] The color-cycled heart rate signal is obtained through a heart rate estimation network.

[0036] In a preferred embodiment, the step of "extracting the face spectrum signal from the region of interest in the spectrum" includes the following sub-steps:

[0037] Select 3 frames from the video sequence to obtain the Fourier spectrum signal;

[0038] The face spectrum signal is obtained by training the spectral depth model function.

[0039] In a preferred embodiment, the step of "weightedly combining the color periodic heart rate signal and the face spectrum signal to determine face liveness" includes the following sub-steps:

[0040] Set the upper and lower threshold values ​​for the color cycle heart rate signal;

[0041] If the color-cycled heart rate signal is between the upper and lower thresholds, and the weighted face spectrum signal is 1, then the face is determined to be a live face.

[0042] This application also discloses a face recognition device in video mode, including:

[0043] The acquisition module is used to acquire face video frames;

[0044] The positioning module is used to locate the forehead region based on the facial video frames;

[0045] A region of interest generation module is used to generate a live region of interest and a spectral region of interest in the forehead region;

[0046] The extraction module is used to extract the color periodic heart rate signal from the live body region of interest and to extract the face spectrum signal from the spectrum region of interest.

[0047] The liveness detection module is used to perform liveness detection of faces by weighted combination of the color periodic heart rate signal and the face spectrum signal;

[0048] The recognition module is used to acquire multi-dimensional facial features through imaging control for enhanced facial recognition.

[0049] In a preferred embodiment, the identification module includes the following sub-modules:

[0050] The image group generation submodule is used to generate image groups through imaging control;

[0051] The face feature group generation submodule is used to obtain face features under different brightness levels based on the image group and generate a face feature group.

[0052] The block brightness feature group generation submodule is used to extract the block brightness feature groups of the face by dividing the face region;

[0053] The synthesis submodule is used to combine the face feature group and the block brightness feature group to synthesize the face feature;

[0054] The face recognition submodule is used to perform face recognition based on the facial features.

[0055] Compared with the prior art, the embodiments of this application include at least the following differences and effects:

[0056] A face recognition method based solely on video combines color-cycled heart rate signals and spectral signals for face liveness detection, making the liveness detection more robust and reliable. Simultaneously, by using imaging control to obtain more reliable and discriminative multi-dimensional facial features for enhanced face recognition, the method can solve the problem of uneven facial imaging under different lighting conditions, thus improving the algorithm's reliability.

[0057] Furthermore, by extracting the color periodic heart rate signal from the frequency domain information, the liveness attribute of the current face can be better determined.

[0058] Furthermore, imaging control is used to obtain more reliable and discriminative facial features for recognition.

[0059] Furthermore, utilizing imaging control to extract facial features under various lighting conditions can significantly improve the robustness of facial recognition algorithms.

[0060] Furthermore, by segmenting the face and supplementing the lighting information of facial features, the problem of uneven facial lighting can be solved, resulting in more reliable facial features.

[0061] The specification of this application contains numerous technical features distributed across various technical solutions. Listing all possible combinations of these technical features (i.e., technical solutions) would make the specification excessively lengthy. To avoid this problem, the various technical features disclosed in the above-described invention, the various technical features disclosed in the following embodiments and examples, and the various technical features disclosed in the accompanying drawings can be freely combined to form various new technical solutions (all of which are considered to have been described in this specification), unless such a combination of technical features is technically infeasible. For example, one example discloses feature A+B+C, and another example discloses feature A+B+D+E. Features C and D are equivalent technical means that serve the same function, and technically only one needs to be used; they cannot be used simultaneously. Feature E can technically be combined with feature C. Therefore, the solution A+B+C+D should not be considered as described because it is technically infeasible, while the solution A+B+C+E should be considered as described. Attached Figure Description

[0062] Figure 1 This is a flowchart illustrating a face recognition method in a video mode according to the first embodiment of this application;

[0063] Figure 2 This is a schematic diagram of the technical solution flow according to a specific embodiment of this application;

[0064] Figure 3 This is a schematic diagram of the structure of a forehead positioning depth network according to a specific embodiment of this application;

[0065] Figure 4 This is a schematic diagram of face imaging under different lighting conditions according to a specific embodiment of this application;

[0066] Figure 5 This is a schematic diagram of facial region division according to a specific embodiment of this application;

[0067] Figure 6 This is a structural schematic diagram of a face recognition device in video mode according to the second embodiment of this application. Detailed Implementation

[0068] In the following description, many technical details are presented to help the reader better understand this application. However, those skilled in the art will understand that the technical solutions claimed in this application can be implemented even without these technical details and various variations and modifications based on the following embodiments.

[0069] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.

[0070] The first embodiment of this application relates to a face recognition method in a video mode. Figure 1 This is a flowchart illustrating the face recognition method in this video format.

[0071] Specifically, such as Figure 1 As shown, the face recognition method in this video mode includes the following steps:

[0072] In step 101, facial video frames are acquired.

[0073] It should be noted that this application proposes a face recognition method based solely on video.

[0074] Then proceed to step 102, where the forehead region is located based on the facial video frames.

[0075] In this embodiment, preferably, a forehead localization algorithm is used to automatically locate the forehead region. The forehead region is obtained through a trained forehead model.

[0076] To enable rapid localization and facilitate deployment on ultra-low-power GPU devices, the low-complexity yolov5n algorithm is preferred.

[0077] Then proceed to step 103, where a live region of interest and a spectral region of interest are generated in the forehead region.

[0078] In this embodiment, preferably, step 103 may further include the following sub-steps:

[0079] The forehead center position is obtained using a forehead localization method, and the living body region of interest is generated at this forehead center position. At the same time, the spectral region of interest is also generated at this forehead center position.

[0080] Specifically, the forehead region (x, y, w, h) is obtained through a trained forehead model, where x and y are the coordinates of the top-left corner of the forehead region, and w and h are the width and height of the forehead region. The algorithm steps are as follows:

[0081] a) Prioritize using the forehead localization method to obtain the center position of the forehead, and generate the live region of interest based on this center position:

[0082] The center of the forehead is (x+w / 2, y+h / 2). The region of interest (ROI) is set to (x+w / 2, y+h / 2, w*r1, h*r1) using a scaling factor r1.

[0083] b) Simultaneously generate a region of interest in the spectrum for extracting the spectral signal:

[0084] Similarly, based on the center of the forehead, a region of interest in the spectrum is generated: (x+w / 2, y+h / 2, w*r2, h*r2), where r2 is the scaling factor of the region of interest in the spectrum.

[0085] Then proceed to step 104, where the color periodic heart rate signal is extracted from the region of interest of the living body, and the face spectrum signal is extracted from the region of interest of the spectrum.

[0086] Color-coded periodic heart rate signals can reflect important liveness features of a face. In some liveness attacks, such as those involving printed facial photographs, the color changes in the region of interest (ROI) of a displayed face differ significantly from those of a normal live face. Therefore, this feature can serve as an important reference signal for liveness detection.

[0087] Human blood circulation often causes subtle, periodic changes in the skin, usually imperceptible to the naked eye, but very similar to the heart rate. When the heart beats, blood flows through the blood vessels; the greater the blood volume, the more light is absorbed, and the less light is reflected from the skin. Therefore, periodic heart rate signals can be calculated by analyzing the time-frequency distribution of the forehead region of interest.

[0088] By extracting the color periodic heart rate signal from the frequency domain information, the liveness attribute of the current face can be better determined.

[0089] In this embodiment, preferably, the step of "extracting the color periodic heart rate signal from the region of interest in the living organism" includes the following sub-steps:

[0090] The video sequence is decomposed into pyramid multi-resolution to obtain images at different scales;

[0091] Temporal bandpass filtering is performed on the image at each scale to obtain several frequency bands of interest;

[0092] For each frequency band, the signal is differentially approximated using Taylor series, and the approximation result is linearly amplified.

[0093] The color-cycled heart rate signal is obtained through a heart rate estimation network.

[0094] The step of "extracting the face spectrum signal from the region of interest in the spectrum" includes the following sub-steps:

[0095] Select 3 frames from the video sequence to obtain the Fourier spectrum signal;

[0096] The face spectrum signal is obtained by training the spectral depth model function.

[0097] Then proceed to step 105, where the color periodic heart rate signal and the face spectrum signal are weighted and combined to determine the liveness of the face.

[0098] In this embodiment, preferably, step 105 may further include the following sub-steps:

[0099] Set the upper and lower threshold values ​​for the color cycle heart rate signal;

[0100] If the color-cycled heart rate signal falls between the upper and lower thresholds, and the weighted face spectrum signal is 1, then the face is determined to be live. The weighting coefficient lambda is configurable. Adding a lambda weighting coefficient to the liveness determination rule allows for more flexible configuration of liveness detection.

[0101] Then proceed to step 106, where multi-dimensional facial features are acquired through imaging control to enhance facial recognition.

[0102] In this embodiment, preferably, step 106 may further include the following sub-steps:

[0103] Image groups are generated through imaging control;

[0104] Based on the image set, facial features under different brightness levels are obtained to generate a facial feature set;

[0105] By segmenting facial regions, brightness feature groups of facial blocks are extracted.

[0106] The facial feature group and the segmented brightness feature group are combined to synthesize the facial feature;

[0107] Facial recognition is performed based on the facial features.

[0108] Furthermore, preferably, the step of "generating an image group through imaging control" may further include the following sub-steps:

[0109] The camera shutter speed, gain, and white balance are locked, and the brightness of the LED fill light is controlled to obtain the image group.

[0110] The step of "combining the face feature group and the segmented brightness feature group to synthesize the face feature" includes the following sub-steps:

[0111] Reliable face features are obtained by combining the face feature group and the block brightness feature group using a feature concatenation method.

[0112] The reliable facial features are then dimensionality-reduced to obtain lower-dimensional facial features.

[0113] Imaging control is used to obtain more reliable and discriminative facial features for recognition. Utilizing imaging control for facial feature extraction under various lighting conditions significantly improves the robustness of facial recognition algorithms. By segmenting the face and supplementing it with lighting information, the problem of uneven facial illumination can be solved, resulting in more reliable facial features.

[0114] This process will then end.

[0115] In summary, this application proposes a face recognition method based solely on video, which combines color periodic heart rate signals and spectral signals for face liveness detection, making the liveness detection more robust and reliable. Furthermore, by using imaging control to obtain more reliable and discriminative multi-dimensional facial features for face recognition, the problem of uneven facial imaging under different lighting conditions can be solved, thus improving the reliability of the algorithm.

[0116] A specific embodiment of this application is described below. Figure 2 This is a flowchart illustrating the technical solution of this specific embodiment. The specific process is described as follows:

[0117] 1. Generate forehead location points, prioritizing the use of forehead localization algorithms to automatically locate the forehead region:

[0118] a) To enable rapid localization and facilitate deployment on ultra-low-power GPU devices, the low-complexity yolov5n algorithm is preferred.

[0119] Figure 3 The structure of a forehead localization deep network was defined. After calibrating the collected dataset, the network was used to train the model.

[0120] Related operating instructions:

[0121] Concat: Tensor concatenation, which expands the dimensions of two tensors. For example, concatenating two tensors 112*112*64 and 112*112*64 results in 112*112*128.

[0122] `add`: Tensors are added together. Tensors are added directly without expanding the dimensions. For example, adding 112*112*64 and 112*112*64 will still result in 112*112*64.

[0123] BottleNeck: Bottleneck layer.

[0124] CSPNet stands for Cross Stage Paritial Network, which mainly addresses the problem of high computational cost in inference from the perspective of network structure design.

[0125] SPP: Space Pyramid Structure.

[0126] UpSample: Upsampling operation.

[0127] Conv: Convolution.

[0128] 2. Generate the ROI (Region of Interest):

[0129] The forehead region (x, y, w, h) is obtained using a trained forehead model, where x and y are the coordinates of the top-left corner of the region, and w and h are the width and height. The algorithm steps are as follows:

[0130] a) Prioritize using forehead localization methods to obtain the center position of the forehead, and generate a live ROI from this center:

[0131] The center position of the forehead (x+w / 2, y+h / 2) is used to set the living ROI region as (x+w / 2, y+h / 2, w*r1, h*r1) using a living ROI scaling factor r1.

[0132] b) Simultaneously generate a spectral ROI for extracting the spectral signal:

[0133] Similarly, based on the center of the forehead, a spectrum ROI (x+w / 2, y+h / 2, w*r2, h*r2) is generated, where r2 is the spectrum ROI scaling factor.

[0134] 3. Extract color-cycled heart rate signals for liveness detection:

[0135] Color-coded periodic heart rate signals can reflect important liveness features of a face. In some liveness attack scenarios, such as printed or screen-displayed facial images, the color changes in the liveness region of interest (ROI) differ significantly from those of a normal live face. Therefore, this feature can serve as an important reference signal for liveness detection.

[0136] Algorithm steps:

[0137] a) Extract the color periodic signal Sc from the ROI region:

[0138] Human blood circulation often causes subtle, periodic changes in the skin, usually imperceptible to the naked eye, but very similar to the heart rate. When the heart beats, blood flows through the blood vessels; the greater the blood volume, the more light is absorbed, and the less light is reflected from the skin. Therefore, periodic signals can be calculated through time-frequency analysis of the forehead's Region of Interest (ROI). Processing steps:

[0139] 1) Spatial color filtering. The video sequence is decomposed into a pyramid multi-resolution structure;

[0140] 2) Temporal filtering. Temporal bandpass filtering is performed on the image at each scale to obtain several frequency bands of interest;

[0141] 3) Amplify the filtering results. For each frequency band, the signal is approximated differentially using Taylor series, and the approximation result is linearly amplified.

[0142] Let I(x,t) represent the pixel value at position x and time t, δ(t) represent the object's displacement, and α represent the magnification factor. Then I(x,0) = f(x) and I(x,t) = f(x + δ(t)). The magnification calculation is as follows:

[0143]

[0144] 4) Obtain the periodic heart rate signal Sc through a heart rate estimation network;

[0145] The RhythmNet method is preferred to calculate the heart rate signal Sc from the image sequence of the forehead live ROI region using a trained depth model: Sc = FL(I(x,y,n)), where FL is the trained heart rate depth model function, I(x,y,n) is the video frame image, and n is the number of video frames used for analysis, usually selected as 75 or more.

[0146] b) Select 3 frames of images to extract the face spectral signal Sf:

[0147] First, select 3 frames from the video sequence to obtain the Fourier spectrum signal. The spectrum calculation formula is as follows:

[0148]

[0149] Where f(x,y) is the input image, (x,y) represents the image coordinates, F(u,v) represents the spatial spectrum image, j is the imaginary unit, M and N are the image width and height, and π is pi.

[0150] Sf = FM(F(u,v,k)), where FM is the trained spectral depth model function, F(u,v,k) is the video spectral image, and k is usually 3. Sf is the attack category, where 0 represents non-live and 1 represents live.

[0151] c) Weighted combination of color periodic signal Sc and spectrum Sf for liveness detection:

[0152] The logic for determining liveness is as follows:

[0153] 1) Set the Sc threshold, T1 is the lower limit of heart rate, usually set to 40, and T2 is the upper limit of heart rate, usually set to 150;

[0154] 2) Obtain Sc and Sf;

[0155] 3) If T1 < Sc < T2 and at the same time lamda * Sf is 1, the current face detection is a live body, and the algorithm can continue.

[0156] Otherwise, it is determined as a non - live body and the algorithm is aborted.

[0157] 4. Enhanced face feature extraction and recognition method combined with imaging control:

[0158] a) Lock the camera shutter, gain, and white balance, and control the brightness of the LED fill light to obtain an image group ImageK:

[0159] ImageK = I(x, y, k), where k usually takes values from 3 to 5, indicating 3 - 5 groups of brightness images.

[0160] b) Obtain the face features under K groups of brightness for the image group ImageK to generate a face feature group Fk:

[0161] Fk = F facemodel (ImageK), where F facemodel is a face recognition deep model, and a deep model based on the arcface method is preferably used. The model depth preferably adopts the ResNet network and is selected according to the computational amount in terms of the model depth.

[0162] c) Divide the face area and extract a segmented face brightness feature group Fkb:

[0163] Figure 4 Shows face imaging under different lighting conditions. The face area is divided, and the division method includes the following steps:

[0164] Step 1: Locate the face feature points. Preferably, a face feature point localization method based on deep learning is used, such as MTCNN, etc.;

[0165] Step 2: Divide M core regions of the face with the eyes, nose, and mouth respectively, as Figure 5 shown;

[0166] Fkb = F light (M(k)), where F light is a brightness feature extraction function, and M(k) is the image of the k - th core region. The extraction method preferably uses histogram equalization and statistically calculates feature information such as information entropy, brightness mean and variance, and brightness median.

[0167] d) Combine the face feature group Fk and the segmented face brightness feature group Fkb:

[0168] i. Preferably use the feature concatenation method concat:

[0169] feat=concat(Fk,Fkb)

[0170] The concat function combines the face feature set Fk and the brightness feature set Fkb. Concat is a tensor concatenation operation that expands the dimensions of the two tensors to obtain the final reliable feature feat that describes the current face.

[0171] i. To further reduce the dimensionality of features, PCA can be used for dimensionality reduction to obtain lower-dimensional facial features:

[0172] From the perspective of feature storage, this method prioritizes PCA to reduce the dimensionality of features, thereby controlling the feature dimension between 128 and 1024.

[0173] PCA stands for Principal Component Analysis. It uses orthogonal transformations to convert a set of potentially correlated variables into a set of linearly uncorrelated variables; these transformed variables are called principal components. This is a commonly used statistical method.

[0174] iii. Facial feature comparison:

[0175] The recognition method uses cosine similarity to calculate the distance between two facial features. The calculation method is as follows:

[0176]

[0177] Ai and Bi represent two facial features, and n is the feature length. Generally, a similarity of cosΘ greater than 0.6 indicates that they are the same person; otherwise, they are different people.

[0178] It should be noted that all embodiments of the present invention can be implemented in software, hardware, firmware, etc. Regardless of whether the present invention is implemented in software, hardware, or firmware, the instruction code can be stored in any type of computer-accessible memory (e.g., permanent or modifiable, volatile or non-volatile, solid-state or non-solid-state, fixed or replaceable media, etc.). Similarly, the memory can be, for example, Programmable Array Logic (PAL), Random Access Memory (RAM), Programmable Read Only Memory (PROM), Read-Only Memory (ROM), Electrically Erasable Programmable ROM (EEPROM), magnetic disk, optical disk, Digital Versatile Disc (DVD), etc.

[0179] The second embodiment of this application relates to a face recognition device in video mode. Figure 6 This is a schematic diagram of the facial recognition device in this video mode.

[0180] Specifically, such as Figure 6 As shown, the face recognition device in this video mode includes:

[0181] The acquisition module is used to acquire face video frames;

[0182] The positioning module is used to locate the forehead region based on the facial video frames;

[0183] A region of interest generation module is used to generate a live region of interest and a spectral region of interest in the forehead region;

[0184] The extraction module is used to extract the color periodic heart rate signal from the live body region of interest and to extract the face spectrum signal from the spectrum region of interest.

[0185] The liveness detection module is used to perform liveness detection of faces by weighted combination of the color periodic heart rate signal and the face spectrum signal;

[0186] The recognition module is used to acquire multi-dimensional facial features through imaging control for enhanced facial recognition.

[0187] In this embodiment, preferably, the identification module may further include the following sub-modules:

[0188] The image group generation submodule is used to generate image groups through imaging control;

[0189] The face feature group generation submodule is used to obtain face features under different brightness levels based on the image group and generate a face feature group.

[0190] The block brightness feature group generation submodule is used to extract the block brightness feature groups of the face by dividing the face region;

[0191] The synthesis submodule is used to combine the face feature group and the block brightness feature group to synthesize the face feature;

[0192] The face recognition submodule is used to perform face recognition based on the facial features.

[0193] It should be noted that this embodiment is a device embodiment corresponding to the first embodiment, and this embodiment can be implemented in conjunction with the first embodiment. The relevant technical details mentioned in the first embodiment remain valid in this embodiment, and will not be repeated here to avoid repetition. Correspondingly, the relevant technical details mentioned in this embodiment can also be applied to the first embodiment.

[0194] All modules mentioned in the various device embodiments of the present invention are logical modules. Physically, a logical module can be a physical module, a part of a physical module, or a combination of multiple physical modules. The physical implementation of these logical modules is not the most important factor; rather, the combination of functions implemented by these logical modules is the key to solving the technical problem proposed by the present invention. Furthermore, to highlight the innovative aspects of the present invention, the above-described device embodiments have not introduced modules that are not closely related to solving the technical problem proposed by the present invention. This does not mean that the above-described device embodiments do not contain other modules.

[0195] It should be noted that in this patent application, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one" does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element. In this patent application, if it refers to performing an action according to an element, it means performing the action at least according to that element, including two cases: performing the action only according to that element, and performing the action according to that element and other elements. Expressions such as "multiple," "repeatedly," and "various" include two, two times, two kinds, and more than two, more than two times, and more than two kinds.

[0196] Although the invention has been illustrated and described with reference to certain preferred embodiments thereof, those skilled in the art will understand that various changes in form and detail may be made therein without departing from the spirit and scope of the invention.

Claims

1. A face recognition method in video mode, characterized in that, Includes the following steps: Acquire face video frames; Based on the facial video frames, locate the forehead area; A live region of interest and a spectral region of interest are generated in the forehead region; The color periodic heart rate signal is extracted from the region of interest of the living organism, and the facial spectrum signal is extracted from the region of interest of the spectrum. The color periodic heart rate signal and the face spectrum signal are weighted and combined to determine face liveness. Enhanced facial recognition is achieved by acquiring multi-dimensional facial features through imaging control; The step of "extracting the color periodic heart rate signal from the region of interest in the living organism" includes the following sub-steps: The video sequence is decomposed into pyramid multi-resolution to obtain images at different scales; Temporal bandpass filtering is performed on the image at each scale to obtain several frequency bands of interest; For each frequency band, the signal is differentially approximated using Taylor series, and the approximation result is linearly amplified. The color-cycled heart rate signal is obtained through a heart rate estimation network; The step of "extracting the face spectrum signal from the region of interest in the spectrum" includes the following sub-steps: Select 3 frames from the video sequence to obtain the Fourier spectrum signal; The face spectrum signal is obtained by training the spectral depth model function; The step of "weightedly combining the color periodic heart rate signal and the face spectrum signal to determine face liveness" includes the following sub-steps: Set the upper and lower threshold values ​​for the color cycle heart rate signal; If the color-cycled heart rate signal is between the upper and lower thresholds, and the weighted face spectrum signal is 1, then the face is determined to be a live face.

2. The face recognition method in video mode as described in claim 1, characterized in that, The step of "acquiring multi-dimensional facial features through imaging control to enhance facial recognition" includes the following sub-steps: Image groups are generated through imaging control; Based on the image set, facial features under different brightness levels are obtained to generate a facial feature set; By segmenting facial regions, brightness feature groups of facial blocks are extracted. The facial feature group and the segmented brightness feature group are combined to synthesize the facial feature; Facial recognition is performed based on the facial features.

3. The face recognition method in video mode as described in claim 2, characterized in that, The step of "generating an image group through imaging control" includes the following sub-steps: The camera shutter speed, gain, and white balance are locked, and the brightness of the LED fill light is controlled to obtain the image group.

4. The face recognition method in video mode as described in claim 2, characterized in that, The step of "combining the face feature group and the segmented brightness feature group to synthesize the face feature" includes the following sub-steps: Reliable face features are obtained by combining the face feature group and the block brightness feature group using a feature concatenation method. The reliable facial features are then dimensionality-reduced to obtain lower-dimensional facial features.

5. The face recognition method in video mode as described in claim 1, characterized in that, The step of "generating a live region of interest and a spectral region of interest in the forehead region" includes the following sub-steps: The forehead center position is obtained using a forehead localization method, and the living body region of interest is generated at this forehead center position. At the same time, the spectral region of interest is also generated at this forehead center position.

6. A face recognition device in video mode, characterized in that, include: The acquisition module is used to acquire face video frames; The positioning module is used to locate the forehead region based on the facial video frames; A region of interest generation module is used to generate a live region of interest and a spectral region of interest in the forehead region; The extraction module is used to extract the color periodic heart rate signal from the live body region of interest and to extract the face spectrum signal from the spectrum region of interest. The liveness detection module is used to perform liveness detection of faces by weighted combination of the color periodic heart rate signal and the face spectrum signal; The recognition module is used to acquire multi-dimensional facial features through imaging control to enhance facial recognition; The extraction module, when extracting the color-cycled heart rate signal from the region of interest in the living organism, includes the following sub-steps: The video sequence is decomposed into pyramid multi-resolution to obtain images at different scales; Temporal bandpass filtering is performed on the image at each scale to obtain several frequency bands of interest; For each frequency band, the signal is differentially approximated using Taylor series, and the approximation result is linearly amplified. The color-cycled heart rate signal is obtained through a heart rate estimation network; When the extraction module extracts the face spectrum signal from the region of interest, it includes the following sub-steps: Select 3 frames from the video sequence to obtain the Fourier spectrum signal; The face spectrum signal is obtained by training the spectral depth model function; The liveness detection module includes the following sub-steps when performing face liveness detection: Set the upper and lower threshold values ​​for the color cycle heart rate signal; If the color-cycled heart rate signal is between the upper and lower thresholds, and the weighted face spectrum signal is 1, then the face is determined to be a live face.

7. The face recognition device in video mode as described in claim 6, characterized in that, The identification module includes the following sub-modules: The image group generation submodule is used to generate image groups through imaging control; The face feature group generation submodule is used to obtain face features under different brightness levels based on the image group and generate a face feature group. The block brightness feature group generation submodule is used to extract the block brightness feature groups of the face by dividing the face region; The synthesis submodule is used to combine the face feature group and the block brightness feature group to synthesize the face feature; The face recognition submodule is used to perform face recognition based on the facial features.

Citation Information

Patent Citations

  • A face recognition method based on living body detection

    CN109409343A

  • Heart rate detection method based on ordinary camera

    CN110384491A