Multi-modal living body detection method

By combining grayscale information from visible light and infrared modes with a multimodal liveness detection method, and dynamically adjusting modal weights, a collaborative defense architecture of physiology, environment, and physics is constructed. This solves the problem of low recognition accuracy of traditional methods in complex environments and achieves high-precision liveness detection.

CN122024337AInactive Publication Date: 2026-05-12GUANGZHOU JEEKUP INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
GUANGZHOU JEEKUP INFORMATION TECH CO LTD
Filing Date
2026-04-10
Publication Date
2026-05-12
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Traditional single-modal or simple fusion techniques cannot effectively distinguish real living organisms, nor can they identify subpixel-level digital synthesis traces or frequency anomalies. As a result, when faced with highly realistic visual dynamic digital forgery attacks, the accuracy of identifying living organisms drops significantly, easily leading to misjudgments or missed detections, which seriously threatens data security.

Method used

A multimodal liveness detection method is adopted. By acquiring grayscale information in visible light and infrared modes, the signal sequence of the forehead region is extracted, filtered and fast Fourier transform is performed, the correlation strength of physiological pulsation and the confidence of environmental mode are calculated, and the texture probability of convolutional neural network is combined to dynamically adjust the mode weights and construct a collaborative defense architecture of physiological-environmental-physical.

Benefits of technology

It effectively identifies highly realistic digital forgeries, improves the accuracy and robustness of liveness detection, and ensures security in complex scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122024337A_ABST
    Figure CN122024337A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of image processing, and particularly relates to a multi-modal in-vivo detection method, which comprises the following steps of: acquiring related data of two paths of signals corresponding to each frame of image of a face of a current detector, calculating a pulsation intensity sequence of the two paths of signals, and performing band-pass filtering and fast Fourier transform to obtain a pulse intensity sequence of the two paths of signals; the pulsation intensity sequence of the filtered two-path time domain signals and the main frequency of the two-path signals are obtained, and the physiological pulsation correlation intensity is calculated; environment modal credibility is calculated according to the gray level change of each modal in each frame of image; and fusing the physiological pulsation correlation intensity, the environmental modal credibility and the texture probability output by the pre-trained convolutional neural network to determine the confidence of the living body. According to the method, through the synergistic effect of the microscopic physiological features, the macroscopic texture and the environmental perception, the recognition precision of resisting the digital forgery attack in the complex illumination environment is effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image processing technology. More specifically, this invention relates to a multimodal liveness detection method. Background Technology

[0002] In the current digital identity authentication scenario, liveness detection technology is the core barrier to defend against identity fraud and protect the security of digital assets. The core purpose of liveness detection is to identify whether the current subject is a real biological living being through technical means, rather than non-living attacks through photos, screen captures, or 3D masks, which seriously threaten personal privacy and digital asset security. Therefore, it is necessary to conduct intelligent detection of biological liveness to effectively defend against new types of digital forgery attacks and ensure the security of real-name authentication.

[0003] Currently, the mainstream liveness detection methods in the industry mainly rely on single-modal texture analysis or command actions. Traditional algorithms typically use Gaussian smoothing filters to extract edge texture features of images to determine whether there are reflection features of non-natural surfaces; or they determine the validity of biometric features by detecting preset actions such as blinking and opening the mouth.

[0004] However, due to the drastic changes in ambient light in different authentication scenarios and the covert nature of forgery techniques, traditional single-modal or simple fusion techniques cannot effectively distinguish real living organisms. For example, images in strong light environments are prone to oversaturation, and traditional fixed logic cannot dynamically adjust the weighting of different modalities and ignores the co-evolutionary relationship of different modalities in microscopic physiological characteristics. It cannot identify subpixel-level digital synthesis traces or frequency anomalies, resulting in a significant decrease in the accuracy of identifying living organisms when faced with highly realistic visual dynamic digital forgery attacks. This can easily lead to misjudgments or missed detections, seriously threatening data security. Summary of the Invention

[0005] To address the aforementioned technical problems of poor modal robustness due to ambient light interference and inaccurate judgments caused by the inability of forged media to simulate deep biological microscopic physiological laws, this invention provides a multimodal liveness detection method, comprising: acquiring each frame of an image containing the face of the current subject, wherein the image contains grayscale information in visible light and infrared modes; for each frame of the image: extracting the forehead region and acquiring visible light signal sequences and infrared signal sequences of the forehead region, using the average of all signal values ​​in each signal sequence as the pulsation intensity of each signal; filtering the pulsation intensity sequences of each signal in all frames of the image, acquiring the filtered pulsation intensity sequences of each signal and performing a Fast Fourier Transform; The process involves: acquiring the main frequency of each signal; determining the physiological pulsation correlation strength based on the difference in main frequencies between two signals and the correlation coefficient of the pulsation intensity sequences of the two filtered time-domain signals; calculating the effective grayscale ratio of the infrared mode in each frame image; constructing dynamic environment weights based on the degree to which the average grayscale of the visible light mode in each frame image deviates from the preset ideal imaging brightness; using the dynamic environment weights to correct the effective grayscale ratio of the infrared mode and determine the environmental mode confidence; inputting all frame images into a pre-trained convolutional neural network to obtain texture probabilities; fusing the physiological pulsation correlation strength, environmental mode confidence, and texture probabilities to determine the liveness confidence; and determining a liveness status when the liveness confidence exceeds a set threshold.

[0006] This invention integrates dual-spectral grayscale information to construct a closed-loop evaluation system from two dimensions: microscopic physiological pulsation and macroscopic environmental modality. It captures intrinsic consistency by utilizing the correlation strength of physiological pulsation, effectively identifying highly realistic digital forgeries. By introducing environmental modality credibility and dynamically adjusting modality weights, it solves the problem of single modality being prone to saturation or high noise under extreme lighting conditions. Finally, by coupling deep texture probability, it achieves multi-dimensional cross-verification of physiological, physical, and environmental aspects, greatly improving the accuracy and robustness of liveness detection in complex scenarios.

[0007] Preferably, the step of extracting the forehead region and collecting the visible light signal sequence and infrared signal sequence of the forehead region includes: extracting the forehead region by locating and identifying it using a facial key point detection algorithm; and collecting the visible light signal and infrared signal of the forehead region using RGB and IR sensors respectively.

[0008] Preferably, obtaining the pulsation intensity sequence of the two filtered signals includes: arranging the pulsation intensity of each signal of all frame images in chronological order of acquisition time, and filtering them using a fourth-order Butterworth bandpass filter to obtain the pulsation intensity sequence of each filtered signal; and setting the filtering frequency range to 0.7Hz to 4Hz.

[0009] This invention employs a fourth-order Butterworth bandpass filter and focuses on the frequency range of 0.7Hz to 4Hz, locking in the normal heart rate range of the human body under resting and exercise conditions. It can effectively eliminate low-frequency drift caused by drastic changes in ambient light and filter out high-frequency electronic noise generated by the sensor, providing high-purity raw data for subsequent physiological signal extraction and ensuring the stability of monitoring.

[0010] Preferably, obtaining the main frequency of each signal includes: using a fast Fourier transform to convert the pulsation intensity sequence of each filtered signal to the frequency domain, and identifying the frequency value of the point where the energy is most concentrated in the frequency domain as the main frequency of each signal.

[0011] Preferably, determining the physiological pulsation correlation strength includes: calculating the Pearson correlation coefficient of the pulsation intensity sequences of the two filtered signals; calculating the ratio of the absolute difference of the dominant frequencies of the two signals to the sum of the dominant frequencies of the two signals, and inputting the ratio as a negative independent variable into the natural exponential function; taking the maximum value of the Pearson correlation coefficient and 0, and calculating the product of the maximum value and the output of the natural exponential function to obtain the physiological pulsation correlation strength.

[0012] This invention measures the synchronization of two modal signals by constructing a correlation formula that includes the degree of resonance of waveform fluctuations and the penalty term for the main frequency deviation. It can keenly identify frequency shifts and phase loss caused by electronic screen refresh rate or rendering delay. By utilizing the consistency of physiological characteristics at the physical level, it effectively intercepts highly realistic visual forgery attacks that are difficult for traditional algorithms to identify.

[0013] Preferably, determining the environmental modality confidence level includes: for each frame of image: the effective grayscale ratio of the infrared modality is equal to the ratio of the standard deviation of the grayscale values ​​of all pixels in the infrared modality to the maximum grayscale standard deviation benchmark value calibrated by the infrared sensor; calculating the absolute difference between the mean of the grayscale values ​​of all pixels in the visible light modality and the preset ideal imaging brightness, summing the preset ideal imaging brightness and the absolute difference, and the dynamic environmental weight is equal to the ratio of the preset ideal imaging brightness and the summation result; multiplying the effective grayscale ratio of the infrared modality by the dynamic environmental weight, and taking the average of the product results of all frames of image to obtain the environmental modality confidence level.

[0014] This invention introduces a contrast standard deviation and preset ideal imaging brightness deviation evaluation mechanism, which enables the system to perceive ambient light quality in real time. By weighting degraded modes or positively incentivizing high-quality modes, dynamic optimization of the liveness detection logic is achieved.

[0015] Preferably, the pre-trained convolutional neural network adopts the MobileNetV3 model.

[0016] Preferably, determining the liveness confidence includes: multiplying the product of the physiological pulsation association strength and the environmental modality confidence by the square root of the sum of the square of the physiological pulsation association strength, the square of the environmental modality confidence, and a preset harmonic parameter to obtain a first ratio; and multiplying the first ratio by the texture probability as the liveness confidence.

[0017] This invention nonlinearly couples microscopic physiological signals with macroscopic physical textures, and only provides positive excitation when physiological rhythms are synchronized and the imaging environment is reliable. At the same time, it introduces harmonic parameters to suppress abnormal fluctuations, ensuring the scientific validity of the final confidence level.

[0018] Preferably, the method further includes: if the liveness confidence level is less than or equal to a set threshold, determining it as a non-liveness attack and sending an abnormal sample alarm, and retaining the current abnormal sample evidence through an interface.

[0019] Preferably, the method further includes: using a binocular synchronous camera for image acquisition.

[0020] The beneficial effects of this invention are as follows: This invention constructs a collaborative defense architecture of physiology, environment, and physics. It utilizes the dual-spectral phase-locking characteristic to identify microscopic forgeries, and ensures that the decision-making focus is optimized in real time according to the lighting environment through a dynamic weight adjustment mechanism. Finally, it performs multimodal deep fusion of deep convolution features and time-frequency domain physical features, thereby improving the accuracy of liveness detection and building a solid security barrier for digital identity authentication. Attached Figure Description

[0021] Figure 1 This is a flowchart illustrating a multimodal liveness detection method according to the present invention; Figure 2 This schematically illustrates the in vivo multimodal physiological pulsation curve; Figure 3 It schematically illustrates the non-in vivo multimodal physiological pulsation curve. Detailed Implementation

[0022] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0023] The specific embodiments of the present invention will now be described in detail with reference to the accompanying drawings.

[0024] This invention discloses a multimodal liveness detection method, referring to... Figure 1 This includes steps S1 to S4: S1. Acquire each frame of image containing the face of the current detector, wherein the image contains grayscale information in visible light mode and infrared mode; for each frame of image: extract the forehead region and collect the visible light signal sequence and infrared signal sequence of the forehead region, and take the average value of all signal values ​​in each signal sequence as the pulsation intensity of each signal.

[0025] It should be noted that the most fundamental difference between living organisms and fake media lies in the microscopic fluctuations caused by blood circulation in deep tissues. Traditional visible light (RGB) cameras are easily obscured by external light, while near-infrared (IR) modes can penetrate the skin surface and capture changes in the absorption of infrared radiation by blood flow. Therefore, by constructing a dual-spectral temporal window and comparing the dynamic evolution of the same spatial location under different spectra, the physical reflection of ambient light on the skin surface can be separated, and the pure biological pulsation signal can be accurately extracted.

[0026] Specifically, considering that the forehead region has a thinner skin layer and a relatively uniform distribution of blood vessels, it can more clearly reflect changes in blood oxygen levels driven by heartbeat; therefore, a binocular synchronous camera is used to acquire video frames containing the face of the current subject at a preset frequency. For each frame, a facial key point detection algorithm is used to locate the forehead region and acquire two signal sequences of the forehead region. The two signal acquisition steps are as follows: the position of the forehead region is sent to two sensors, RGB and IR, and the sensors acquire visible light signal sequences and infrared signal sequences of the forehead region, respectively.

[0027] It should be added that, in this embodiment of the invention, the preset frequency is 40 frames / second to obtain biological pulsation signals with sufficient periodicity, and each frame of image acquired by the binocular synchronous camera includes grayscale information in both visible light and infrared modes.

[0028] Furthermore, the mean of all signal values ​​in each signal sequence within the region of each frame image is calculated as the pulsation intensity of each signal in each frame image.

[0029] At this point, the pulsation intensity of each signal in each frame of the current detector's image has been obtained.

[0030] S2. Obtain the pulsation intensity sequences of the two filtered signals and perform Fast Fourier Transform on them respectively to obtain the main frequency of each signal; determine the physiological pulsation correlation strength based on the difference in the main frequencies of the two signals and the correlation coefficient of the pulsation intensity sequences of the two filtered time-domain signals.

[0031] It should be noted that in a real living organism, because blood oxygen concentration is driven by the heartbeat, the fluctuations in the green channel of the RGB signal and the absorption fluctuations in the near-infrared channel have extremely high phase-locking characteristics in the time dimension; while forged synthetic frames often have rendering delays, or the moiré patterns generated by the refresh rate of the electronic screen in the two sensors are incoherent; therefore, by analyzing the coupling degree of the two signals in the time and frequency domain, possible forged signals can be determined from a physical perspective.

[0032] Specifically, the pulsation intensity of each signal from all frames is arranged sequentially according to the acquisition time and input into a fourth-order Butterworth bandpass filter. Filtering is performed according to a preset filtering frequency range to obtain the pulsation intensity sequence of each filtered signal. The Pearson correlation coefficient between the pulsation intensity sequences of the two filtered signals is then calculated. It should be noted that in this embodiment, the filtering frequency range is set to 0.7Hz to 4Hz, which corresponds to the normal heart rate range of a human body at rest and during exercise, effectively eliminating low-frequency drift caused by drastic changes in ambient light and high-frequency electronic noise from the sensor.

[0033] The fast Fourier transform is used to convert the pulsation intensity sequence of each filtered signal to the frequency domain, and the frequency value of the point where the energy is most concentrated in the frequency domain is identified as the main frequency of each signal.

[0034] The physiological pulsation correlation strength is determined based on the difference in the main frequencies of the two signals and the correlation coefficient of the pulsation intensity sequences of the filtered two time-domain signals; the physiological pulsation correlation strength satisfies the expression:

[0035] In the formula, The correlation strength of physiological pulsation; Pearson correlation coefficient between the pulsation intensity sequence of the filtered visible light signal and the pulsation intensity sequence of the infrared signal; , The dominant frequencies of visible light and infrared signals; To obtain the maximum value; It is a natural exponential function.

[0036] in, The value reflects the degree of resonance between the two filtered signals of the current detector in terms of waveform fluctuation trend. The larger the value, the more synchronized the physiological rhythms detected by the two spectra of the current detector are in terms of time evolution. This effectively eliminates the misleading effect of single signal caused by local environmental reflection or random jump. The maximum value function ensures that the system only provides excitation to positive physiological correlation. This reflects the relative symmetrical frequency difference rate of the two filtered signals from the current detector. Using the outer natural exponential function, a non-linear smoothing frequency penalty term is formed. When the two main frequencies perfectly coincide, it indicates that the physiological pulsation signals measured by the two sensors are completely identical. At this point, the penalty term... Approaching 1; if the object being detected is a screen photograph, due to refresh rate and synthesis error, the two signals often exhibit frequency offset, resulting in a decrease in the penalty term value, thereby lowering the overall correlation strength. By multiplying and cascading the above-mentioned index representing the resonance degree of the time-domain waveform with the penalty term, a time-frequency dual-domain joint gating mechanism is established. Only when the two signals are highly correlated in the time-domain waveform and completely overlap in the frequency-domain energy main frequency will the correlation strength of the current detector's physiological pulsation increase significantly.

[0037] At this point, the correlation strength of the current test subject's physiological pulsation is obtained.

[0038] S3. Calculate the effective grayscale ratio of the infrared mode in each frame of the image, and determine the reliability of the environmental mode by combining the degree of deviation of the average grayscale of the visible light mode from the preset ideal imaging brightness.

[0039] It should be noted that in practical biometric applications, RGB sensor modalities rely on the quality of ambient light. Excessive light can cause pixel saturation, leading to the loss of skin texture features, while insufficient light can induce severe shot noise, masking subtle physiological pulsations. IR sensor modalities also suffer from limited dynamic range under outdoor infrared interference. Therefore, by analyzing the grayscale distribution characteristics of the region of interest in the image in real time and evaluating the ability of each modality to identify information, this perception mechanism ensures that data from degraded modalities are no longer blindly trusted. Instead, it dynamically seeks the optimal decision-making focus based on image quality, thereby solving the problem of misjudgment caused by drastic environmental changes.

[0040] Specifically, the effective grayscale ratio of the infrared mode in each frame of the image is calculated. The effective grayscale ratio is equal to the ratio of the standard deviation of the grayscale values ​​of all pixels in the infrared mode in each frame of the image to the maximum grayscale standard deviation benchmark value calibrated by the infrared sensor.

[0041] A dynamic environment weight is constructed based on the degree to which the average grayscale value of the visible light mode in each frame deviates from the preset ideal imaging brightness. This dynamic environment weight is then used to correct the effective grayscale ratio of the infrared mode, thus determining the environment mode reliability. The environment mode reliability satisfies the expression:

[0042] In the formula, For environmental modal credibility; For the first The effective grayscale percentage of infrared modes in a frame image; For the first The mean grayscale value of all pixels in the visible light mode of a frame image; To preset the ideal imaging brightness; To take the absolute value; , This refers to the index value and total number of captured frames.

[0043] in, Reflects the first The effective grayscale ratio of the infrared mode in the frame image is used as the basis for the credibility of the physical information of the current frame image, since infrared imaging mainly relies on the reflection of the supplementary light and can extract deep skin features. The larger the ratio, the clearer the texture details and the stronger the resolution of the current frame infrared image area. Indicates the first The dynamic environment weights of a frame image reflect the current detector's position in the first frame. The degree of visible light illumination deviation in a frame image is considered. When the average grayscale of the visible light region perfectly matches the preset ideal imaging brightness, the dynamic environment weight approaches 1; if there is strong light overexposure or extreme darkness causing the actual grayscale to deviate significantly from the ideal brightness, the dynamic environment weight will be less than 1. The dynamic environment weight increases significantly, and at this point, the dynamic environment weight approaches 0. The basic credibility score is corrected by the dynamic environment weight, ensuring that under extremely poor lighting conditions, the dynamic environment weight will spontaneously penalize and lower the overall score of the frame image, playing the role of environmental filtering and security degradation, and ensuring that the finally calculated environmental modal credibility is always based on high-quality data.

[0044] It should be added that the gray values ​​of human pixels in the infrared mode cannot be completely consistent. That is, the standard deviation and the maximum standard deviation of the gray values ​​of pixels in the infrared mode are both greater than 0. The ideal imaging brightness refers to the ideal imaging brightness of most standard skin in the RGB space, which is generally around 128. In this embodiment, 128 is used.

[0045] At this point, the reliability of the current detector's environmental modality is obtained.

[0046] S4. Input all frame images into a pre-trained convolutional neural network to obtain texture probabilities; fuse physiological pulse correlation strength, environmental modality confidence and texture probabilities to determine the liveness confidence; in response to the liveness confidence being greater than a set threshold, determine that the person is alive.

[0047] It should be noted that the difficulty in determining the liveness of an object lies in distinguishing between highly realistic forgeries and real living objects. Due to the nonlinear reflective properties of forgery media such as LCD panels or photographic paper, their macroscopic texture features are fundamentally different from those of real skin. Therefore, by nonlinearly coupling the correlation strength of microscopic physiological pulsations with the credibility of macroscopic environmental modes, and introducing depth texture probability for final verification, a multi-dimensional evaluation space is constructed, thereby improving the accuracy of the final evaluation.

[0048] Specifically, the texture probability is obtained by inputting all captured frame images into a pre-trained convolutional neural network. The network extracts subtle texture features of the images through cascaded convolutional layers. At the end of the network, the extracted high-dimensional features are fused using the Softmax activation function to output the texture probability. The closer the texture probability is to 1, the more the image conforms to the feature distribution of a real living organism in the cross-modal texture dimension.

[0049] In this embodiment of the invention, the pre-trained convolutional neural network is MobileNetV3. The MobileNetV3 model is trained from a large number of real and fake live samples. The training process is existing technology and will not be described in detail here.

[0050] The liveness confidence is determined by fusing physiological pulsation correlation strength, environmental modality confidence, and texture probability; the liveness confidence satisfies the expression:

[0051] In the formula, For liveness confidence; , , For physiological pulsation correlation strength, environmental modality credibility, and texture probability; These are the preset harmonic parameters.

[0052] Among them, the construction of the liveness confidence score is based on the classic Tikhonov regularization and two-dimensional Euclidean distance theory in mathematics, and a non-linear Euclidean norm gating mechanism is constructed. It acts as a dual-gating mechanism, where the numerator term will only generate a significant positive excitation when the correlation strength of the current tester's physiological pulsation is greater and the confidence level of the environmental modality is higher. The penalty term, constructed based on Euclidean distance, reflects the suppression logic for abnormal fluctuations. Its function is to prevent the score from being affected when the correlation strength of physiological pulsations or the credibility of environmental modalities reach extreme maxima. Out of range; texture probability As a final correction for physical authenticity, this means that even if the physiological fluctuation signal passes the frequency check, if the deep network identifies that the current image has obvious screen edges or paper texture, the smaller... The value will still affect the final score. Significantly reduced.

[0053] It should be added that the harmonic parameters The regularization cutoff threshold is derived rigorously from the nonlinear decay rate of the system against highly realistic digital forgery attacks. If the harmonic parameter is less than 0.1, the denominator penalty term is too small, and the system will exhibit a microscopic amplification effect when processing the slight sensor noise generated by real liveness detection, causing the decision curve to exhibit a step-like shape, which can easily lead to false rejections by legitimate users. If the harmonic parameter is greater than 1, the regularization becomes overly smooth, and the curve becomes too flat, causing the full-score output value of real liveness detection to be excessively diluted. For example, when the harmonic parameter is 1.5, This results in insufficient security margin between the threshold and the interception defense line. In scenarios with extremely high security requirements, the preferred value in this embodiment is 0.5 to ensure that the score has a stronger suppressive effect on edge forged samples. Implementers can fine-tune the score within the range of [0.1,1] according to the miss rate index under different security levels.

[0054] Furthermore, after obtaining the liveness confidence score, a threshold is set. When the liveness confidence score is less than or equal to the set threshold, it indicates that the probability of the current detector being a live person is low. Therefore, the current detector is judged to be a non-live attack, which may be a digital forgery attack. An abnormal sample alarm is immediately sent, and the current abnormal sample evidence is retained through the interface. Conversely, if the liveness confidence score of the current detector is greater than the set threshold, the current detector is judged to be a live person, which is a normal situation and no alarm is triggered.

[0055] It should be added that the threshold is set based on statistical learning and boundary optimization of a large-scale multimodal adversarial dataset. During the system calibration and pre-training phase, a test set containing a large number of real live samples under complex working conditions and various highly realistic digital forgery samples is used. These samples are input into the nonlinear fusion model of this invention to obtain the probability distribution characteristics of the liveness confidence. By plotting the subject feature curve, minimizing the error rate is used as the optimization objective. 0.4 is selected as the optimal threshold in the distribution of scores for real and fake samples. This threshold can ensure that the vast majority of negatively correlated samples and low-correlation edge samples are effectively intercepted. Therefore, 0.4 is used as the set threshold in this embodiment. However, in actual industrial deployment, implementers can adaptively fine-tune the set threshold according to the security level of the specific application scenario.

[0056] For example, the present invention uses a facial landmark algorithm to locate the forehead region, where the skin is relatively thin and blood vessels are evenly distributed, as a physiological response sampling area. Figure 2 The curves represent in vivo multimodal physiological pulsation, with the horizontal axis representing the sampling sequence and the vertical axis representing the pulsation intensity. The visible light and infrared signals extracted from the in vivo physiological response sampling area, after bandpass filtering, exhibit highly regular periodic sinusoidal fluctuations, representing the real blood oxygenation level changes driven by the same heartbeat frequency, proving that the tested subjects have cross-modal intrinsic consistency of the in vivo circulatory system.

[0057] and Figure 3 As a non-living, multimodal physiological pulsation curve, although the attack image is visually highly deceptive, the extracted visible light and infrared signal curves are extremely chaotic, exhibiting irregular noise jumps and failing to synchronize with the microscopic fluctuations caused by deep tissue blood circulation. Due to the lack of physiological pulsation consistency in non-living media, its final evaluation value is far below the judgment threshold, thus being accurately identified and blocked, effectively defending against high-precision digital forgery attacks.

Claims

1. A multimodal liveness detection method, characterized in that, include: Acquire each frame of image containing the face of the current detector, the image containing grayscale information in visible light mode and infrared mode; For each frame of image: extract the forehead region and acquire the visible light signal sequence and infrared signal sequence of the forehead region, and take the average of all signal values ​​in each signal sequence as the pulsation intensity of each signal. The pulse intensity sequence of each signal in all frames of images is filtered, and the pulse intensity sequence of each filtered signal is obtained and subjected to fast Fourier transform to obtain the main frequency of each signal. The physiological pulse correlation strength is determined based on the difference between the main frequencies of the two signals and the correlation coefficient of the pulse intensity sequences of the two filtered time-domain signals. Calculate the effective grayscale ratio of the infrared mode in each frame of the image. Based on the degree to which the average grayscale of the visible light mode in each frame of the image deviates from the preset ideal imaging brightness, construct a dynamic environment weight. Use the dynamic environment weight to correct the effective grayscale ratio of the infrared mode and determine the reliability of the environment mode. Input all frame images into a pre-trained convolutional neural network to obtain texture probabilities; The correlation strength of physiological pulsation, the confidence of environmental modality, and the texture probability are fused to determine the liveness confidence; when the liveness confidence is greater than a set threshold, the person is determined to be alive.

2. The multimodal liveness detection method according to claim 1, characterized in that, The extraction of the forehead region and the acquisition of visible light and infrared signal sequences from the forehead region include: The forehead region is extracted and identified using a facial landmark detection algorithm; visible light and infrared signals within the forehead region are collected by two sensors, RGB and IR, respectively.

3. The multimodal liveness detection method according to claim 1, characterized in that, The acquisition of the pulsation intensity sequence of the two filtered signals includes: The pulsation intensity of each signal in all frames of images is arranged sequentially according to the acquisition time, and then filtered using a fourth-order Butterworth bandpass filter to obtain the pulsation intensity sequence of each signal after filtering; and the filtering frequency range is set to 0.7Hz to 4Hz.

4. The multimodal liveness detection method according to claim 1, characterized in that, The acquisition of the main frequency of each signal includes: Fast Fourier Transform is used to convert the pulsation intensity sequence of each filtered signal to the frequency domain, and the frequency value of the point where the energy is most concentrated in the frequency domain is identified as the main frequency of each signal.

5. The multimodal liveness detection method according to claim 1, characterized in that, The determination of physiological pulsation correlation strength includes: Calculate the Pearson correlation coefficient of the pulsation intensity sequence of the two filtered signals; calculate the ratio of the absolute difference of the dominant frequencies of the two signals to the sum of the dominant frequencies of the two signals, and input the ratio as a negative independent variable into the natural exponential function; take the maximum value of the Pearson correlation coefficient and 0, and calculate the product of the maximum value and the output of the natural exponential function to obtain the physiological pulsation correlation strength.

6. The multimodal liveness detection method according to claim 1, characterized in that, The determination of environmental modal credibility includes: For each frame of image: the effective grayscale ratio of the infrared mode is equal to the ratio of the standard deviation of the grayscale values ​​of all pixels in the infrared mode to the maximum grayscale standard deviation benchmark value calibrated by the infrared sensor; calculate the absolute difference between the mean of the grayscale values ​​of all pixels in the visible light mode and the preset ideal imaging brightness, sum the preset ideal imaging brightness and the absolute difference, and the dynamic environment weight is equal to the ratio of the preset ideal imaging brightness and the summation result; multiply the effective grayscale ratio of the infrared mode by the dynamic environment weight, and take the average of the product results of all frames of images to obtain the environmental mode confidence level.

7. The multimodal liveness detection method according to claim 1, characterized in that, The pre-trained convolutional neural network uses the MobileNetV3 model.

8. The multimodal liveness detection method according to claim 1, characterized in that, The determination of the liveness confidence level includes: The product of the physiological pulsation correlation strength and the environmental modality confidence is divided by the square root of the sum of the square of the physiological pulsation correlation strength, the square of the environmental modality confidence, and the preset harmonic parameter to obtain a first ratio; the product of the first ratio and the texture probability is used as the liveness confidence.

9. The multimodal liveness detection method according to claim 1, characterized in that, The method further includes: if the liveness confidence level is less than or equal to a set threshold, it is determined to be a non-liveness attack and an abnormal sample alarm is sent, and the current abnormal sample evidence is retained through the interface.

10. The multimodal liveness detection method according to claim 1, characterized in that, The method also includes: using a binocular synchronous camera for image acquisition.