Multi-modal face recognition method and system based on deep learning

Through multimodal data fusion and iterative optimization of deep learning methods, the problem of biometric loss in traditional face recognition technology in occluded environments is solved, and efficient identity verification and resistance to forgery attacks in medical environments are achieved.

CN120673461APending Publication Date: 2025-09-19CHONGQING UNIV
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510910857.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-02
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

Traditional facial recognition technology faces the problem of biometric loss caused by obstructions in complex application environments, especially in medical environments. Traditional methods find it difficult to restore the subcutaneous blood vessel distribution characteristics in occluded areas, and the separation of liveness detection and identity recognition cannot meet the needs of passive verification, making it vulnerable to attacks using photos and 3D masks.

Method used

A multimodal face recognition method based on deep learning is adopted. Through the synchronous acquisition of 3D structured light depth map, near-infrared subcutaneous vascular image and PPG spectrum map, a dual-branch generator is constructed to complete the occluded area. A complete face image is generated under physiological constraints. The accurate extraction and verification of biometric features are achieved through iterative optimization and adaptive loss function adjustment.

Benefits of technology

It effectively improves the success rate of face recognition in occluded conditions, resists forgery attacks, enhances the robustness and recognition accuracy of the system in low-light environments, and ensures the biological plausibility and accuracy of the generated images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120673461A_ABST
    Figure CN120673461A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-modal face recognition method and system based on deep learning, and relates to the technical field of face recognition, 3D structured light, near-infrared images and PPG blood flow signals are synchronously fused, occlusion completion is achieved through a double-branch generation network, PPG spectrum constraints generate blood vessel distribution, and biological feature authenticity is ensured; low-confidence sub-regions are divided based on a structural similarity algorithm, the weight of a loss function is dynamically adjusted, only problem regions are iteratively generated, and calculation redundancy is reduced; the depth feature of the complemented image and the blood vessel frequency spectrum main frequency band overlapping rate are extracted, the identity authenticity is doubly verified, and 3D mask and photo attacks are resisted; the performance bottleneck of traditional single-mode identification in a shielding scene is broken through, closed-loop optimization is driven through data, precision, efficiency and safety are considered, and the method is particularly suitable for high-requirement scenes such as medical treatment and finance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of face recognition technology, and in particular to a multimodal face recognition method and system based on deep learning. Background Art

[0002] Facial recognition technology, a key branch of biometric identification, has been widely used in real-world settings. Traditional methods primarily rely on visible light images or video streams for feature extraction and matching. However, these methods face multiple technical bottlenecks in complex environments. In medical settings, obstructions such as masks and respirators cause visible light image features to be lost, and traditional methods relying on a single modality (such as RGB cameras) struggle to recover key biometric features. Mainstream solutions separate liveness detection (such as blinking and mouth opening) from identity verification into separate processes, failing to meet the needs of passive verification scenarios, such as those involving anesthetized patients, and are vulnerable to attacks using photographs and 3D masks. To address these issues, existing technologies urgently need improvement. Summary of the Invention

[0003] In response to the shortcomings of the existing technology, the present invention provides a multimodal face recognition method and system based on deep learning.

[0004] In order to achieve the above object, the technical solution of the present invention is as follows:

[0005] In a first aspect, the present invention discloses a multimodal face recognition method based on deep learning, comprising the following steps:

[0006] Acquire synchronized multimodal data of the target user, the synchronized multimodal data including a 3D structured light depth map, a near-infrared subcutaneous vascular image including vascular distribution characteristics, and a PPG spectrum map; the PPG spectrum map is generated by short-time Fourier transform of the PPG blood flow pulse signal;

[0007] Positioning the occlusion area on the 3D structured light depth map to generate a binary occlusion mask;

[0008] Inputting the binary occlusion mask, the near-infrared subcutaneous vascular image, and the PPG spectrum into an occlusion completion network including a dual-branch generator to generate a complete visible light image of the face;

[0009] The completed area of ​​the complete face visible light image is divided into several sub-regions, and the similarity score between each sub-region and the unobstructed standard template is calculated using a structural similarity algorithm;

[0010] Determining whether the similarity score is lower than a preset similarity threshold, if so, determining that the confidence of the sub-region is insufficient, and adjusting the weight distribution of the loss function of the occlusion completion network;

[0011] Perform local iterative generation in this sub-region using the adjusted loss function weight distribution, recalculate the similarity score after each iteration, until the similarity threshold judgment standard or the maximum number of iterations is met, and obtain the complete visible light image of the face after iteration;

[0012] Extract the depth feature vector and vascular texture spectrum of the complete face visible light image, calculate the cosine similarity between the depth feature vector and the registered features in the preset face database; and simultaneously calculate the main frequency band overlap rate between the vascular texture spectrum and the real-time PPG spectrum;

[0013] If the cosine similarity reaches a first preset threshold and the main frequency band overlap rate reaches a second preset threshold, an identity verification success instruction is output.

[0014] In a second aspect, the present invention discloses a multimodal face recognition system based on deep learning, which is applied with the above-mentioned multimodal face recognition method based on deep learning, including:

[0015] A multimodal synchronous acquisition module, configured to acquire synchronized multimodal data of the target user, including a 3D structured light depth map, a near-infrared subcutaneous vascular image containing vascular distribution characteristics, and a PPG spectrum map; the PPG spectrum map is generated by short-time Fourier transform of the PPG blood flow pulse signal;

[0016] an occlusion detection module, configured to locate an occlusion area on the 3D structured light depth map and generate a binary occlusion mask;

[0017] A completion generation module, configured to input the binary occlusion mask, the near-infrared subcutaneous vascular image, and the PPG spectrum into an occlusion completion network comprising a dual-branch generator to generate a complete visible light image of the face;

[0018] The similarity calculation module is used to divide the completed area of ​​the complete face visible light image into several sub-areas and calculate the similarity score between each sub-area and the unobstructed standard template using a structural similarity algorithm;

[0019] A confidence evaluation module is used to determine whether the similarity score is lower than a preset similarity threshold. If so, the confidence of the sub-region is determined to be insufficient, and the weight distribution of the loss function of the occlusion completion network is adjusted;

[0020] An iterative optimization module is used to perform local iterative generation in the sub-region using the adjusted loss function weight distribution, recalculate the similarity score after each iterative generation, and obtain the complete visible light image of the face after the iteration until the similarity threshold judgment standard or the maximum number of iterations is met;

[0021] The feature extraction and cross-modal comparison module is used to extract the depth feature vector and vascular texture spectrum of the complete facial visible light image, calculate the cosine similarity between the depth feature vector and the registered features in the preset facial database, and simultaneously calculate the main frequency band overlap rate between the vascular texture spectrum and the real-time PPG spectrum;

[0022] The decision control module is configured to output an identity verification success instruction when the cosine similarity reaches a first preset threshold and the main frequency band overlap rate reaches a second preset threshold.

[0023] Compared with the prior art, the present invention has the following beneficial effects:

[0024] 1. By integrating 3D structured light, near-infrared, and PPG signal data, a dual-branch generative adversarial network is used to complete occluded areas. The PPG spectrum is used as a physiological constraint to force the generation of vascular textures to conform to biological laws, ensuring the biological plausibility of the completed areas and effectively improving the success rate of face recognition in the presence of occlusion.

[0025] 2. Geometric and physiological features are calculated independently and must meet the criteria simultaneously for verification to be successful. Forged vascular textures cannot match dynamic blood flow signals and effectively resist attacks such as photos and 3D masks. Forgers must simultaneously simulate the three-dimensional geometric structure, static vascular distribution, and dynamic blood flow pulse spectrum, which increases the technical difficulty exponentially.

[0026] 3. Dynamically adjust the threshold parameters based on historical successful verification data, combined with the adaptive correction coefficient of ambient light intensity. In low-light scenarios, the energy distribution of the main frequency bands of vascular texture and PPG spectrum is more significant, and the attenuation of visible light characteristics is compensated by adjusting the gain coefficient. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] The disclosure of the present invention is described with reference to the accompanying drawings. It should be understood that the drawings are for illustrative purposes only and are not intended to limit the scope of protection of the present invention. In the drawings, the same reference numerals are used to refer to the same components. Among them:

[0028] Figure 1 A diagram showing the steps of the method of the present invention;

[0029] Figure 2 It is a data flow diagram of the present invention;

[0030] Figure 3 Updating flow chart for the first preset threshold and the second preset threshold of the present invention;

[0031] Figure 4 It is a system module diagram of the present invention. DETAILED DESCRIPTION

[0032] It is easy to understand that according to the technical solution of the present invention, without changing the essential spirit of the present invention, a person skilled in the art can propose a variety of interchangeable structural modes and implementation modes. Therefore, the following specific embodiments and drawings are only exemplary descriptions of the technical solution of the present invention and should not be regarded as the entire invention or as a limitation or restriction of the technical solution of the present invention.

[0033] Application Overview

[0034] In traditional medical facial recognition systems, reliance on single-modal visible light data makes it difficult to effectively recover biometric features from occluded areas. When a patient wears a respiratory mask or medical patch, the occluded area exceeds a critical threshold, leading to inaccurate deep feature extraction. Traditional pixel completion algorithms are unable to reconstruct a physiologically consistent subcutaneous vascular distribution. Independently operating liveness detection modules rely on active behavioral interaction and are unable to simultaneously verify the spatiotemporal correlation between blood flow pulse signals and visible light features, significantly increasing the probability of successful static forgery attacks.

[0035] For example, in an intensive care unit identity verification scenario, patients wearing non-invasive ventilators are completely covered by a silicone mask from the nose to the jaw. Facial images captured by RGB cameras lack texture information around the mouth, and the image completion model based on a generative adversarial network can only generate false images that conform to the geometric structure but lack the distribution of capillaries. Because the physiological signal acquisition module lacks a constraint relationship with the visual completion process, the generated facial red membrane area does not carry the periodic blood flow fluctuation characteristics corresponding to the real PPG signal, resulting in a phase shift in the energy distribution of the main frequency band of the spectrum during subsequent cross-modal comparison.

[0036] If these issues are not addressed, the biometric recognition system's false positive rate under occlusion conditions will exceed the permissible range of medical device safety standards, and will fail Class III certification. Attackers could exploit this flaw by using an injection-molded mask to forge a patient's facial features and implant pre-recorded PPG signals to bypass liveness detection. Such security vulnerabilities could allow unauthorized individuals to gain access to sensitive medical data, potentially compromising patient privacy and posing risks of misoperation in clinical decision-making systems.

[0037] Faced with the above problems, this application first explores how to deeply integrate multimodal physiological features with the visual completion process. Traditional methods only consider geometric structure continuity when completing images, and do not introduce the physiological constraints of subcutaneous blood vessel distribution and blood flow signals, resulting in a lack of biometric authenticity in the generated area. In response to this, this application attempts to use vascular images and PPG spectra as input constraints for the generation network, and establish the correlation between vascular texture and blood flow signals through frequency domain feature mapping. At the same time, in response to the need for dynamic adjustment, this application studies how to adaptively adjust the optimization direction of the generation network based on the confidence evaluation results of the local area, so that the completion process can simultaneously meet the dual verification standards of geometric structure and physiological characteristics.

[0038] like Figure 1 、 Figure 2 As shown, this application proposes a multimodal face recognition method based on deep learning, which includes the following steps:

[0039] Acquire synchronized multimodal data of the target user, including 3D structured light depth map, near-infrared subcutaneous vascular image containing vascular distribution characteristics, and PPG spectrum map; PPG spectrum map is generated by short-time Fourier transform of PPG blood flow pulse signal; synchronized multimodal data refers to a data set containing three-dimensional depth information, subcutaneous vascular distribution and blood flow spectrum characteristics collected synchronously through the same clock source. Specifically, it can be achieved by collaborative collection of 3D structured light camera, near-infrared camera and PPG sensor to solve the problem of missing biometric features in traditional single visible light modality in occluded scenes.

[0040] The occluded areas of the 3D structured light depth map are located to generate a binary occlusion mask. The binary occlusion mask refers to a binary matrix that locates the occluded areas of the face through 3D depth information analysis. Specifically, it can be generated by combining depth map edge detection and region growing algorithm with morphological operations to accurately guide the generation network to reconstruct the missing areas.

[0041] The binary occlusion mask, near-infrared subcutaneous vascular image, and PPG spectrum map are input into the occlusion completion network containing a dual-branch generator to generate a complete visible light image of the face; the dual-branch generator of the occlusion completion network refers to an adversarial generation architecture that simultaneously processes image completion and physiological feature constraints. Specifically, a neural network structure with an image completion branch and a physiological constraint branch in parallel can be adopted to achieve physiological consistency reconstruction of the occluded area by fusing vascular texture and blood flow spectrum features.

[0042] The completed area of ​​the visible light image of the complete face is divided into several sub-areas, and the similarity score between each sub-area and the unobstructed standard template is calculated using the structural similarity algorithm; the structural similarity algorithm (SSIM) refers to an evaluation method that quantifies the visual similarity between the generated area and the standard template. Specifically, it can be implemented by calculating the weighted product of the three components of brightness, contrast and structural similarity, which is used to dynamically detect local generation quality and trigger iterative optimization.

[0043] Determine whether the similarity score is lower than the preset similarity threshold. If so, the confidence of the sub-region is determined to be insufficient, and the weight distribution of the loss function of the occlusion completion network is adjusted;

[0044] In this sub-region, local iterative generation is performed using the adjusted loss function weight distribution. The similarity score is recalculated after each iterative generation until the similarity threshold judgment standard or the maximum number of iterations is met, and the complete face visible light image after iteration is obtained; local iterative generation refers to the repeated generation process after weight adjustment for low-confidence sub-regions. Specifically, it can be achieved by adaptively adjusting the loss function weights and combining it with the gradient descent algorithm, and the reconstruction accuracy of the occluded area is improved through multiple rounds of local optimization.

[0045] The depth feature vector and vascular texture spectrum of the iterative complete facial visible light image are extracted, and the cosine similarity between the depth feature vector and the registered features in the preset face database is calculated; at the same time, the main frequency band overlap rate of the vascular texture spectrum and the real-time PPG spectrum is calculated; the main frequency band overlap rate refers to the degree of matching between the vascular texture spectrum and the real-time PPG spectrum in the main energy frequency band. Specifically, it can be achieved by calculating the intersection ratio of the energy distribution of the two sets of spectra within the preset frequency band, which is used to verify the consistency of the physiological activity of the biometric characteristics.

[0046] If the cosine similarity reaches a first preset threshold and the main frequency band overlap rate reaches a second preset threshold, an identity verification success instruction is output.

[0047] The core innovation of this application lies in constructing a multimodal data collaborative occlusion completion mechanism. By fusing 3D depth information, subcutaneous vascular characteristics and cross-modal correlation of blood flow spectra, visual reconstruction and physiological feature verification of occluded areas are simultaneously achieved in the generative adversarial network, effectively solving the problem of feature loss and liveness verification separation in severe occlusion scenarios in traditional methods.

[0048] This solution achieves effective completion and identity verification of occluded areas through multimodal data fusion and iterative optimization. It utilizes near-infrared subcutaneous vascular images and PPG signals as physiological constraints to ensure the biometric authenticity of the completed image. A local iterative generation mechanism further improves completion quality and recognition accuracy.

[0049] This application further proposes that the training of the occlusion completion network includes:

[0050] Construct a dual-branch structure based on a generative adversarial network. The dual-branch generator includes an image completion branch and a physiological constraint branch.

[0051] The image completion branch of the generator takes the binary occlusion mask and the near-infrared subcutaneous vascular image as input, and the discriminator compares the pixel differences between the generated image and the real unoccluded area;

[0052] The physiological constraint branch inputs the PPG spectrum map into the generator, extracts the main frequency band energy distribution through the spectral attention mechanism, and constrains the correlation coefficient between the vascular texture of the generated image and the PPG spectrum map to be no less than the preset correlation coefficient threshold.

[0053] The image completion branch employs an encoder-decoder architecture. The encoder fuses the occlusion mask with the near-infrared image to generate a latent feature vector, and the decoder reconstructs the complete visible light image based on the feature vector. The image completion branch first processes the occlusion mask and near-infrared image separately through a two-channel convolutional layer to generate a preliminary fused feature map. Residual dense blocks are then used to perform multi-scale feature extraction on the fused features, generating a high-dimensional feature tensor. The decoder progressively upsamples the high-dimensional feature tensor to the original resolution through transposed convolutional layers, outputting the completed visible light image. The physiological constraint branch feeds the PPG spectrogram into a spectral attention module, which extracts the spectral energy distribution through a short-time Fourier transform and generates normalized spectral weights using a softmax function. The spectral weights are then dot-multiplied with the intermediate feature map from the image completion branch to suppress non-physiologically relevant feature channels. During training, a physiological consistency loss function calculates the Pearson correlation coefficient between the frequency-domain energy distribution of the generated image's vascular texture and the PPG spectrogram in real time. When the correlation coefficient falls below 0.85, backpropagation is used to adjust the convolution kernel parameters of the generator. The discriminator adopts a multi-scale receptive field design to perform local discrimination on the generated image at four scales from 16×16 to 256×256. Each scale outputs a confidence mask of the corresponding resolution and performs adversarial training with the real image area.

[0054] As a preferred embodiment, the solution of this application is specifically implemented as follows:

[0055] Specifically, the dual-branch generator uses a U-Net architecture, with the encoder and decoder each consisting of four convolutional blocks. The input to the image completion branch is a 256x256x2 tensor containing a binary occlusion mask and a near-infrared subcutaneous vascular image. The input to the physiological constraint branch is a 256x1 PPG spectrum vector. The spectral attention mechanism, consisting of a fully connected layer and a softmax activation function, calculates the weights for each frequency band of the PPG spectrum. The generator output is a 256x256x3 RGB image. The discriminator uses a PatchGAN architecture and outputs a 70x70 discriminant result matrix.

[0056] During training, the generator was first trained for 100 epochs using the Adam optimizer with a learning rate of 1e-4. The discriminator was then trained for 50 epochs with a learning rate of 1e-5, keeping the generator parameters fixed. Finally, the generator and discriminator were jointly trained for 100 epochs with a learning rate of 1e-5. The batch size was set to 16, and the training data consisted of 10,000 pairs of unobstructed face images, corresponding occlusion masks, and simultaneously acquired PPG signals. The default correlation coefficient threshold was set to 0.8.

[0057] This application further proposes a loss function based on a dual-branch structure of a generative adversarial network, including a pixel reconstruction loss function and a physiological consistency loss function.

[0058] The pixel reconstruction loss function uses the L1 distance to calculate the pixel difference between the generated image and the true unobstructed area. This distance function is robust to outliers and can maintain smooth transitions of edge structures. The physiological consistency loss function measures the linear correlation between the spectral energy distribution of the generated vascular texture and the PPG spectrum using the Pearson correlation coefficient. The correlation coefficient threshold is set at 0.85 to ensure the effective transmission of physiological features. The two loss functions are combined through weighted summation during backpropagation to form a joint optimization objective. The initial weight of the pixel reconstruction loss is 1.0, and the initial weight of the physiological consistency loss is 0.7. The weight ratio is adaptively adjusted to balance the optimization strength of spatial details and physiological features.

[0059] During the generator training phase, the image completion branch first constrains the pixel accuracy of the generated image in the visible region using a pixel reconstruction loss, ensuring that the grayscale values ​​of the completed region flow seamlessly with those of adjacent regions. The physiological constraint branch simultaneously extracts the energy distribution of the main frequency band in the PPG spectrum and quantifies the degree of matching of the spectral features of the generated vascular texture using the Pearson correlation coefficient. For example, when the PPG spectrum exhibits a bimodal characteristic in the 1.5-2.5Hz range, the physiological consistency loss forces the correlation coefficient of the energy distribution of the generated image in this frequency band to be above 0.9. These two loss functions are simultaneously applied to the generator's parameter updates via a gradient descent algorithm, ensuring that the completed region maintains the geometric characteristics of the structured light depth map while conforming to the physiological characteristics of the near-infrared vascular image and PPG signal. This joint optimization mechanism effectively avoids vascular orientation distortion or PPG feature mismatch caused by single-pixel reconstruction, improving the biometric reliability of the completed image in the subsequent feature extraction phase.

[0060] Through the above technical solution, the present application achieves dual constraints on the quality and physiological characteristics of the generated image. The pixel reconstruction loss ensures the visual similarity between the generated image and the real image, while the physiological consistency loss ensures the consistency of the vascular texture in the generated image with the actual physiological signal. This dual constraint mechanism improves the authenticity and credibility of the generated image and effectively prevents the generation of forged images. At the same time, due to the use of spectrum-based physiological feature comparison, the system's robustness to changes in different lighting, posture, etc. is enhanced. In addition, the loss function design also provides a reliable evaluation indicator for subsequent iterative optimization, which is conducive to further improving the quality of the completed image.

[0061] The present application further proposes that adjusting the weight distribution of the loss function of the occlusion completion network includes: adjusting the physiological consistency loss weight to: initial value + preset adjustment amplitude × (preset similarity threshold - similarity score).

[0062] The initial value represents the baseline weight of the physiological consistency loss in the total loss, the preset adjustment range is a linear adjustment coefficient determined based on experimental data, the preset similarity threshold is a pre-set structural similarity criterion, and the similarity score is the real-time similarity calculation result between the subregion and the unobstructed standard template. The adjustment range of the physiological consistency loss weight is positively correlated with the degree to which the similarity score falls below the preset threshold. When the similarity score falls below the preset threshold, the difference will be linearly increased by the preset adjustment range.

[0063] Specifically, when the similarity score of a sub-region is lower than a preset threshold, the region is judged to have insufficient confidence. At this time, the physiological consistency loss weight is adjusted to the initial value plus the product of the preset adjustment amplitude and the score deviation, so that the weight increases with the increase of the score deviation. For example, when the preset adjustment amplitude is 0.1 and the similarity score is 0.5 lower than the threshold, the physiological consistency loss weight will increase by 0.05. This adjustment process forces the network to prioritize the optimization of physiological feature consistency in low-confidence areas, thereby guiding the generator to enhance the correlation between vascular texture and PPG spectrum during the iterative process. Through the linear relationship between weight and score deviation, differentiated optimization control of the completion area is achieved, ensuring that the strength of physiological feature constraints is dynamically matched with the local completion quality requirements.

[0064] In the occlusion completion network, a physiological consistency loss function is used to measure the consistency between the spectral energy distribution of the generated vascular texture and the PPG spectrogram. When the similarity score of a subregion falls below a preset threshold, the weight of the physiological consistency loss for that region is increased to strengthen the match between the generated image and the actual physiological features.

[0065] For example, if the initial physiological consistency loss weight is set to 0.5, the preset adjustment range is 0.1, and the preset similarity threshold is 0.8, then if the similarity score of a sub-region is 0.6, the adjusted physiological consistency loss weight is: 0.5 + 0.1 × (0.8 - 0.6) = 0.52.

[0066] Through this dynamic adjustment mechanism, differentiated loss function optimization can be performed for the generation quality of different sub-regions, thereby improving the accuracy and reliability of the overall completion effect.

[0067] Through the above technical solution, this application achieves adaptive adjustment of the weights of the loss function of the occlusion completion network. This allows for differentiated processing of areas with varying degrees of occlusion, improving the overall quality of the completed image and the consistency of physiological features. Furthermore, this solution enhances the robustness of the face recognition system to partial occlusion, improving recognition accuracy in complex application environments.

[0068] like Figure 3As shown in FIG. , this is a flowchart for updating the first preset threshold and the second preset threshold of the present application; the present application further proposes: counting the cosine similarity and main frequency band overlap rate data successfully verified within a preset time period, and updating the first preset threshold and the second preset threshold by a sliding window average algorithm, including the following steps:

[0069] Count the cosine similarity dataset and main frequency band overlap rate dataset of all successfully verified records within a preset time period;

[0070] The sliding window mean algorithm is applied to the cosine similarity dataset and the main frequency band overlapping wave dataset respectively to calculate the first arithmetic mean and the second arithmetic mean of the data in the window;

[0071] The first preset threshold is updated to the first arithmetic mean × cosine similarity attenuation coefficient, and the second preset threshold is updated to the second arithmetic mean × overlap rate gain coefficient. Both the cosine similarity attenuation coefficient and the overlap rate gain coefficient are dynamically adjusted according to the ambient light intensity.

[0072] For example, the sliding window mean algorithm uses a time series window length of 10-15 minutes, a window step of 1 minute, and each window contains at least 20 successfully verified records. The cosine similarity decay coefficient is set to 0.9-1.1, and the overlap rate gain coefficient is set to 1.1-1.4. When the ambient light sensor detects a drop in light intensity of 50 lux, the system automatically decreases the decay coefficient by 0.05 and increases the gain coefficient by 0.05. The arithmetic mean calculation excludes extreme values ​​in the first and last 5% of the window.

[0073] Specifically, the system reads the successful verification data within the last 10 minutes every 1 minute. After excluding outliers, the average cosine similarity in the window is calculated to be 0.92, and the average overlap rate of the main frequency band is 85%. Based on the current ambient light intensity of 200 lux, the cosine similarity attenuation coefficient is selected as 0.95 and the overlap rate gain coefficient is selected as 1.2. The new first preset threshold is updated to 0.92×0.95=0.874, and the second preset threshold is updated to 85%×1.2=102%. When the ambient light drops to 150 lux, the attenuation coefficient is adjusted to 0.90 and the gain coefficient is adjusted to 1.25, so that the threshold parameters are automatically calibrated with environmental changes to maintain the stability of biometric recognition.

[0074] Through the above technical solution, the present application realizes adaptive adjustment of the threshold. This improves the robustness of the face recognition system under different environmental conditions. Furthermore, by introducing the ambient light intensity as a regulating factor, the threshold adjustment is more accurately adapted to the actual application scenario. Specifically, when the lighting conditions are poor, appropriately lowering the cosine similarity threshold and increasing the main frequency band overlap rate threshold can effectively reduce the recognition misjudgment rate. On the contrary, when the lighting conditions are good, increasing the cosine similarity threshold and lowering the main frequency band overlap rate threshold can improve the security of the system. Therefore, the solution of the present application improves the environmental adaptability of the system and the user experience while ensuring recognition accuracy.

[0075] This application further proposes that the cosine similarity attenuation coefficient ranges from 0.9 to 1.1, the overlap rate gain coefficient ranges from 1.1 to 1.4, and when the ambient light intensity drops by 50 lux, the cosine similarity attenuation coefficient decreases by 0.05 and the overlap rate gain coefficient increases by 0.05.

[0076] The coefficient range is based on the correlation between light intensity and biometric recognizability, verified by experimental data. Specifically, as light intensity decreases, visible light image quality deteriorates, leading to a decrease in the reliability of cosine similarity. Lowering the attenuation coefficient relaxes the matching tolerance of visible light features, while increasing the gain coefficient enhances the weight of vascular spectral features. The coefficient's step adjustment of 0.05 is determined by the extreme value of the first-order derivative of the recognition success rate curve in the light gradient experiment, ensuring smooth parameter adjustment under light fluctuations.

[0077] For example, under the condition of an initial ambient light intensity of 500 lux, the cosine similarity attenuation coefficient is set to 1.0, and the overlap rate gain coefficient is set to 1.2. When the ambient light intensity drops to 450 lux, the cosine similarity attenuation coefficient is adjusted to 0.95, and the overlap rate gain coefficient is adjusted to 1.25. Furthermore, if the ambient light intensity continues to drop to 400 lux, the cosine similarity attenuation coefficient is adjusted to 0.9, and the overlap rate gain coefficient is adjusted to 1.3. In the scenario of a sudden drop in ambient light in the operating room, this mechanism can avoid the false rejection of legitimate users due to visible light image noise, while at the same time preventing forgery attacks by enhancing the strength of liveness feature verification. The parameter adjustment process is implemented using linear interpolation to ensure smooth threshold transitions without sudden changes during continuous lighting changes.

[0078] By dynamically adjusting these two coefficients, the system can adapt to face recognition needs in different lighting environments. In poor lighting conditions, the accuracy and robustness of recognition are balanced by reducing the cosine similarity requirement and increasing the main frequency band overlap requirement.

[0079] Through the above technical solution, this application achieves adaptive adjustment of the facial recognition system to different lighting environments. As lighting conditions change, the system can dynamically adjust the recognition threshold to maintain a high recognition accuracy rate. Furthermore, by introducing an overlap rate gain coefficient, the system's reliance on physiological characteristics is enhanced, improving anti-counterfeiting capabilities. This adaptive mechanism enables the system to maintain stable recognition performance in a variety of complex lighting environments, significantly improving the reliability and practicality of facial recognition technology in practical applications.

[0080] This application further proposes a method for collecting synchronous multimodal data: controlling the collection timing of the 3D structured light camera, near-infrared camera and PPG sensor through a single clock source, and the collection time deviation of each modal data is less than a preset number of bits.

[0081] A single clock source, using a hardware-level clock synchronization chip, outputs a unified pulse signal to the three sensors to trigger data acquisition. The pulse signal frequency is set to 100Hz, and the trigger error is controlled within ±0.1ms. The preset bit count is set to microsecond accuracy based on data fusion requirements, specifically ensuring that the timestamp difference between the data frames of each modality does not exceed 500μs. Under this clock source synchronization mechanism, the three sensors simultaneously initiate data acquisition within every 100ms period. The 3D structured light camera emits a coded light spot to acquire depth maps, the near-infrared camera illuminates the skin surface with an 850nm wavelength to capture vascular images, and the PPG sensor detects blood flow signals using a photodiode. Because the time deviation of the clock source trigger signal is controlled within 500μs, the depth map, vascular image, and PPG signal remain strictly synchronized in the temporal dimension, preventing spatial registration errors between modalities caused by minor facial movements. When the data enters the occlusion completion network, the subcutaneous vascular distribution and PPG spectral features acquired at the same time are accurately aligned, ensuring that the vascular texture of the generated image is temporally and spatially consistent with the physiological signal spectrum. This solution reduces the complexity of the software alignment algorithm through hardware-level synchronization. In the mask occlusion scenario, the time-aligned multimodal data can provide the generator with accurate physiological feature constraints, so that the generation of vascular texture in the completed area conforms to the actual hemodynamic laws.

[0082] Through the above technical solution, this application achieves high-precision synchronous acquisition of multimodal data. Specifically, by controlling multiple acquisition devices with a single clock source, drift and cumulative errors between multiple independent clock sources are avoided. At the same time, the use of a high-speed data acquisition card to record timestamp information further ensures the accuracy of data synchronization. This high-precision synchronous acquisition method provides a reliable data foundation for subsequent multimodal fusion recognition, effectively improving the accuracy and robustness of the recognition algorithm.

[0083] This application further proposes that the extraction of deep feature vectors includes: using a pre-trained ResNet-34 model to extract the feature vectors of the complete facial visible light image, and performing L2 normalization on the feature vectors; the registered feature vectors in the preset face database are extracted through the same model and stored as floating-point tensors.

[0084] The pre-trained ResNet-34 model achieves deep feature extraction through a residual learning structure. Its 34-layer convolutional network architecture can capture the global structure and local detail features of facial images. L2 normalization processing maps the feature vector to the unit hypersphere space, eliminating the vector modulus fluctuation caused by differences in illumination intensity. Floating-point tensor storage uses a 32-bit single-precision format to ensure the consistency of the numerical accuracy of the registered feature vector and the real-time extracted vector. When the completed visible light image of the complete face is input into the ResNet-34 model, it undergoes five stages of downsampling operations and outputs a feature vector of dimension 512. The jump connection structure in the third stage enhances the ability to capture edge features of the completed area. L2 normalization is used to divide the value of each dimension of the feature vector by the vector modulus, so that the cosine similarity calculation is converted into the angle measure between vectors, avoiding the similarity calculation deviation caused by pixel intensity fluctuations during the image completion process. The preset face database uses the same ResNet-34 model parameters as the recognition end. During the registration phase, normalized floating-point tensors are generated. Each identity verification step directly uses these tensor data for matrix multiplication, enabling high-speed feature matching. This solution reduced the feature matching error rate to 3.2% in a test set with a mask occlusion rate of 60%, validating its effectiveness.

[0085] As a preferred embodiment, the solution of the present application is specifically implemented as follows: the ResNet-34 model is used in the extraction process of the deep feature vector, and the model is migrated to the face recognition task after pre-training on the ImageNet dataset. The input complete face visible light image is adjusted to a resolution of 224×224 pixels, and is input into the convolutional layer of ResNet-34 after mean normalization. The last fully connected layer of the model is removed, and a feature vector with a dimension of 512 is output. The feature vector is L2 normalized to constrain the values ​​of each dimension to the unit sphere space. The registration feature vector in the preset face database is extracted using the same model architecture. During the registration stage, the user's unobstructed frontal image is collected and input into the ResNet-34 model after preprocessing. The output 512-dimensional feature vector is stored in the database index table in 32-bit floating-point tensor format.

[0086] Through the above technical solution, this application effectively solves the technical problem of increased matching errors caused by inconsistent feature spaces in multimodal face recognition systems. By uniformly adopting pre-trained deep neural networks for feature extraction and implementing standardized vector processing, the consistency of feature distribution between real-time collected data and registered data is ensured. This technical approach significantly improves the robustness of feature vectors to local area loss in mask occlusion scenarios, allowing the completed visible light image to be accurately mapped to the vector space of registered features, thereby improving the reliability of cross-modal biometric matching.

[0087] The present application further proposes triggering a dynamic re-collection mechanism when the cosine similarity or main frequency band overlap rate data does not reach a corresponding threshold, including:

[0088] Control the six-degree-of-freedom robotic arm to adjust the camera acquisition angle and re-collect multimodal data within the range of ±30° horizontally and ±15° vertically;

[0089] The data will be re-collected for a new round of recognition process.

[0090] As a preferred embodiment, the solution of the present application is specifically implemented as follows: when it is detected during the identity verification process that the cosine similarity or the main frequency band overlap rate does not reach the preset threshold, the dynamic re-acquisition mechanism is triggered. By controlling the rotation axis and translation axis of the six-degree-of-freedom robotic arm to be linked, the camera of the multimodal acquisition device is deflected in the horizontal and vertical directions, and multi-angle data acquisition is completed according to the preset path. The adjusted multimodal data is transmitted to the completion generation module for image reconstruction, and the vascular texture spectrum extraction and cross-modal comparison process are re-executed to generate a new verification result. If the re-acquired data meets the preset threshold condition, an identity verification success instruction is output; if the condition is still not met, multiple rounds of acquisition angle adjustment are continuously triggered until the maximum number of attempts is reached.

[0091] Through the above technical solution, this application effectively solves the problem of incomplete biometric data caused by partial occlusion or acquisition angle deviation. By dynamically adjusting the acquisition angle, it can adaptively acquire multimodal data from unobstructed facial areas in passive verification scenarios, enhancing the simultaneous verification capabilities of liveness detection and identity recognition. This mechanism can avoid feature extraction failures caused by restricted user posture, improving the robustness and fault tolerance of identity verification systems in complex medical environments.

[0092] like Figure 4 As shown, the present application further proposes a multimodal face recognition system based on deep learning, including a multimodal synchronous acquisition module, an occlusion detection module, a completion generation module, a similarity calculation module, a confidence assessment module, an iterative optimization module, a feature extraction and cross-modal comparison module, and a decision control module.

[0093] The multimodal synchronous acquisition module controls the 3D structured light camera, near-infrared camera and PPG sensor through a single clock source, reducing the acquisition time deviation to less than 1 millisecond. Hardware-level clock synchronization ensures the spatiotemporal alignment of the 3D depth map, near-infrared image and PPG signal, eliminating time drift of multi-source data.

[0094] The occlusion detection module uses an edge detection algorithm to locate the occlusion contour of the 3D structured light depth map and generates a binary mask with a pixel resolution of 512×512; it performs morphological analysis on the concave features of the nose bridge area in the depth map to generate accurate mask occlusion area markers.

[0095] The dual-branch generator of the completion generation module includes an image completion branch and a physiological constraint branch. The image completion branch inputs the mask and near-infrared image to generate the initial completion result. The physiological constraint branch extracts the energy distribution of the 0.5-5Hz main frequency band of the PPG spectrum through the spectral attention mechanism. When the dual-branch structure generates visible light images, the physiological constraint branch forces the correlation coefficient between the vascular texture and the PPG spectrum to be no less than 0.75, ensuring that the completion area conforms to physiological characteristics.

[0096] The similarity calculation module divides the completed area into several grid cells and uses the SSIM algorithm to evaluate the local structural similarity between each area and the standard template.

[0097] When the confidence assessment module detects that the SSIM value of the eye area is lower than the threshold, it triggers the weight adjustment mechanism to strengthen the physiological constraints of vascular texture.

[0098] The iterative optimization module adopts a region-restricted adversarial training strategy, fine-tuning parameters only on low-confidence sub-regions to avoid degradation of global generation quality.

[0099] The feature extraction and cross-modal comparison module uses the ResNet-34 model to extract L2-normalized 2048-dimensional feature vectors, and the main frequency band overlap rate is calculated using Hamming window weighted spectral integration; the feature extraction and cross-modal comparison module uses a dual-factor verification mechanism of cosine similarity and main frequency band overlap rate.

[0100] The decision control module uses the FPGA chip to realize real-time processing of threshold judgment and instruction output; after the dual thresholds are met, the decision control module outputs the identity verification success instruction.

[0101] Through the above technical solution, this application can effectively solve the problem of missing biometric features caused by facial occlusion in medical scenarios. Through multimodal data fusion and dynamic weight adjustment mechanisms, it can achieve simultaneous verification of liveness characteristics while maintaining identity recognition accuracy, avoiding the security vulnerabilities caused by the step-by-step processing in traditional methods. The dual-branch structure of the generative adversarial network ensures that the completed image conforms to the actual blood flow characteristics through physiological spectrum constraints, suppressing adversarial sample attacks and improving the reliability and security of the system in passive verification scenarios.

[0102] The technical scope of the present invention is not limited to the contents of the above description. Those skilled in the art can make various deformations and modifications to the above embodiments without departing from the technical idea of ​​the present invention, and these deformations and modifications should all fall within the protection scope of the present invention.

Claims

1. A multimodal face recognition method based on deep learning, characterized by: The steps include: Acquire synchronized multimodal data of the target user, the synchronized multimodal data including a 3D structured light depth map, a near-infrared subcutaneous vascular image including vascular distribution characteristics, and a PPG spectrum map; the PPG spectrum map is generated by short-time Fourier transform of the PPG blood flow pulse signal; Positioning the occlusion area on the 3D structured light depth map to generate a binary occlusion mask; Inputting the binary occlusion mask, the near-infrared subcutaneous vascular image, and the PPG spectrum into an occlusion completion network including a dual-branch generator to generate a complete visible light image of the face; The completed area of ​​the complete face visible light image is divided into several sub-regions, and the similarity score between each sub-region and the unobstructed standard template is calculated using a structural similarity algorithm; Determining whether the similarity score is lower than a preset similarity threshold, if so, determining that the confidence of the sub-region is insufficient, and adjusting the weight distribution of the loss function of the occlusion completion network; Perform local iterative generation in this sub-region using the adjusted loss function weight distribution, recalculate the similarity score after each iteration, until the similarity threshold judgment standard or the maximum number of iterations is met, and obtain the complete visible light image of the face after iteration; Extract the depth feature vector and vascular texture spectrum of the complete face visible light image, calculate the cosine similarity between the depth feature vector and the registered features in the preset face database; and simultaneously calculate the main frequency band overlap rate between the vascular texture spectrum and the real-time PPG spectrum; If the cosine similarity reaches a first preset threshold and the main frequency band overlap rate reaches a second preset threshold, an identity verification success instruction is output.

2. The multimodal face recognition method based on deep learning according to claim 1, characterized in that: The training of the occlusion completion network includes: Construct a dual-branch structure based on a generative adversarial network. The dual-branch generator includes an image completion branch and a physiological constraint branch. The image completion branch of the generator takes the binary occlusion mask and the near-infrared subcutaneous blood vessel image as input, and the discriminator compares the pixel differences between the generated image and the real unoccluded area; The physiological constraint branch inputs the PPG spectrum map into the generator, extracts the main frequency band energy distribution through the spectrum attention mechanism, and constrains the correlation coefficient between the vascular texture of the generated image and the PPG spectrum map to be no less than a preset correlation coefficient threshold.

3. The multimodal face recognition method based on deep learning according to claim 2, characterized in that: The loss function of the dual-branch structure based on the generative adversarial network includes: Pixel reconstruction loss function, which calculates the L1 distance between the generated image and the real unobstructed area; The physiological consistency loss function calculates the Pearson correlation coefficient between the spectrum energy distribution of the generated vascular texture and the PPG spectrum map.

4. The multimodal face recognition method based on deep learning according to claim 1, characterized in that: The adjusting the weight distribution of the loss function of the occlusion completion network includes: adjusting the physiological consistency loss weight to: initial value+preset adjustment amplitude×(preset similarity threshold-similarity score).

5. The multimodal face recognition method based on deep learning according to claim 1, characterized in that: Also includes: Counting the cosine similarity and main frequency band overlap rate data successfully verified within a preset time period, and updating the first preset threshold and the second preset threshold by a sliding window average algorithm, includes the following steps: Count the cosine similarity dataset and main frequency band overlap rate dataset of all successfully verified records within a preset time period; Apply the sliding window mean algorithm to the cosine similarity dataset and the main frequency band overlap rate dataset respectively, and calculate the first arithmetic mean and the second arithmetic mean of the data in the window; The first preset threshold is updated to the first arithmetic mean×cosine similarity attenuation coefficient, and the second preset threshold is updated to the second arithmetic mean×overlap rate gain coefficient. The cosine similarity attenuation coefficient and the overlap rate gain coefficient are both dynamically adjusted according to the ambient light intensity.

6. The multimodal face recognition method based on deep learning according to claim 5, characterized in that: The cosine similarity attenuation coefficient has a value range of 0.9-1.1, and the overlap rate gain coefficient has a value range of 1.1-1.

4. When the ambient light intensity decreases by 50 lux, the cosine similarity attenuation coefficient decreases by 0.05, and the overlap rate gain coefficient increases by 0.

05.

7. The multimodal face recognition method based on deep learning according to claim 1, characterized in that: The synchronous multimodal data is collected in the following manner: The acquisition timing of the 3D structured light camera, near-infrared camera and PPG sensor is controlled by a single clock source, and the acquisition time deviation of each modality data is less than the preset number of bits.

8. The multimodal face recognition method based on deep learning according to claim 1, characterized in that: The extraction of the depth feature vector includes: Extracting feature vectors from the complete face visible light image using a pre-trained ResNet-34 model, and performing L2 normalization on the feature vectors; The registered feature vectors in the preset face database are extracted through the same model and stored as floating-point tensors.

9. The multimodal face recognition method based on deep learning according to claim 1, characterized in that: When the cosine similarity or main frequency band overlap rate data does not reach the corresponding threshold, a dynamic re-collection mechanism is triggered, including: Control the six-degree-of-freedom robotic arm to adjust the camera acquisition angle and re-collect multimodal data within the range of ±30° horizontally and ±15° vertically; The data will be re-collected for a new round of recognition process.

10. A multimodal face recognition system based on deep learning, characterized by: The multimodal face recognition method based on deep learning according to any one of claims 1 to 9 is applied, comprising: A multimodal synchronous acquisition module, configured to acquire synchronized multimodal data of the target user, including a 3D structured light depth map, a near-infrared subcutaneous vascular image containing vascular distribution characteristics, and a PPG spectrum map; the PPG spectrum map is generated by short-time Fourier transform of the PPG blood flow pulse signal; an occlusion detection module, configured to locate an occlusion area on the 3D structured light depth map and generate a binary occlusion mask; A completion generation module, configured to input the binary occlusion mask, the near-infrared subcutaneous vascular image, and the PPG spectrum into an occlusion completion network comprising a dual-branch generator to generate a complete visible light image of the face; The similarity calculation module is used to divide the completed area of ​​the complete face visible light image into several sub-areas and calculate the similarity score between each sub-area and the unobstructed standard template using a structural similarity algorithm; A confidence evaluation module is used to determine whether the similarity score is lower than a preset similarity threshold. If so, the confidence of the sub-region is determined to be insufficient, and the weight distribution of the loss function of the occlusion completion network is adjusted; An iterative optimization module is used to perform local iterative generation in the sub-region using the adjusted loss function weight distribution, recalculate the similarity score after each iterative generation, and obtain the complete visible light image of the face after the iteration until the similarity threshold judgment standard or the maximum number of iterations is met; The feature extraction and cross-modal comparison module is used to extract the depth feature vector and vascular texture spectrum of the complete facial visible light image, calculate the cosine similarity between the depth feature vector and the registered features in the preset facial database, and simultaneously calculate the main frequency band overlap rate between the vascular texture spectrum and the real-time PPG spectrum; The decision control module is configured to output an identity verification success instruction when the cosine similarity reaches a first preset threshold and the main frequency band overlap rate reaches a second preset threshold.

Citation Information

Cited By

  • Face auditing method and system based on feature vector comparison

    CN120976996A

  • A face verification method and system based on feature vector comparison

    CN120976996B

  • Face image anti-counterfeiting detection method and device based on frequency spectrum reconstruction

    CN121354195A