A multi-modal dual-camera face depth detection method and system

CN121096003BActive Publication Date: 2026-09-22SHENZHEN YUDUN TIMES ELECTRONICS CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511292616.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-11
Publication Date
2026-09-22
Estimated Expiration
2045-09-11

AI Technical Summary

Technical Problem

[0005]此外,传统活体检测方法通常侧重于面部的运动或颜色特征,忽略了皮肤纹理、热分布等细节特征的利用

Benefits of technology

[0018]本发明提供一种基于多模态双摄人脸深度检测方法及系统,所述方法通过S1:通过多模态双摄采集原始图像组,依据相机几何模型与原始图像组推导极线约束关系并计算视差偏移值,对原始图像组进行插值变换后输出对齐图像集;S2:基于对齐图像集进行跨模态联合特征提取,通过计算多模态拓扑张力因子和特征图几何扩张度,融合生成模态协调膨胀系数,完成初级异常筛查;S3:依据对齐图像集生成人脸结构三维点云,结合时序分析计算多模态评估指标;S4:基于对齐图像集推导皮肤真实性指数,融合模态协调膨胀系数、多模态评估指标和皮肤真实性指数,通过多参数加权决策函数计算真实性置信度评分,进行真实性判别,产生的有益效果包括:

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121096003B_ABST
    Figure CN121096003B_ABST
Patent Text Reader

Abstract

The application discloses a kind of based on multimodal dual camera face depth detection method and system, it is related to image recognition field.The method includes: S1: through multimodal dual camera synchronous acquisition image information calculation parallax offset value, interpolation output alignment image set;S2: based on alignment image set by calculating multimodal topological tension factor and feature map geometric expansion, fusion generates modal coordination inflation coefficient and carries out preliminary judgment;S3: according to alignment image set generates face structure three-dimensional point cloud, calculates multimodal evaluation index;S4: according to alignment image set calculation skin authenticity index, combine modal coordination inflation coefficient, multimodal evaluation index and skin authenticity index construct authenticity confidence score function, carries out authenticity discrimination.By comprehensive evaluation facial structure, texture and heat distribution, effectively improve the accuracy and robustness of liveness detection, adapt to the high security requirement under a variety of environments.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image recognition, specifically to a method and system for face depth detection based on multimodal dual-camera setup. Background Technology

[0002] Facial recognition and liveness detection technologies have wide applications in the field of biometrics. However, existing technologies mainly rely on single-modal images, such as RGB, infrared, or depth images. These methods often face significant challenges under environmental conditions such as changes in lighting, facial expressions, and occlusion. RGB images are generally greatly affected by changes in lighting and angle in face recognition; while infrared images can work effectively in low-light environments, they are insufficient in capturing facial details; and depth images provide three-dimensional spatial information, but are weak in processing facial expressions and subtle changes.

[0003] To overcome these problems, facial recognition technology based on multimodal information fusion has gradually become a research hotspot in recent years. By combining multiple modal information such as RGB, infrared, and depth images, the advantages of different modalities can be leveraged to improve the accuracy and robustness of recognition. Especially in complex environments, multimodal fusion can effectively improve the accuracy of facial recognition and solve common problems such as lighting and expression changes.

[0004] Currently, multimodal image fusion methods mainly focus on image alignment, feature extraction, and deep learning network fusion. Although some progress has been made, how to improve detection accuracy while ensuring real-time performance, especially in dynamic environments, remains an urgent issue to be addressed.

[0005] Furthermore, traditional liveness detection methods typically focus on facial movement or color features, neglecting the utilization of detailed features such as skin texture and thermal distribution. In recent years, researchers have begun to explore liveness detection methods based on biological features such as skin microstructure and thermal distribution, which provides a new direction for improving the authenticity of facial recognition.

[0006] Against this backdrop, combining multimodal image information for deep face detection and liveness recognition can provide more accurate and stable technical support in various environments, and is particularly suitable for high-requirement scenarios such as security monitoring and identity verification. Summary of the Invention

[0007] In view of the shortcomings of the prior art described above, the purpose of this invention is to provide a multimodal dual-camera face depth detection method and system to solve the above-mentioned technical problems.

[0008] To achieve the above objectives, the present invention provides the following technical solution: a multimodal dual-camera face depth detection method, comprising: S1: Acquire the original image set through multimodal dual cameras, derive the epipolar constraint relationship based on the camera geometric model and the original image set, calculate the disparity offset value, and output the aligned image set after interpolation transformation of the original image set; S2: Based on the aligned image set, cross-modal joint feature extraction is performed. By calculating the multimodal topological tension factor and the geometric dilatation of the feature map, the modal coordination expansion coefficient is fused to complete the primary anomaly screening. S3: Generate a 3D point cloud of the face structure based on the aligned image set, and calculate multimodal evaluation indicators by combining temporal analysis; S4: Based on the aligned image set, the skin authenticity index is derived, and the modal compatibility dilation coefficient, multimodal evaluation index and skin authenticity index are integrated. The authenticity confidence score is calculated through a multi-parameter weighted decision function to determine authenticity.

[0009] The present invention is further configured such that S1 includes: The original image includes: RGB image, depth image and infrared image of the face region, which are simultaneously acquired by an RGB-D camera and a separate IR camera; Based on the geometric information of the depth image and the camera calibration parameters, the pixel-level transformation relationship from the infrared image to the RGB image coordinate system is calculated and set as the disparity offset value; The infrared image is transformed based on the parallax offset value to obtain a corrected infrared image that is aligned with the RGB image space; The corrected infrared image is mapped to the coordinate system of the RGB image using an interpolation algorithm, generating a registered infrared image that is aligned with the pixels of the RGB image and the depth image. An aligned image set is constructed by combining RGB images, depth images, and registered infrared images.

[0010] The present invention is further configured such that S2 includes: Based on the aligned image set, the mid-level modal feature maps of each modality are extracted using a feature encoder. Calculate the second-order differential feature differences between corresponding positions in each middle-layer semantic feature map, and construct a modal difference topological tensor to characterize the changing tension of local geometric structure; By combining the consistency of gradient directions between modes, the modal difference topological tensor is weighted and normalized to obtain the tension value between mode pairs; By aggregating and nonlinearly normalizing the tension values ​​between all modal pairs in the spatial domain, the multimodal topological tension factor is finally obtained.

[0011] The present invention is further configured to generate a multi-scale feature map set by performing a multi-scale transformation operation on the mid-level modal feature map; Based on a multi-scale feature map set, key response regions identified by edge detection algorithms are extracted from the feature layer on a scale-by-scale basis. Calculate the local structural response characteristics within the critical response region and construct a structural differential anisotropy index that reflects local structural changes. Based on multi-scale feature maps, a Laplacian convolution kernel is introduced to calculate the high-frequency perturbation response intensity of key response regions; By combining the structural differential anisotropy index with the high-frequency perturbation response intensity, the local geometric expansion response characteristics at the current scale are obtained; Within the key response regions of feature maps at all scales, the local geometric expansion response features are accumulated and aggregated across scales and spatial domains to obtain the feature map geometric expansion degree used to characterize the amplitude of structural perturbation and the trend of edge expansion.

[0012] The present invention is further configured to integrate the multimodal topological tension factor and the feature map geometric expansion to form a modal coordination expansion coefficient that comprehensively evaluates structural complexity and modal alignment. The modal coordination dilatation coefficient is judged based on a preset preliminary judgment threshold. When the modal coordination dilatation coefficient is lower than the preliminary judgment threshold, it is judged as a non-real face.

[0013] The present invention is further configured such that S3 includes: Based on the depth image of the aligned image set, back projection calculation is performed using RGB image intrinsic parameters to generate a 3D point cloud of the face structure corresponding to the RGB camera coordinates; Based on the three-dimensional point cloud of the human face structure, the temporal information of the infrared image is fused to calculate multimodal evaluation indicators including structural restoration, dynamic topological coherence and infrared perturbation consistency. The structural restoration degree is constructed by comparing the geometric deviation between the 3D point cloud of the human face structure and the point cloud model reconstructed based on neural networks.

[0014] The present invention is further configured to calculate density weights based on the density of the local neighborhood of the three-dimensional point cloud of the human face structure; Based on the 3D point cloud of the face structure in continuous frames, the iterative nearest point registration algorithm is used to analyze the temporal motion trajectory of the corresponding points, and the degree of structural deformation is calculated by combining the density weighting coefficient to construct dynamic topological coherence. Based on the local gradient changes of dynamic sensitive regions in infrared images, the edge gradient operator is used to extract the perturbation differences of key regions between adjacent frames and calculate the infrared perturbation consistency.

[0015] The present invention is further configured such that S4 includes: Based on the 3D point cloud of the human face structure, a skin texture undulation index is constructed by calculating the degree of difference in local normal vectors in the neighborhood of each pixel. The facial thermal intensity distribution is extracted based on the registered infrared images in the aligned image set. The temperature fluctuation amplitude within each preset facial region is calculated, and the average amplitude of temperature fluctuation in each region is calculated to construct the skin thermal distribution uniformity index. A skin authenticity index is constructed by integrating two types of indicators: skin texture undulation index and skin thermal distribution uniformity index.

[0016] The present invention is further configured to construct a confidence score function for authenticity by substituting a multimodal evaluation index, a modal coordination expansion coefficient, and a skin authenticity index into a preset multi-parameter scoring formula. The authenticity of a face is determined based on the output of the authenticity confidence score function. If the authenticity confidence score is greater than the judgment threshold, the face is judged as a real face; otherwise, it is judged as a non-real face.

[0017] The present invention also provides a multimodal dual-camera face depth detection system, the system comprising: Image acquisition and registration module: Simultaneously acquires RGB images, depth images and infrared images through multimodal dual cameras, calculates the parallax offset value using the depth image, corrects the epipolar lines of the infrared image to the RGB coordinate system, and outputs an aligned image set after interpolation; Joint feature extraction module: Extracts joint features based on aligned image sets, and performs preliminary judgment by calculating multimodal topological tension factor and feature map geometric dilatation and fusing them to generate modal coordination dilatation coefficient; 3D face image modeling module: Generates 3D point cloud of face structure based on aligned image set and calculates multimodal evaluation index; Liveness detection and authenticity assessment module: Calculates the skin authenticity index based on the aligned image set, and constructs an authenticity confidence scoring function by combining the modal compatibility dilation coefficient, multimodal evaluation index and skin authenticity index to assess authenticity.

[0018] This invention provides a multimodal dual-camera face depth detection method and system. The method comprises: S1: acquiring a set of original images using a multimodal dual-camera system; deriving epipolar constraint relationships and calculating disparity offset values ​​based on the camera geometry model and the original image set; and outputting an aligned image set after interpolation transformation of the original image set; S2: performing cross-modal joint feature extraction based on the aligned image set; calculating the multimodal topological tension factor and feature map geometric dilatation; fusing these to generate a modal coordination dilatation coefficient; and completing initial anomaly screening; S3: generating a 3D point cloud of the face structure based on the aligned image set; and calculating multimodal evaluation indicators using temporal analysis; and S4: deriving a skin authenticity index based on the aligned image set; fusing the modal coordination dilatation coefficient, multimodal evaluation indicators, and the skin authenticity index; and calculating an authenticity confidence score using a multi-parameter weighted decision function for authenticity determination. The beneficial effects include: Multimodal fusion enhances recognition accuracy: By simultaneously acquiring RGB, depth, and infrared images, and utilizing the depth image for parallax mapping and the infrared image for epipolar correction, more facial feature information can be obtained. Multimodal image fusion not only solves the performance degradation problem of single-modal imaging under conditions of illumination, angle, and occlusion, but also improves the recognition accuracy of microscopic features such as facial expression changes and skin texture details.

[0019] Efficient 3D modeling and perturbation response analysis of facial structure: This invention generates high-precision 3D point clouds of faces based on aligned image sets, and through the calculation of multimodal evaluation indicators, it can effectively capture subtle changes and irregularities in facial structure, providing a more detailed basis for facial authenticity judgment and improving the accuracy of liveness detection.

[0020] Integrated Application of Skin Features and Thermal Distribution Information: This invention introduces the skin texture undulation index and the skin thermal distribution uniformity index, combined with the modal compatibility expansion coefficient and multimodal evaluation indicators, to comprehensively assess the realism of the skin. This innovative index not only considers changes in skin surface texture and color difference but also integrates thermal imaging technology, enabling the system to maintain high liveness detection accuracy in various environments.

[0021] The above description is only an overview of the technical solution of this application. In order to better understand the technical means of this application and to implement it in accordance with the contents of the specification, and to make the above and other objects, features and advantages of this application more obvious and understandable, the following are specific embodiments of this application. Attached Figure Description

[0022] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. In the drawings: Figure 1 A flowchart illustrating a multimodal dual-camera face depth detection method is shown as an exemplary embodiment of the present invention. Figure 2 This is a schematic diagram of a multimodal dual-camera face depth detection system, which is an exemplary embodiment of the present invention. Detailed Implementation

[0023] The embodiments of the present invention will be described below with reference to the accompanying drawings and preferred embodiments. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be understood that the preferred embodiments are only for illustrating the present invention and not for limiting the scope of protection of the present invention.

[0024] It should be noted that the illustrations provided in the following embodiments are only schematic representations of the basic concept of the present invention. Therefore, the drawings only show the components related to the present invention and are not drawn according to the actual number, shape and size of the components in the actual implementation. In the actual implementation, the form, quantity and proportion of each component can be arbitrarily changed, and the layout of the components may also be more complex.

[0025] In the following description, numerous details are explored to provide a more thorough explanation of embodiments of the invention. However, it will be apparent to those skilled in the art that embodiments of the invention may be practiced without these specific details. In other embodiments, well-known structures and devices are shown in block diagram form rather than in detail to avoid obscuring embodiments of the invention.

[0026] Example 1: A multimodal dual-camera face depth detection method, such as Figure 1 As shown, it includes: S1: Acquire the original image set through multimodal dual cameras, derive the epipolar constraint relationship based on the camera geometric model and the original image set, calculate the disparity offset value, and output the aligned image set after interpolation transformation of the original image set; S2: Based on the aligned image set, cross-modal joint feature extraction is performed. By calculating the multimodal topological tension factor and the geometric dilatation of the feature map, the modal coordination expansion coefficient is fused to complete the primary anomaly screening. S3: Generate a 3D point cloud of the face structure based on the aligned image set, and calculate multimodal evaluation indicators by combining temporal analysis; S4: Based on the aligned image set, the skin authenticity index is derived, and the modal compatibility dilation coefficient, multimodal evaluation index and skin authenticity index are integrated. The authenticity confidence score is calculated through a multi-parameter weighted decision function to determine authenticity.

[0027] The present invention is further configured such that S1 includes: The original image includes: RGB image, depth image and infrared image of the face region, which are simultaneously acquired by an RGB-D camera and a separate IR camera; Based on the geometric information of the depth image and the camera calibration parameters, the pixel-level transformation relationship from the infrared image to the RGB image coordinate system is calculated and set as the disparity offset value; The infrared image is transformed based on the parallax offset value to obtain a corrected infrared image that is aligned with the RGB image space; The corrected infrared image is mapped to the coordinate system of the RGB image using an interpolation algorithm, generating a registered infrared image that is aligned with the pixels of the RGB image and the depth image. An aligned image set is constructed by combining RGB images, depth images, and registered infrared images. Specifically, an RGB-D camera is used to acquire RGB and depth images, while a separate IR camera acquires infrared images. The parallax offset value represents the pixel difference required to map a pixel in an infrared image to its corresponding position in an RGB image on the image plane. It is a displacement value per pixel. The specific calculation logic is as follows: where is the parallax offset value; is the depth image, which is the image captured by the RGB-D camera, representing the depth distance from the object to the camera; is the camera focal length, which is related to the imaging size and distortion correction, and is obtained by the camera's intrinsic parameters at the factory; is the camera distance, which is the physical distance between the infrared and RGB cameras; is the depth map gradient, representing the degree of change in the depth image, used to capture edge information, and is obtained by calculating the first derivative of the depth image pixels; is the edge enhancement coefficient, used to enhance the correction accuracy of the infrared image at object boundaries and avoid blur fusion. Its value ranges from 0.5 to 2.0, with a default value of 1.0. Higher values ​​result in higher processing accuracy, and specific values ​​can be adaptively modified according to the actual deployment scenario; is the L2 norm of the depth map gradient, used to measure edge strength. A corrected infrared image, coplanar with the RGB image, is obtained by offsetting the pixel values ​​of the infrared image based on the parallax offset value. The corrected infrared image is then remapped according to the depth information using a bilinear interpolation function to make it correspond to the structure of the RGB image. The specific calculation logic is as follows: ,in, To register infrared images; These are the pixel coordinates of the RGB image; This is the disparity offset value corresponding to the pixel coordinates of the RGB image, used to correct the pixels in the infrared image to the coordinates of the RGB image; It is a bilinear interpolation function used to estimate the grayscale values ​​of infrared images at non-integer coordinate positions, ensuring that the aligned image is continuous, smooth, and free of jagged edges. Infrared images are raw, unprocessed infrared images captured by an infrared camera. The aligned image set contains images in three modalities: RGB images. Depth images Register infrared images By aligning the three modal images to the same coordinate space, a unified multimodal input image triplet is constructed. This addresses the difference in viewing angles between infrared and RGB images caused by different imaging systems, ensuring spatial consistency and accuracy in subsequent modal feature extraction.

[0028] The present invention is further configured such that S2 includes: Based on the aligned image set, the mid-level modal feature maps of each modality are extracted using a feature encoder. Calculate the second-order differential feature differences between corresponding positions in each middle-layer semantic feature map, and construct a modal difference topological tensor to characterize the changing tension of local geometric structure; By combining the consistency of gradient directions between modes, the modal difference topological tensor is weighted and normalized to obtain the tension value between mode pairs; By aggregating and nonlinearly normalizing the intermodal tension values ​​in the spatial domain, the multimodal topological tension factor is finally obtained. Specifically, the fused mid-level modal feature map is obtained by inputting the three-modal images from the aligned image set into a unified shared feature encoder, such as ResNet, MobileNet, or Swin Transformer, which falls within the scope of existing technology and will not be elaborated here. The modal difference topological tensor is used to measure the local "cross curvature difference" of the modal pair at each pixel; the calculation logic of the modal difference topological tensor is as follows: ,in, For modal difference topology tensor, it represents the degree of structural inconsistency between the two modes in a mode pair at the same point. The larger the value, the stronger the geometric inconsistency of the mode pair at this point. For a modality pair, it contains images of three modalities from an aligned image set. Each pair of these forms a mode pair; The number of channels represents the feature dimension, typically ranging from 64 to 512. The first in the middle-level modal feature map The first one on the passage The value at the pixel position; This is the second-order mixed derivative, used to measure the local curvature of the texture at the current point. Modal topological tension factor calculation logic: ,in, This is the modal topological tension factor, used to characterize the fusion tension in the spatial structure of three modal images: RGB images, depth images, and registered infrared images. It is used when inconsistencies appear in the texture, edges, or local geometry of different modal images at the same location. The value will increase significantly, which is used to determine whether the modality fusion is natural and whether there are traces of forgery. The tension value between modal pairs represents the tension value between a modal pair used to highlight regions of structural inconsistency between two modal images. The gradient vector of the mode is obtained by... Calculated For the middle-level modal feature map The value at the pixel position; To ensure directional consistency between modal pairs, a larger value indicates a higher degree of convergence in the directions of the two modes, proving better structural alignment; conversely, a smaller value indicates a more severe difference in directions.

[0029] The present invention is further configured to generate a multi-scale feature map set by performing a multi-scale transformation operation on the mid-level modal feature map; Based on a multi-scale feature map set, key response regions identified by edge detection algorithms are extracted from the feature layer on a scale-by-scale basis. Calculate the local structural response characteristics within the critical response region and construct a structural differential anisotropy index that reflects local structural changes. Based on multi-scale feature maps, a Laplacian convolution kernel is introduced to calculate the high-frequency perturbation response intensity of key response regions; By combining the structural differential anisotropy index with the high-frequency perturbation response intensity, the local geometric expansion response characteristics at the current scale are obtained; Within the key response regions of all scale feature maps, local geometric expansion response features are accumulated and aggregated across scales and spatial domains to obtain the geometric expansion degree of the feature map used to characterize the amplitude of structural perturbation and the trend of edge expansion. Specifically, multi-scale feature maps are generated by performing pyramid-style multi-scale transformation on the mid-level modal feature maps, and all the obtained multi-scale feature maps are aggregated to construct a multi-scale feature map set. This invention does not limit the multi-scale transformation processing method; it can use spatial downsampling, perforated convolution transformation, pyramid convolution structure, or block self-attention mechanism based on Swin / ViT, etc., which are within the scope of existing technology and will not be elaborated here. In each scale feature map, an edge detection algorithm is used to extract key response regions for subsequent perturbation analysis. This invention does not limit the edge detection algorithm; it can use traditional algorithms such as Canny and LoG, or it can be replaced with a more efficient neural edge detection module according to the actual application scenario. The specific selection can be adaptively modified according to the actual deployment scenario. The boundary of the key region is determined by adjusting the boundary threshold. The boundary threshold is preset in advance, and the value range of the boundary threshold is as follows. However, a value greater than 8 may indicate structural synthesis, over-sharpening, or edge forgery regions. A value below 1 is recommended to ensure a natural structure without abnormal perturbations. Generating multi-scale feature maps and identifying key response regions are existing technologies and will not be elaborated upon here. Feature map geometric dilatation is an index that comprehensively models the amplitude and directional differences of edge perturbations in local structures and high-frequency texture features based on multi-scale feature maps. It is used to capture potential structural forgery traces or edge dilatation anomalies in images. The calculation logic for feature map geometric dilatation is as follows: ,in, The geometric expansion of the feature map; The hierarchy of multi-scale feature maps; As a critical response area, it is based on the first The multi-scale feature map of the layer is used to extract key edge regions through an edge detection algorithm; The pixel values ​​are the feature map pixel values, specifically the multi-scale values ​​output by the CNN encoder. Layer feature maps at pixels Pixel value at; For multi-scale feature maps in The second-order partial derivative in the direction, specifically expressed as the multi-scale feature map in The trend of edge change in direction; For multi-scale feature maps in The second-order partial derivative in the direction, specifically expressed as the multi-scale feature map in The trend of edge change in direction; This is the Laplacian operator, used to reflect the degree of grayscale transition between a pixel and its neighborhood; The differential anisotropy index is used to characterize the difference in the "expansion rate" of image edges in the horizontal and vertical directions. The larger the value, the more inconsistent the texture curvature of the point is in different directions, which often appears in the stretched areas of the synthetic or fake regions. The smaller the value or the closer it is to zero, the more similar the "change trend" of the point is in the horizontal and vertical directions. It represents the high-frequency perturbation response intensity, used to detect the perturbation intensity in high-frequency regions of an image, and to amplify high-frequency perturbations and synthesize misaligned responses in detail. Based on the local geometric expansion response characteristics at the current scale, by combining the structural difference anisotropy index and the high-frequency perturbation response intensity, the potential unnaturalness of the image can be detected from two dimensions: directional inconsistency and high-frequency abrupt change. The geometric dilatation of the feature map is obtained by summarizing the local geometric dilatation response features of all levels.

[0030] The present invention is further configured to integrate the multimodal topological tension factor and the feature map geometric expansion to form a modal coordination expansion coefficient that comprehensively evaluates structural complexity and modal alignment. The modal coordination dilation coefficient (MCD) is used to determine whether an image is a real face if it is lower than a preset threshold. Specifically, the MCD, as a fusion coefficient, is used to jointly evaluate the internal complexity of an image and the alignment quality between modalities. It is a static recognition result; a larger MCD indicates a natural structure and good fusion, making it more likely to be a real face, while a smaller MCD indicates an unreliable image or abnormal fusion, making it more likely to be a face image or a simulated face mask. The calculation logic for the MCD is as follows: ,in, The modal compatibility expansion coefficient; It is the multimodal topological tension factor; This refers to the geometric dilatation coefficient of the feature map. The judgment threshold can be obtained by statistically analyzing real faces and fake samples. The default setting is between 0.3 and 0.75. The higher the value, the higher the sensitivity of recognition. The default setting is 0.5. The specific value can be adaptively adjusted according to different deployment environments. When the modal dilatation coefficient is greater than the judgment threshold, the judgment result is not modified and the data is passed to the next stage. When the modal dilatation coefficient is less than or equal to the judgment threshold, the judgment result is set as a fake face, the judgment result is directly fed back, and a warning signal is generated.

[0031] The present invention is further configured such that S3 includes: Based on the depth image of the aligned image set, back projection calculation is performed using RGB image intrinsic parameters to generate a 3D point cloud of the face structure corresponding to the RGB camera coordinates; Based on the three-dimensional point cloud of the human face structure, the temporal information of the infrared image is fused to calculate multimodal evaluation indicators including structural restoration, dynamic topological coherence and infrared perturbation consistency. The structural restoration score is constructed by comparing the geometric deviation between the 3D point cloud of the facial structure and the point cloud model reconstructed based on a neural network. Specifically, the 3D point cloud of the facial structure is a set of 3D coordinates of the face surface obtained by backprojection using RGB images, depth images, and camera intrinsic parameters. It reflects the 3D structure of the real face in the current frame and is the basis for subsequent comparison, reconstruction, and perturbation analysis. The camera intrinsic parameters are calibrated at the time of camera manufacturing and are core parameters used to describe the relationship of the camera's imaging set. Calculating the 3D point cloud of the facial structure through backprojection is an existing technology and will not be elaborated further here. By inputting the depth image into the point cloud reconstruction network, the corresponding fitted point cloud is output. The same coordinate index is selected according to the spatial correspondence. This scheme does not limit the point cloud reconstruction network; methods including PIFuHD and 3D-CNN can be used to obtain the corresponding fitted point cloud. The structural restoration score is an indicator used to measure the error between the 3D facial structure reconstructed from the input image and the real structure. The smaller the value, the more accurate the reconstruction, and the more likely the image is to be a real face rather than a fake image. The structural restoration score calculation logic is as follows: ,in, For structural resilience; The total number of points in the 3D point cloud of the human face structure; The effective pixel set is the index of effective points within the facial region, obtained through face segmentation or contour detection algorithms, which is an existing technology and will not be elaborated further. The three-dimensional coordinates of pixels in the three-dimensional point cloud of the face structure are generated by back projection. To fit the 3D coordinates of pixels in the point cloud, the output of the point cloud reconstruction network is used. It is the three-dimensional Euclidean distance norm, used to calculate the Euclidean distance between pixels in the three-dimensional point cloud of each face structure and pixels in the fitted point cloud. Using the cube power is to amplify large errors and enhance the penalty for outliers.

[0032] The present invention is further configured to calculate density weights based on the density of the local neighborhood of the three-dimensional point cloud of the human face structure; Based on the 3D point cloud of the face structure in continuous frames, the iterative nearest point registration algorithm is used to analyze the temporal motion trajectory of the corresponding points, and the degree of structural deformation is calculated by combining the density weighting coefficient to construct dynamic topological coherence. Based on the local gradient changes in dynamically sensitive regions of infrared images, edge gradient operators are used to extract perturbation differences in key regions between adjacent frames, and infrared perturbation consistency is calculated. Specifically, density weights are used to emphasize the coherence of topologically stable regions; the density weight calculation logic is as follows: ,in, For pixels Density weights; For pixels The neighborhood point set is calculated by counting the points within the region whose distance is less than a fixed unit. This fixed unit is preset and defaults to 2. This is a smoothing parameter used to control the range of density calculations; it is typically set to 1~5. The three-dimensional coordinates of pixels in a three-dimensional point cloud of a human face structure; Pixels in a 3D point cloud of facial structure The three-dimensional coordinates of the neighborhood point. Dynamic topological coherence is used to determine whether the same point moves stably in three-dimensional space across multiple frames of images, or whether there are unnatural jitters, jumps, or other behaviors, thereby determining whether the input image sequence represents a real live human face; Dynamic topological coherence calculation logic: ,in, For dynamic topological coherence; Frame number represents the total number of consecutive image frames; The total number of points in the 3D point cloud of the human face structure; The L2 norm squared represents the square of the Euclidean distance, used to measure the spatial variation of the same point in adjacent frames. Infrared perturbation consistency measures whether the temperature gradient changes in dynamic facial sensitive areas, such as the eyes, nose tip, and lips, are continuous, natural, and consistent between adjacent frames. Abnormal or inconsistent changes in these areas often indicate video forgery, composite masking, image mapping, infrared masks, and other fraudulent activities. The infrared perturbation consistency calculation logic is as follows: ,in, For infrared perturbation consistency; For dynamic sensitive regions, the regions most easily changed with facial expressions and movements in the facial infrared image are generally selected, such as the area around the eyes, the bridge / tip of the nose, and the lip area. These regions can be obtained using facial landmark detection algorithms, such as dlib, MediaPipe, and DeepFace. To register infrared images medium pixel The gradient is used to reflect the difference in thermal structure between each pixel and its neighborhood in the current frame. It can be obtained by the Sobel operator, the Prewitt operator, or the differential convolution kernel, etc. The acquisition method is the existing technology, which will not be elaborated on in detail. The time difference is the interval between adjacent frames. It is a fixed time interval preset during system deployment, with a default value of 1.

[0033] The present invention is further configured such that S4 includes: Based on the 3D point cloud of the human face structure, a skin texture undulation index is constructed by calculating the degree of difference in local normal vectors in the neighborhood of each pixel. The facial thermal intensity distribution is extracted based on the registered infrared images in the aligned image set. The temperature fluctuation amplitude within each preset facial region is calculated, and the average amplitude of temperature fluctuation in each region is calculated to construct the skin thermal distribution uniformity index. A skin realism index is constructed by integrating two types of indicators: skin texture undulation index and skin thermal distribution uniformity index. Specifically, the skin texture undulation index is a three-dimensional geometric indicator that measures whether the skin's microstructure is continuous, natural, and realistic. Starting from the normal vector variability, it calculates the local perturbation intensity of the normal vector on the surface of the three-dimensional point cloud to assess the skin roughness and pore texture clarity in the facial region, completing the detection of fake faces caused by the lack of three-dimensional microstructure in cases such as fake faces, printed photos, screen replays, and mask synthesis. The calculation logic of the skin texture undulation index is as follows: ,in, This refers to the skin texture undulation index; The total number of points in the 3D point cloud of the human face structure; For pixels The set of neighborhood points; The normal vector is the pixel. The corresponding normal vector is mainly used to describe the local surface direction of a 3D point. The unit normal vector can be obtained by performing covariance analysis or principal component analysis (PCA) on the neighborhood set of each pixel point, and is used to characterize the normal orientation of the local structure. This process belongs to the prior art and will not be described in detail in this invention. For pixels The corresponding unit normal vector. The skin thermal distribution uniformity index is a statistical measure of the temperature fluctuation range of healthy facial functional areas based on registered infrared image analysis, quantifying the uniformity or disorder of skin thermal distribution; the calculation logic of the skin thermal distribution uniformity index is as follows: ,in, The skin heat distribution uniformity index; This represents the preset total number of facial regions; This is a thermal intensity map, representing the facial thermal intensity distribution of the registered infrared image. Because the grayscale value in the image output by the infrared camera has a monotonic relationship with the temperature, the grayscale value of the registered infrared image can be directly used as the thermal intensity map. To pre-define the facial region mask, the pre-define facial regions include: the forehead, glabella and bridge of the nose in the T-zone, the jaw and left and right cheeks in the U-zone. The region mask of the pre-define facial region is obtained by using facial key point localization technology, including Dlib and MediaPipe, which are existing technologies and will not be described in detail here. The Skin Authenticity Index is used to calculate the local temperature difference amplitude. It obtains the temperature fluctuation amplitude within a region by calculating the difference between the highest and lowest temperatures, thus characterizing the presence of local anomalies. The average of the local temperature difference amplitudes across all regions yields the overall facial skin thermal distribution uniformity index. The Skin Authenticity Index is a unified scoring model for realistic face detection constructed through the weighted fusion of two sub-indicators. The calculation logic for the Skin Authenticity Index is as follows: ,in, The skin authenticity index indicates that the higher the value, the more unnatural the skin appears or that there are signs of forgery. This is the texture weight, used to balance the contribution of the skin texture undulation index to the skin health score. The default value is 1. If you value forgery detection, it is recommended to use a value of 1.5. The sum of the texture weight and the heat distribution weight must be 2. This is the heat distribution weight, used to balance the contribution of the skin heat distribution equilibrium index to the skin health score. The default value is 1, and the sum of the texture weight and the heat distribution weight must be 2.

[0034] The present invention is further configured to construct a confidence score function for authenticity by substituting a multimodal evaluation index, a modal coordination expansion coefficient, and a skin authenticity index into a preset multi-parameter scoring formula. The authenticity of a face is determined based on the output of the authenticity confidence scoring function. A face is considered real if the authenticity confidence score is greater than a threshold; otherwise, it is considered fake. Specifically, the authenticity confidence scoring function is a non-linear fusion function used for facial authenticity confidence scoring. It integrates multiple key parameters from multiple modalities, scales, and spatial structure levels, ultimately outputting a normalized authenticity confidence score, which is compared with the threshold to determine whether a face is real or fake. The calculation logic of the authenticity confidence scoring function is as follows: ,in, The output of the confidence score function is the confidence score. The modal compatibility expansion coefficient; For infrared perturbation consistency, The 12th power representing the uniformity of infrared perturbations is used to indicate a strong amplification of authenticity. For structural resilience; For dynamic topological coherence; As a skin authenticity index; As an incentive factor, because the infrared perturbation of real faces is small, it can exponentially improve the authenticity confidence score; As a penalty term, the larger the product of the three terms, the less truthful it is. This is used to balance the various penalties and prevent any one from becoming too dominant. This is used to compress the exponential space, suppress extreme values, and improve the numerical stability of the model. The decision threshold is used to ultimately determine whether a face image captured by the dual cameras is a real face. When the authenticity confidence score is greater than the decision threshold, the face image captured by the dual cameras is determined to be a real face; otherwise, it is determined to be a non-real face. The default decision threshold is set to 0.6. To meet the requirements of some deployment environments, a dual-threshold decision mode can be set. When the authenticity confidence score is greater than decision threshold one, the face image captured by the dual cameras is determined to be a real face. When the authenticity confidence score is less than or equal to decision threshold one but greater than decision threshold two, the face image captured by the dual cameras is determined to be suspicious and requires manual verification. When the authenticity confidence score is less than or equal to decision threshold two, the face image captured by the dual cameras is determined to be a non-real face. The default decision threshold is set to 0.75, and the default decision threshold is set to 0.4. The higher the decision threshold value, the higher the decision accuracy. The specific values ​​should be adaptively modified according to the actual deployment needs.

[0035] Example 2: Please see Figure 2 An exemplary multimodal dual-camera face depth detection system includes: Image acquisition and registration module: Acquires original image sets through multimodal dual cameras, derives epipolar constraint relationship based on camera geometric model and original image set and calculates disparity offset value, outputs aligned image set after interpolation transformation of original image set; Joint feature extraction module: Based on the aligned image set, it performs cross-modal joint feature extraction. By calculating the multimodal topological tension factor and the geometric dilatation of the feature map, it fuses and generates the modal coordinated dilatation coefficient to complete the primary anomaly screening. 3D face image modeling module: Generates 3D point cloud of face structure based on aligned image set, and calculates multimodal evaluation index by combining temporal analysis; Liveness detection and authenticity assessment module: Based on the aligned image set, the skin authenticity index is derived, and the modal compatibility dilation coefficient, multimodal evaluation index and skin authenticity index are integrated. The authenticity confidence score is calculated through a multi-parameter weighted decision function to determine authenticity.

[0036] It should be noted that the multimodal dual-camera face depth detection system and the multimodal dual-camera face depth detection method provided in the above embodiments belong to the same concept. The specific operation methods of each module and unit have been described in detail in the method embodiments and will not be repeated here. In practical applications, the multimodal dual-camera face depth detection system provided in the above embodiments can be assigned to different functional modules as needed, that is, the internal structure of the system can be divided into different functional modules to complete all or part of the functions described above. This is not a limitation here.

[0037] The above embodiments can be implemented, in whole or in part, by software, hardware, firmware, or any other combination thereof. When implemented using software, the above embodiments can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more sets of available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium. A semiconductor medium can be a solid-state drive.

[0038] It should be understood that the term "and / or" in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. A and B can be singular or plural. Additionally, the character " / " in this article generally indicates an "or" relationship between the preceding and following related objects, but it can also represent an "and / or" relationship. Please refer to the context for a more accurate understanding.

[0039] In this application, "at least one" means one or more, and "more than one" means two or more. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single or multiple items. For example, at least one of a, b, or c can mean: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or multiple.

[0040] It should be understood that in the various embodiments of this application, the order of the above-mentioned processes does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0041] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0042] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0043] In the several embodiments provided in this application, it should be understood that the disclosed system can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.

[0044] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0045] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0046] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0047] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A face depth detection method based on multimodal dual-camera setup, characterized in that, include: S1: Acquire the original image set through multimodal dual cameras, derive the epipolar constraint relationship based on the camera geometric model and the original image set, calculate the disparity offset value, and output the aligned image set after interpolation transformation of the original image set; S2: Cross-modal joint feature extraction is performed based on aligned image sets. By calculating the multimodal topological tension factor and the geometric dilatation of the feature maps, a modal coordination expansion coefficient is generated to complete the initial anomaly screening. The multimodal topological tension factor characterizes the fusion tension in the spatial structure of RGB images, depth images, and registered infrared images. When inconsistencies appear in the texture, edges, or local geometry of different modal images at the same location, the multimodal topological tension factor value increases significantly, used to determine whether the modal fusion is natural and whether there are forgery traces. The feature map geometric dilatation calculation logic is as follows: ,in, The geometric expansion of the feature map; The hierarchy of multi-scale feature maps; As a critical response area, it is based on the first The multi-scale feature map of the layer is used to extract key edge regions through an edge detection algorithm; The pixel values ​​are the feature map pixel values, specifically the multi-scale values ​​output by the CNN encoder. Layer feature maps at pixels Pixel value at; For multi-scale feature maps in The second-order partial derivative in the direction, specifically expressed as the multi-scale feature map in The trend of edge change in direction; For multi-scale feature maps in The second-order partial derivative in the direction, specifically expressed as the multi-scale feature map in The trend of edge change in direction; This is the Laplacian operator, used to reflect the degree of grayscale transition between a pixel and its neighborhood; The differential anisotropy index is used to characterize the difference in the "expansion rate" of image edges in the horizontal and vertical directions. The larger the value, the more inconsistent the texture curvature of the pixels in different directions, which often appears in the stretched areas of the synthetic or fake regions. The smaller the value or the closer it is to zero, the more similar the "change trend" of the pixels in the horizontal and vertical directions. It represents the high-frequency perturbation response intensity, used to detect the perturbation intensity in high-frequency regions of an image, and to amplify high-frequency perturbations and synthesize misaligned responses in detail. Based on the local geometric expansion response characteristics at the current scale, by combining the structural difference anisotropy index and the high-frequency perturbation response intensity, the potential unnaturalness of the image can be detected from two dimensions: directional inconsistency and high-frequency abrupt change. The geometric dilatation of the feature map is obtained by summarizing the local geometric dilatation response features of all levels; the modal coordination dilatation coefficient is used as a fusion coefficient to jointly evaluate the internal complexity of the image and the alignment quality between modalities. It is a static recognition result. The larger the value, the more natural the structure and the better the fusion, and the more likely it is to be a real face. The smaller the value, the more likely the image is unreliable or the fusion is abnormal, and the more likely it is to be a face image or a simulated face mask. S3: Generate a 3D point cloud of facial structure based on the aligned image set, and calculate multimodal evaluation metrics using temporal analysis. This includes back-projection calculations using RGB image intrinsic parameters based on the depth image from the aligned image set to generate a 3D point cloud of facial structure corresponding to RGB camera coordinates. Then, fuse temporal information from infrared images into the 3D point cloud of facial structure to calculate multimodal evaluation metrics including structural restoration, dynamic topological coherence, and infrared perturbation consistency. Specifically, structural restoration is constructed by comparing the geometric deviation between the 3D point cloud of facial structure and the point cloud model reconstructed based on a neural network. Structural restoration measures the difference between the 3D facial structure reconstructed from the input image and the true structure. One indicator is the accuracy of the reconstruction; the smaller the value, the more accurate the reconstruction and the more likely the image is to be a real human face rather than a fake image. Dynamic topological coherence is used to determine whether the same point moves stably in three-dimensional space under multi-frame 3D point cloud images of facial structure, or whether there are unnatural jitters, jumps, or other behaviors, thereby determining whether the input image sequence is a real live human face. Infrared perturbation consistency is used to measure whether the temperature gradient changes of dynamic sensitive areas of the face in the infrared image, such as the eyes, nose tip, and lips, are continuous, natural, and consistent between adjacent registered infrared image frames. If these areas change abnormally or inconsistently, it is often a manifestation of forgery behaviors such as video forgery, synthetic masking, image mapping, and infrared masks. S4: Based on the aligned image set, a skin authenticity index is derived. This index integrates the modal coordination dilation coefficient, multimodal evaluation metrics, and the skin authenticity index. A multi-parameter weighted decision function is used to calculate the authenticity confidence score for authenticity determination. This includes: constructing a skin texture undulation index based on the 3D point cloud of the face structure by calculating the difference in local normal vectors within the neighborhood of each pixel; extracting facial thermal intensity distribution from the registered infrared images in the aligned image set, calculating the temperature fluctuation amplitude within each preset facial region, and calculating the average amplitude of temperature fluctuation in each region to construct a skin thermal distribution equilibrium index; and constructing a skin authenticity index by integrating the skin texture undulation index and the skin thermal distribution equilibrium index. The skin authenticity index is a unified scoring model for real face detection constructed through the weighted fusion of the two sub-indicators. A larger value indicates more unnatural skin or the presence of forgery traces. The authenticity confidence scoring function is a nonlinear fusion function for facial authenticity confidence scoring, integrating multiple key parameters from multimodal, multi-scale, and spatial structural levels. The final output is a normalized real confidence score, which is compared with a judgment threshold to determine whether the face is real or fake.

2. The face depth detection method based on multimodal dual-camera according to claim 1, characterized in that, S1 includes: The original image set includes: RGB images, depth images, and infrared images of the face region, which are simultaneously acquired by an RGB-D camera and a separate IR camera; Based on the geometric information of the depth image and the camera calibration parameters, the pixel-level transformation relationship from the infrared image to the RGB image coordinate system is calculated and set as the disparity offset value; The infrared image is transformed based on the parallax offset value to obtain a corrected infrared image that is aligned with the RGB image space; The corrected infrared image is mapped to the coordinate system of the RGB image using an interpolation algorithm, generating a registered infrared image that is aligned with the pixels of the RGB image and the depth image. An aligned image set is constructed by combining RGB images, depth images, and registered infrared images.

3. The face depth detection method based on multimodal dual-camera according to claim 1, characterized in that, S2 includes: Based on the aligned image set, the mid-level modal feature maps of each modality are extracted using a feature encoder. Calculate the second-order differential feature differences between corresponding positions in each middle-layer semantic feature map, and construct a modal difference topological tensor to characterize the changing tension of local geometric structure; By combining the consistency of gradient directions between modes, the modal difference topological tensor is weighted and normalized to obtain the tension value between mode pairs; By aggregating and nonlinearly normalizing the tension values ​​between all modal pairs in the spatial domain, the multimodal topological tension factor is finally obtained.

4. The face depth detection method based on multimodal dual-camera according to claim 3, characterized in that, A multi-scale feature map set is generated by performing multi-scale transformation operations on the mid-level modal feature map; Based on a multi-scale feature map set, key response regions identified by edge detection algorithms are extracted from the feature layer on a scale-by-scale basis. Calculate the local structural response characteristics within the critical response region and construct a structural differential anisotropy index that reflects local structural changes. Based on multi-scale feature maps, a Laplacian convolution kernel is introduced to calculate the high-frequency perturbation response intensity of key response regions; By combining the structural differential anisotropy index with the high-frequency perturbation response intensity, the local geometric expansion response characteristics at the current scale are obtained; Within the key response regions of feature maps at all scales, the local geometric expansion response features are accumulated and aggregated across scales and spatial domains to obtain the feature map geometric expansion degree used to characterize the amplitude of structural perturbation and the trend of edge expansion.

5. The face depth detection method based on multimodal dual-camera according to claim 4, characterized in that, By integrating the multimodal topological tension factor and the geometric expansion degree of the feature map, a modal coordination expansion coefficient is formed that comprehensively evaluates structural complexity and modal alignment. The modal coordination dilatation coefficient is judged based on a preset preliminary judgment threshold. When the modal coordination dilatation coefficient is lower than the preliminary judgment threshold, it is judged as a non-real face.

6. The face depth detection method based on multimodal dual-camera according to claim 1, characterized in that, Density weights are calculated based on the density of the local neighborhood of a 3D point cloud of a human face structure. Based on the 3D point cloud of the face structure in continuous frames, the iterative nearest point registration algorithm is used to analyze the temporal motion trajectory of the corresponding points, and the degree of structural deformation is calculated by combining the density weighting coefficient to construct dynamic topological coherence. Based on the local gradient changes of dynamic sensitive regions in infrared images, the edge gradient operator is used to extract the perturbation differences of key regions between adjacent frames and calculate the infrared perturbation consistency.

7. The face depth detection method based on multimodal dual-camera according to claim 1, characterized in that, Based on multimodal evaluation indicators, modal coordination expansion coefficient, and skin authenticity index, a confidence score function for authenticity is jointly constructed by substituting them into a preset multi-parameter scoring formula. The authenticity of a face is determined based on the output of the authenticity confidence score function. If the authenticity confidence score is greater than the judgment threshold, the face is judged as a real face; otherwise, it is judged as a non-real face.

8. A multimodal dual-camera face depth detection system, used to implement the multimodal dual-camera face depth detection method according to any one of claims 1-7, characterized in that, include: Image acquisition and registration module: Acquires original image sets through multimodal dual cameras, derives epipolar constraint relationship based on camera geometric model and original image set and calculates disparity offset value, outputs aligned image set after interpolation transformation of original image set; Joint feature extraction module: Based on the aligned image set, it performs cross-modal joint feature extraction. By calculating the multimodal topological tension factor and the geometric dilatation of the feature map, it fuses and generates the modal coordinated dilatation coefficient to complete the primary anomaly screening. 3D face image modeling module: Generates 3D point cloud of face structure based on aligned image set, and calculates multimodal evaluation index by combining temporal analysis; Liveness detection and authenticity assessment module: Based on the aligned image set, the skin authenticity index is derived, and the modal compatibility dilation coefficient, multimodal evaluation index and skin authenticity index are integrated. The authenticity confidence score is calculated through a multi-parameter weighted decision function to determine authenticity.

Citation Information

Patent Citations

  • Multi-modal face anti-counterfeiting detection method and device based on cross-modal fusion, equipment and medium

    CN117437677A

  • One-way door control method, system and equipment based on image recognition

    CN120375503A