Multimodal face Anti-counterfeiting model training method, and face Anti-counterfeiting recognition method
By using IR images and RGB images collected by binocular cameras to calculate depth images, and combining three modal data to train multimodal anti-counterfeiting models, the problem of low reliability of traditional face anti-counterfeiting recognition technology is solved, and higher recognition accuracy and security are achieved.
Patent Information
- Application Number
- PCT/CN2023/137791
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-10-30
- Filing Date
- 2023-12-11
- Publication Date
- 2025-05-08
AI Technical Summary
Traditional facial anti-counterfeiting recognition technology relies on single modal data (such as IR images, RGB images or Depth images) for feature extraction, which poses security risks, such as photo attacks, video attacks and 3D model attacks, resulting in low reliability.
IR images and RGB images were collected by a binocular camera, depth images were calculated, and multimodal anti-counterfeiting model was trained in combination with three modal data (IR images, RGB images and depth images) to perform face alignment and feature extraction.
Improve the accuracy and security of facial anti-counterfeiting recognition, and enhance the resistance to photo attacks, video attacks and 3D model attacks.
Smart Images

Figure CN2023137791_08052025_PF_FP_ABST
Abstract
Description
Multimodal face anti-counterfeiting model training and face anti-counterfeiting recognition method Technical Field
[0001] The present invention relates to the field of computer vision technology, and in particular to a multimodal face anti-counterfeiting model training and face anti-counterfeiting recognition method. Background Art
[0002] Currently, facial recognition technology is widely used in various fields, such as mobile phone unlocking, access control systems, and payment systems. However, traditional facial recognition technology relies on facial modality data from IR images, RGB images, or depth images to extract features and classify whether a face is genuine. This poses security risks, such as photo attacks, video attacks, and 3D model attacks, and traditional facial recognition technology is not very reliable.
[0003] Summary of the Invention
[0004] The present invention provides a multimodal face anti-counterfeiting model training and face anti-counterfeiting recognition method to solve the technical problem of low reliability of traditional face anti-counterfeiting recognition technology.
[0005] In order to solve the above technical problems, an embodiment of the present invention provides a multimodal face anti-counterfeiting model training method, comprising:
[0006] A plurality of facial images are captured using a binocular camera as a facial image set; wherein the dual-mode module includes an IR lens and an RGB lens aligned left and right; the facial image set includes the plurality of IR images and an RGB image corresponding to each IR image captured at the same time and scene;
[0007] Calculating a corresponding depth image based on each set of IR images and corresponding RGB images in the face image set, and adding the obtained depth image to the face image set to obtain an updated face image set;
[0008] Performing face alignment on all images in the updated face image set to obtain an aligned face dataset;
[0009] A multimodal anti-counterfeiting model is trained based on the aligned face dataset to obtain a trained multimodal face anti-counterfeiting model.
[0010] As a preferred solution, the step of calculating the corresponding depth image based on each set of IR images and the corresponding RGB images in the face image set includes:
[0011] Preprocessing the IR image and the RGB image to obtain a processed IR image and an RGB image;
[0012] performing image stereo correction on the processed IR image and the RGB image to obtain a corrected IR image and RGB image;
[0013] Obtain the feature points of each set of corrected IR images and the corresponding RGB images through the facial feature point detection algorithm;
[0014] Calculating a matching matrix between the IR image and the corresponding RGB image based on the feature points of the IR image and the corresponding RGB image, and obtaining a pixel correspondence relationship between the IR image and the corresponding RGB image based on the matching matrix;
[0015] Calculating a disparity map of the IR image and the corresponding RGB image based on a pixel correspondence between the IR image and the corresponding RGB image;
[0016] A corresponding depth map is calculated according to the IR image and the disparity map of the corresponding RGB image.
[0017] As a preferred solution, the step of calculating the corresponding depth map based on the IR image and the disparity map of the corresponding RGB image includes:
[0018] Determine disparity values of corresponding pixels in the IR image and the corresponding RGB image according to the disparity map of the IR image and the corresponding RGB image;
[0019] Obtain the focal length of the binocular camera, the baseline distance between the left and right cameras, and the horizontal coordinates of the pixel points in the IR image and the corresponding RGB image;
[0020] The corresponding depth map is calculated according to the disparity value of the corresponding pixel point, the focal length of the binocular camera, the baseline distance between the left and right cameras, and the horizontal coordinate of the pixel point.
[0021] As a preferred solution, performing face alignment on all images in the updated face image set to obtain an aligned face dataset includes:
[0022] Performing facial feature point detection on all images in the updated facial image set to obtain feature point coordinates corresponding to each facial image and a standard facial image; wherein the standard facial image is a preset facial image with specified feature points;
[0023] Calculating a corresponding perspective transformation matrix based on the coordinates of the feature points corresponding to each facial image and the standard face and the coordinates of the feature points of the standard face; and calculating a facial image after alignment of each facial image based on the perspective transformation matrix;
[0024] The aligned face images are cropped to the same size as the standard face images to obtain an aligned face dataset.
[0025] As a preferred solution, the multimodal anti-counterfeiting model includes: an IR unimodal anti-counterfeiting sub-model, an RGB unimodal anti-counterfeiting sub-model, a deep unimodal anti-counterfeiting sub-model, and a multimodal anti-counterfeiting sub-model; wherein the IR unimodal anti-counterfeiting sub-model, the RGB unimodal anti-counterfeiting sub-model, and the deep unimodal anti-counterfeiting sub-model have the same structure;
[0026] The step of training a multimodal anti-counterfeiting model based on the aligned face dataset to obtain a trained multimodal face anti-counterfeiting model comprises:
[0027] Training the IR unimodal anti-counterfeiting sub-model according to the IR images in the aligned face dataset to obtain a trained IR unimodal anti-counterfeiting sub-model;
[0028] Training the RGB single-modal anti-counterfeiting sub-model according to the RGB images in the aligned face dataset to obtain a trained RGB single-modal anti-counterfeiting sub-model;
[0029] Training the deep unimodal anti-counterfeiting sub-model according to the depth images in the aligned face dataset to obtain a trained deep unimodal anti-counterfeiting sub-model;
[0030] Inputting the IR image of the comparison face dataset into the trained IR unimodal anti-counterfeiting sub-model to obtain the output of the trained IR unimodal anti-counterfeiting sub-model;
[0031] Inputting the RGB image of the comparison face dataset into the trained RGB single-modal anti-counterfeiting sub-model to obtain the output of the trained RGB single-modal anti-counterfeiting sub-model;
[0032] Inputting the depth image of the comparison face dataset into the trained deep unimodal anti-counterfeiting sub-model to obtain the output of the trained deep unimodal anti-counterfeiting sub-model;
[0033] The multimodal anti-counterfeiting sub-model is trained based on the output of the trained IR unimodal anti-counterfeiting sub-model, the output of the trained RGB unimodal anti-counterfeiting sub-model, and the output of the trained depth unimodal anti-counterfeiting sub-model as inputs to obtain a trained multimodal anti-counterfeiting sub-model.
[0034] A trained multimodal face anti-counterfeiting model is generated based on the trained IR unimodal anti-counterfeiting sub-model, the trained RGB unimodal anti-counterfeiting sub-model, the trained depth unimodal anti-counterfeiting sub-model and the trained multimodal anti-counterfeiting sub-model.
[0035] As a preferred solution, the loss functions of the IR unimodal anti-counterfeiting sub-model, the RGB unimodal anti-counterfeiting sub-model, the depth unimodal anti-counterfeiting sub-model and the multimodal anti-counterfeiting sub-model are all binary classification loss functions.
[0036] As a preferred solution, the multimodal anti-counterfeiting model includes: an IR unimodal anti-counterfeiting sub-model, an RGB unimodal anti-counterfeiting sub-model, a deep unimodal anti-counterfeiting sub-model, and a multimodal anti-counterfeiting sub-model; wherein the IR unimodal anti-counterfeiting sub-model, the RGB unimodal anti-counterfeiting sub-model, and the deep unimodal anti-counterfeiting sub-model have the same structure; the input of the multimodal anti-counterfeiting sub-model is the output of the IR unimodal anti-counterfeiting sub-model, the RGB unimodal anti-counterfeiting sub-model, and the deep unimodal anti-counterfeiting sub-model;
[0037] The step of training a multimodal anti-counterfeiting model based on the aligned face dataset to obtain a trained multimodal face anti-counterfeiting model comprises:
[0038] The IR image, RGB image and depth image in the aligned face dataset are respectively input into the IR unimodal anti-counterfeiting sub-model, RGB unimodal anti-counterfeiting sub-model and depth unimodal anti-counterfeiting sub-model, and the outputs of the obtained IR unimodal anti-counterfeiting sub-model, RGB unimodal anti-counterfeiting sub-model and depth unimodal anti-counterfeiting sub-model are input into the multimodal anti-counterfeiting sub-model. The IR unimodal anti-counterfeiting sub-model, RGB unimodal anti-counterfeiting sub-model, depth unimodal anti-counterfeiting sub-model and multimodal anti-counterfeiting sub-model are trained at the same time to obtain a trained multimodal face anti-counterfeiting model.
[0039] As a preferred solution, the loss function of the multimodal face anti-counterfeiting model is: loss = p i ×LOSSI+p r ×LOSSR+p d ×LOSSD+p m ×LOSSM;
[0040] Among them, LOSSI is the loss value of the IR unimodal anti-counterfeiting sub-model, LOSSR is the loss value of the RGB unimodal anti-counterfeiting sub-model, LOSSD is the loss value of the depth unimodal anti-counterfeiting sub-model, LOSSM is the loss value of the multimodal anti-counterfeiting sub-model, and p i is the weighted coefficient of the loss value of the preset IR single-modal anti-counterfeiting sub-model, p r is the weighted coefficient of the loss value of the preset RGB single-modal anti-counterfeiting sub-model, p d is the weighted coefficient of the loss value of the preset deep single-modal anti-counterfeiting sub-model, p m is the weighting coefficient of the preset multimodal sub-model loss value.
[0041] Based on the above embodiment, another embodiment of the present invention provides a face anti-counterfeiting recognition method, comprising:
[0042] Acquire facial images captured by a binocular camera; wherein the dual-mode module includes an IR lens and an RGB lens; the facial image set includes a plurality of IR images and an RGB image corresponding to each IR image captured at the same time and scene;
[0043] The facial image is input into a multimodal facial anti-counterfeiting model so that the multimodal facial anti-counterfeiting model performs anti-counterfeiting identification on the facial image to obtain an anti-counterfeiting identification result; wherein, the multimodal facial anti-counterfeiting model is trained according to the above-mentioned multimodal facial anti-counterfeiting model training method.
[0044] Compared with the prior art, the embodiments of the present invention have the following beneficial effects:
[0045] The present invention uses a binocular camera to collect a number of facial images as a facial image set; wherein the dual-mode module includes an IR lens and an RGB lens; the facial image set includes a number of IR images and an RGB image corresponding to each IR image, collected at the same time and scene; based on the internal and external parameter matrices of the IR lens and the RGB lens, a depth image of each group of IR images and the corresponding RGB image in the facial image set is calculated, and the obtained depth image is added to the facial image set to obtain an updated facial image set; facial alignment is performed on all images in the updated facial image set to obtain a multimodal aligned facial dataset; a multimodal anti-counterfeiting model is trained based on the multimodal aligned facial dataset to obtain a trained multimodal facial anti-counterfeiting model. The present invention combines two facial modal data, IR images and RGB images, to generate a depth image, and trains a multimodal anti-counterfeiting model based on three facial modal data, IR images, RGB images, and depth images. Using the multimodal anti-counterfeiting model trained in this way for facial anti-counterfeiting recognition can improve the accuracy and security of facial anti-counterfeiting recognition. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] FIG1 is a flow chart of a multimodal face anti-counterfeiting model training method provided by one embodiment of the present invention;
[0047] FIG2 is a schematic diagram of the structure of a multimodal face anti-counterfeiting model provided by one embodiment of the present invention;
[0048] FIG3 is a schematic diagram of the structure of another multimodal face anti-counterfeiting model provided by an embodiment of the present invention;
[0049] FIG4 is a flow chart of a face anti-counterfeiting recognition method provided in accordance with an embodiment of the present invention. DETAILED DESCRIPTION
[0050] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0051] Example 1
[0052] Please refer to FIG1 , which is a flowchart of a multimodal face anti-counterfeiting model training method provided by one embodiment of the present invention, including:
[0053] S1. Capture a plurality of facial images as a facial image set using a binocular camera; wherein the dual-mode module includes an IR lens and an RGB lens aligned left and right; the facial image set includes the plurality of IR images and an RGB image corresponding to each IR image, captured at the same moment and in the same scene.
[0054] In this embodiment, the binocular camera can simultaneously capture the left and right perspectives of the same scene, obtaining images captured by the left and right cameras, one of which is an IR image and the other is an RGB image. The IR lens of the binocular camera collects n IR images, and the RGB lens of the binocular camera collects n RGB images to form a face image set, which is recorded as: {image_date_A|ir i ,rgb i ,i=1~n};
[0055] Among them, ir i Represents the i-th IR image, rgb i Represents the i-th RGB image.
[0056] It should be noted that the i-th IR image and the i-th RGB image are two images taken by the binocular camera at the same moment and in the same scene.
[0057] S2. Calculate a corresponding depth image based on each set of IR images and corresponding RGB images in the face image set, and add the obtained depth image to the face image set to obtain an updated face image set.
[0058] It should be noted that the set of IR images and corresponding RGB images refers to the IR image and RGB image of the same scene captured at the same time. In this embodiment, the corresponding depth image is calculated from the IR image and the corresponding RGB image, which is more cost-effective than directly using a depth camera to obtain the depth image.
[0059] In a preferred embodiment, the step of calculating the corresponding depth image based on each set of IR images and the corresponding RGB images in the face image set includes:
[0060] Preprocessing the IR image and the RGB image to obtain a processed IR image and an RGB image;
[0061] performing image stereo correction on the processed IR image and the RGB image to obtain a corrected IR image and RGB image;
[0062] Obtain the feature points of each set of corrected IR images and the corresponding RGB images through the facial feature point detection algorithm;
[0063] Calculating a matching matrix between the IR image and the corresponding RGB image based on the feature points of the IR image and the corresponding RGB image, and obtaining a pixel correspondence relationship between the IR image and the corresponding RGB image based on the matching matrix;
[0064] Calculating a disparity map of the IR image and the corresponding RGB image based on a pixel correspondence between the IR image and the corresponding RGB image;
[0065] A corresponding depth map is calculated according to the IR image and the disparity map of the corresponding RGB image.
[0066] In this embodiment, the IR image and the RGB image are preprocessed. The preprocessing steps include operations such as descaling and correction to eliminate camera lens distortion and align the left and right camera images. Stereo correction is performed on the images to align the left and right camera images horizontally for subsequent disparity calculation. A face detection algorithm is used to obtain feature points for each corrected IR image and corresponding RGB image. Each feature point in the IR image corresponds to a feature point in the corresponding RGB image, forming a feature point group. Each set of IR images and corresponding RGB images contains at least eight feature point groups. Based on the feature points of the IR image and the corresponding RGB image, a matching matrix is calculated between the IR image and the corresponding RGB image. Based on the matching matrix, the correspondence between the pixels of the IR image and the corresponding RGB image is obtained. Based on the correspondence between the pixels of the IR image and the corresponding RGB image, a disparity map is calculated for the IR image and the RGB image. A corresponding depth map is calculated based on the disparity map of the IR image and the RGB image.
[0067] The updated face dataset is recorded as: {image_data_B|ir i ,rgb i ,depth i ,i=1~n};
[0068] Among them, ir i Represents the i-th IR image, rgb i Represents the i-th RGB image, depth i Represents the i-th depth image.
[0069] It should be noted that depth i It is based on ir i and RGB i Calculated, that is, depth i It's ir i and RGB i The corresponding depth image.
[0070] In a preferred embodiment, the calculating the corresponding depth map according to the disparity map of the IR image and the corresponding RGB image includes:
[0071] Determine disparity values of corresponding pixels in the IR image and the corresponding RGB image according to the disparity map of the IR image and the corresponding RGB image;
[0072] Obtain the focal length of the binocular camera, the baseline distance between the left and right cameras, and the horizontal coordinates of the pixel points in the IR image and the corresponding RGB image;
[0073] The corresponding depth map is calculated according to the disparity value of the corresponding pixel point, the focal length of the binocular camera, the baseline distance between the left and right cameras, and the horizontal coordinate of the pixel point.
[0074] In this embodiment, the corresponding depth map is calculated based on the triangulation principle according to the disparity value of the corresponding pixel point, the focal length of the binocular camera, the baseline distance between the left and right cameras, and the horizontal coordinate of the pixel point. The expression of the depth map is:
[0075] Among them, DepthPixel represents the depth image, DepthPixel i,j Indicates the depth value of the i,j pixel position, FocalLength indicates the focal length of the camera, Baseline indicates the baseline distance between the left and right cameras, Disparity i,j Indicates the disparity value of the i, j pixel position, PointIR i,j Indicates the horizontal position of the IR camera at pixel position i,j, PointRGB i,j Indicates the horizontal position of the RGB camera at pixel position i,j.
[0076] S3. Perform face alignment on all images in the updated face image set to obtain an aligned face dataset.
[0077] In a preferred embodiment, performing face alignment on all images in the updated face image set to obtain an aligned face dataset includes:
[0078] Performing facial feature point detection on all images in the updated facial image set to obtain feature point coordinates corresponding to each facial image and a standard facial image; wherein the standard facial image is a preset facial image with specified feature points;
[0079] Calculating a corresponding perspective transformation matrix based on the coordinates of the feature points corresponding to each facial image and the standard face and the coordinates of the feature points of the standard face; and calculating a facial image after alignment of each facial image based on the perspective transformation matrix;
[0080] The aligned face images are cropped to the same size as the standard face images to obtain an aligned face dataset.
[0081] In this embodiment, facial feature point detection is performed on all images in the updated facial image set to obtain the feature point coordinates corresponding to each facial image and the standard facial image; the standard facial image size is 112*112 pixels, and 5 feature point coordinates are preset, and these 5 feature point coordinates are marked as Slandmarks. i , i = 1 to 5. The five feature points of the standard face are set at the left eye, right eye, nose tip, left corner of the mouth and right corner of the mouth respectively. Then, face detection is performed on the face image to obtain the feature points corresponding to the standard face image, which are the left eye, right eye, nose tip, left corner of the mouth and right corner of the mouth of the face image. The coordinates of the above feature points of the face image are obtained and recorded as Dlandmaks i , i = 1 to 5. According to Slandmaks i With Dlandmaks i , calculate the perspective transformation matrix between each face image and the standard face, and calculate the face image after each face image is aligned with the standard face image based on the perspective transformation matrix. The aligned face image is cropped to the same size of 112*112 pixels as the standard face image to obtain an aligned face dataset, which includes the aligned and cropped face images corresponding to all images in the face image set. The comparison face dataset is recorded as: {image_data_C|ir i ,rgb i ,depth i ,i=1~n};
[0082] Among them, ir i Represents the IR image after alignment and cropping to the specified size, rgb iIndicates the RGB image after the i-th image is aligned and cropped to the specified size, depth i Represents the i-th depth image after alignment and cropping to the specified size.
[0083] S4. Training a multimodal anti-counterfeiting model based on the aligned face dataset to obtain a trained multimodal face anti-counterfeiting model.
[0084] In a preferred embodiment, please refer to FIG2 , which is a schematic diagram of the structure of a multimodal face anti-counterfeiting model provided by one embodiment of the present invention, including: an IR unimodal anti-counterfeiting sub-model, an RGB unimodal anti-counterfeiting sub-model, a depth unimodal anti-counterfeiting sub-model, and a multimodal anti-counterfeiting sub-model; wherein the IR unimodal anti-counterfeiting sub-model, the RGB unimodal anti-counterfeiting sub-model, and the depth unimodal anti-counterfeiting sub-model have the same structure;
[0085] The step of training a multimodal anti-counterfeiting model based on the aligned face dataset to obtain a trained multimodal face anti-counterfeiting model comprises:
[0086] Training the IR unimodal anti-counterfeiting sub-model according to the IR images in the aligned face dataset to obtain a trained IR unimodal anti-counterfeiting sub-model;
[0087] Training the RGB single-modal anti-counterfeiting sub-model according to the RGB images in the aligned face dataset to obtain a trained RGB single-modal anti-counterfeiting sub-model;
[0088] Training the deep unimodal anti-counterfeiting sub-model according to the depth images in the aligned face dataset to obtain a trained deep unimodal anti-counterfeiting sub-model;
[0089] Inputting the IR image of the comparison face dataset into the trained IR unimodal anti-counterfeiting sub-model to obtain the output of the trained IR unimodal anti-counterfeiting sub-model;
[0090] Inputting the RGB image of the comparison face dataset into the trained RGB single-modal anti-counterfeiting sub-model to obtain the output of the trained RGB single-modal anti-counterfeiting sub-model;
[0091] Inputting the depth image of the comparison face dataset into the trained deep unimodal anti-counterfeiting sub-model to obtain the output of the trained deep unimodal anti-counterfeiting sub-model;
[0092] The multimodal anti-counterfeiting sub-model is trained based on the output of the trained IR unimodal anti-counterfeiting sub-model, the output of the trained RGB unimodal anti-counterfeiting sub-model, and the output of the trained depth unimodal anti-counterfeiting sub-model as inputs to obtain a trained multimodal anti-counterfeiting sub-model.
[0093] A trained multimodal face anti-counterfeiting model is generated based on the trained IR unimodal anti-counterfeiting sub-model, the trained RGB unimodal anti-counterfeiting sub-model, the trained depth unimodal anti-counterfeiting sub-model and the trained multimodal anti-counterfeiting sub-model.
[0094] It should be noted that the IR unimodal anti-counterfeiting sub-model, the RGB unimodal anti-counterfeiting sub-model, and the deep unimodal anti-counterfeiting sub-model have the same structure. This is so that when performing feature fusion, the three modalities undergo the same feature extraction steps to ensure the consistency of feature fusion, that is, the fused features do not favor any one modality. The structure adopted is a general backbone, such as mobienet, resnet and other series, and they are aligned and pruned. Since the data collected is aligned facial images, the SE attention mechanism can better capture the key features of the face during feature fusion, thereby improving the performance and robustness of the model. The multimodal anti-counterfeiting sub-module adopts a general backbone and its variant structure.
[0095] In this embodiment, the IR unimodal anti-counterfeiting sub-model (the dotted box in the upper left corner of Figure 2) is trained according to the IR image in the aligned face dataset to obtain the trained IR unimodal anti-counterfeiting sub-model; the RGB unimodal anti-counterfeiting sub-model (the middle dotted box in Figure 2) is trained according to the RGB image in the aligned face dataset to obtain the trained RGB unimodal anti-counterfeiting sub-model; the depth unimodal anti-counterfeiting sub-model (the dotted box in the upper right corner of Figure 2) is trained according to the depth image in the aligned face dataset to obtain the trained depth unimodal anti-counterfeiting sub-model.
[0096] Since the IR feature extraction module, RGB feature extraction module and Depth feature extraction module in the figure are basically able to identify real and fake faces, the anti-counterfeiting identification results have a high degree of credibility. Therefore, when training the multimodal anti-counterfeiting sub-model, it is necessary to remove the fully connected layer outputs of the IR unimodal anti-counterfeiting sub-model, the RGB unimodal anti-counterfeiting sub-model and the depth unimodal anti-counterfeiting sub-model, and then lock the feature extraction weights of the IR unimodal anti-counterfeiting sub-model, the RGB unimodal anti-counterfeiting sub-model and the depth unimodal anti-counterfeiting sub-model. When training the multimodal anti-counterfeiting sub-module based on the aligned face dataset, the specific process is: input the IR image of the comparison face dataset into the trained IR unimodal anti-counterfeiting sub-model , obtain the output of the trained IR unimodal anti-counterfeiting sub-model; input the RGB image of the comparison face dataset into the trained RGB unimodal anti-counterfeiting sub-model to obtain the output of the trained RGB unimodal anti-counterfeiting sub-model; input the depth image of the comparison face dataset into the trained depth unimodal anti-counterfeiting sub-model to obtain the output of the trained depth unimodal anti-counterfeiting sub-model; according to the output of the trained IR unimodal anti-counterfeiting sub-model, the output of the trained RGB unimodal anti-counterfeiting sub-model and the output of the trained depth unimodal anti-counterfeiting sub-model, as the input of the multimodal anti-counterfeiting sub-model, train the multimodal anti-counterfeiting sub-model to obtain the trained multimodal anti-counterfeiting sub-model. According to the trained IR unimodal anti-counterfeiting sub-model, the trained RGB unimodal anti-counterfeiting sub-model, the trained depth unimodal anti-counterfeiting sub-model and the trained multimodal anti-counterfeiting sub-model, generate the trained multimodal face anti-counterfeiting model as shown in Figure 3.
[0097] In a preferred embodiment, the loss functions of the IR unimodal anti-counterfeiting sub-model, the RGB unimodal anti-counterfeiting sub-model, the depth unimodal anti-counterfeiting sub-model and the multimodal anti-counterfeiting sub-model are all binary classification loss functions.
[0098] In this embodiment, the loss functions of the IR unimodal anti-counterfeiting sub-model, the RGB unimodal anti-counterfeiting sub-model, the depth unimodal anti-counterfeiting sub-model, and the multimodal anti-counterfeiting sub-model adopt the cross entropy loss function.
[0099] In a preferred embodiment, please refer to Figure 3, which is a schematic structural diagram of another multimodal face anti-counterfeiting model provided by one embodiment of the present invention, including: an IR unimodal anti-counterfeiting sub-model, an RGB unimodal anti-counterfeiting sub-model, a depth unimodal anti-counterfeiting sub-model, and a multimodal anti-counterfeiting sub-model; wherein the IR unimodal anti-counterfeiting sub-model, the RGB unimodal anti-counterfeiting sub-model, and the depth unimodal anti-counterfeiting sub-model have the same structure; the input of the multimodal anti-counterfeiting sub-model is the output of the IR unimodal anti-counterfeiting sub-model, the RGB unimodal anti-counterfeiting sub-model, and the depth unimodal anti-counterfeiting sub-model;
[0100] The step of training a multimodal anti-counterfeiting model based on the aligned face dataset to obtain a trained multimodal face anti-counterfeiting model comprises:
[0101] The IR image, RGB image and depth image in the aligned face dataset are respectively input into the IR unimodal anti-counterfeiting sub-model, RGB unimodal anti-counterfeiting sub-model and depth unimodal anti-counterfeiting sub-model, and the outputs of the obtained IR unimodal anti-counterfeiting sub-model, RGB unimodal anti-counterfeiting sub-model and depth unimodal anti-counterfeiting sub-model are input into the multimodal anti-counterfeiting sub-model. The IR unimodal anti-counterfeiting sub-model, RGB unimodal anti-counterfeiting sub-model, depth unimodal anti-counterfeiting sub-model and multimodal anti-counterfeiting sub-model are trained at the same time to obtain a trained multimodal face anti-counterfeiting model.
[0102] In this embodiment, there is no need to train the IR unimodal anti-counterfeiting sub-model, the RGB unimodal anti-counterfeiting sub-model and the depth unimodal anti-counterfeiting sub-model separately. Instead, the entire model shown in FIG3 is trained based on the aligned face dataset.
[0103] In a preferred embodiment, the loss function of the multimodal face anti-counterfeiting model is: loss = p i ×LOSSI+p r ×LOSSR+p d ×LOSSD+p m ×LOSSM;
[0104] Among them, LOSSI is the loss value of the IR unimodal anti-counterfeiting sub-model, LOSSR is the loss value of the RGB unimodal anti-counterfeiting sub-model, LOSSD is the loss value of the depth unimodal anti-counterfeiting sub-model, LOSSM is the loss value of the multimodal anti-counterfeiting sub-model, and p i is the weighted coefficient of the loss value of the preset IR single-modal anti-counterfeiting sub-model, p r is the weighted coefficient of the loss value of the preset RGB single-modal anti-counterfeiting sub-model, p d is the weighted coefficient of the loss value of the preset deep single-modal anti-counterfeiting sub-model, p m is the weighting coefficient of the preset multimodal sub-model loss value.
[0105] In this embodiment, the weighting coefficient p i is 0.1, p r is 0.1, p d is 0.1, p m is 0.7.
[0106] Example 2
[0107] Please refer to FIG4 , which is a flowchart of a face anti-counterfeiting recognition method provided by one embodiment of the present invention, including:
[0108] Acquire facial images captured by a binocular camera; wherein the dual-mode module includes an IR lens and an RGB lens; the facial image set includes a plurality of IR images and an RGB image corresponding to each IR image captured at the same time and scene;
[0109] The facial image is input into a multimodal facial anti-counterfeiting model so that the multimodal facial anti-counterfeiting model performs anti-counterfeiting identification on the facial image to obtain an anti-counterfeiting identification result; wherein, the multimodal facial anti-counterfeiting model is trained according to the above-mentioned multimodal facial anti-counterfeiting model training method.
[0110] The specific embodiments described above further illustrate the objectives, technical solutions, and beneficial effects of the present invention. It should be understood that the above descriptions are merely specific embodiments of the present invention and are not intended to limit the scope of protection of the present invention. In particular, it should be noted that any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included within the scope of protection of the present invention for those skilled in the art.
Claims
1. A multimodal face anti-counterfeiting model training method, characterized in that: include: A plurality of face images are collected by a binocular camera as a face image set; wherein the dual-mode module includes an IR lens and an RGB lens arranged in left-right alignment; the face image set includes a plurality of IR images and an RGB image collected at the same time and scene corresponding to each IR image; According to each group of IR images and corresponding RGB images in the face image set, a corresponding depth image is calculated, and the obtained depth image is added to the face image set to obtain an updated face image set; Performing face alignment on all images in the updated face image set to obtain an aligned face data set; A multimodal anti-counterfeiting model is trained according to the aligned face data set to obtain a trained multimodal face anti-counterfeiting model.
2. The multimodal face anti-counterfeiting model training method according to claim 1, characterized in that: The step of calculating a corresponding depth image according to each group of IR images and the corresponding RGB image in the face image set includes: Preprocessing the IR image and the RGB image to obtain a processed IR image and RGB image; Performing image stereo correction on the processed IR image and the RGB image to obtain a corrected IR image and RGB image; Through the facial feature point detection algorithm, the feature points of each group of corrected IR images and the corresponding RGB images are obtained; Calculating a matching matrix of the IR image and the corresponding RGB image according to feature points of the IR image and the corresponding RGB image, and obtaining a corresponding relationship between pixels of the IR image and the corresponding RGB image according to the matching matrix; Calculating a disparity map of the IR image and the corresponding RGB image according to a correspondence relationship between pixels of the IR image and the corresponding RGB image; A corresponding depth map is calculated according to the disparity map of the IR image and the corresponding RGB image.
3. The multimodal face anti-counterfeiting model training method according to claim 2, characterized in that: The step of calculating a corresponding depth map according to the disparity map of the IR image and the corresponding RGB image includes: Determine the disparity values of corresponding pixels in the IR image and the corresponding RGB image according to the disparity maps of the IR image and the corresponding RGB image; Obtain the focal length of the binocular camera, the baseline distance between the left and right cameras, and the horizontal coordinates of the pixel points in the IR image and the corresponding RGB image; The corresponding depth map is calculated according to the disparity value of the corresponding pixel point, the focal length of the binocular camera, the baseline distance between the left and right cameras, and the horizontal coordinate of the pixel point.
4. The multimodal face anti-counterfeiting model training method according to claim 1, characterized in that: The step of performing face alignment on all images in the updated face image set to obtain an aligned face data set includes: Performing facial feature point detection on all images in the updated facial image set to obtain feature point coordinates corresponding to each facial image and a standard facial image; wherein the standard facial image is a preset facial image with specified feature points; Calculate the corresponding perspective transformation matrix according to the coordinates of the feature points corresponding to each face image and the standard face and the coordinates of the feature points of the standard face; calculate the face image after each face image is aligned according to the perspective transformation matrix; The aligned face image is cropped to the same size as the standard face image to obtain an aligned face dataset.
5. The multimodal face anti-counterfeiting model training method according to claim 1, characterized in that: The multimodal anti-counterfeiting model includes: an IR unimodal anti-counterfeiting sub-model, an RGB unimodal anti-counterfeiting sub-model, a deep unimodal anti-counterfeiting sub-model and a multimodal anti-counterfeiting sub-model; wherein the IR unimodal anti-counterfeiting sub-model, the RGB unimodal anti-counterfeiting sub-model and the deep unimodal anti-counterfeiting sub-model have the same structure; The step of training a multimodal anti-counterfeiting model according to the aligned face data set to obtain a trained multimodal face anti-counterfeiting model includes: According to the IR images in the aligned face dataset, the IR single-modal anti-counterfeiting sub-model is trained to obtain a trained IR single-modal anti-counterfeiting sub-model; According to the RGB images in the aligned face dataset, the RGB single-modal anti-counterfeiting sub-model is trained to obtain a trained RGB single-modal anti-counterfeiting sub-model; According to the depth image in the aligned face dataset, training the deep single-modal anti-counterfeiting sub-model to obtain a trained deep single-modal anti-counterfeiting sub-model; Inputting the IR image of the comparison face data set into the trained IR single-modal anti-counterfeiting sub-model to obtain the output of the trained IR single-modal anti-counterfeiting sub-model; Inputting the RGB image of the comparison face data set into the trained RGB single-modal anti-counterfeiting sub-model to obtain the output of the trained RGB single-modal anti-counterfeiting sub-model; Inputting the depth image of the comparison face data set into the trained deep single-modal anti-counterfeiting sub-model to obtain the output of the trained deep single-modal anti-counterfeiting sub-model; According to the output of the trained IR unimodal anti-counterfeiting sub-model, the output of the trained RGB unimodal anti-counterfeiting sub-model and the output of the trained depth unimodal anti-counterfeiting sub-model, as the input of the multimodal anti-counterfeiting sub-model, the multimodal anti-counterfeiting sub-model is trained to obtain a trained multimodal anti-counterfeiting sub-model; A trained multimodal face anti-counterfeiting model is generated according to the trained IR unimodal anti-counterfeiting sub-model, the trained RGB unimodal anti-counterfeiting sub-model, the trained depth unimodal anti-counterfeiting sub-model and the trained multimodal anti-counterfeiting sub-model.
6. The multimodal face anti-counterfeiting model training method according to claim 5, characterized in that: The loss functions of the IR unimodal anti-counterfeiting sub-model, the RGB unimodal anti-counterfeiting sub-model, the depth unimodal anti-counterfeiting sub-model and the multimodal anti-counterfeiting sub-model are all binary classification loss functions.
7. The multimodal face anti-counterfeiting model training method according to claim 1, characterized in that: The multimodal anti-counterfeiting model includes: an IR unimodal anti-counterfeiting sub-model, an RGB unimodal anti-counterfeiting sub-model, a deep unimodal anti-counterfeiting sub-model and a multimodal anti-counterfeiting sub-model; wherein the IR unimodal anti-counterfeiting sub-model, the RGB unimodal anti-counterfeiting sub-model and the deep unimodal anti-counterfeiting sub-model have the same structure; the input of the multimodal anti-counterfeiting sub-model is the output of the IR unimodal anti-counterfeiting sub-model, the RGB unimodal anti-counterfeiting sub-model and the deep unimodal anti-counterfeiting sub-model; The step of training a multimodal anti-counterfeiting model according to the aligned face data set to obtain a trained multimodal face anti-counterfeiting model includes: The IR image, RGB image and depth image in the aligned face dataset are respectively input into the IR unimodal anti-counterfeiting sub-model, RGB unimodal anti-counterfeiting sub-model and deep unimodal anti-counterfeiting sub-model, and the outputs of the IR unimodal anti-counterfeiting sub-model, RGB unimodal anti-counterfeiting sub-model and deep unimodal anti-counterfeiting sub-model are input into the multimodal anti-counterfeiting sub-model. The IR unimodal anti-counterfeiting sub-model, RGB unimodal anti-counterfeiting sub-model, deep unimodal anti-counterfeiting sub-model and multimodal anti-counterfeiting sub-model are trained simultaneously to obtain a trained multimodal face anti-counterfeiting model.
8. The multimodal face anti-counterfeiting model training method according to claim 7, characterized in that: The loss function of the multimodal face anti-counterfeiting model is: loss = p i ×LOSSI+p r ×LOSSR+p d ×LOSSD+p m ×LOSSM; Among them, LOSSI is the loss value of the IR unimodal anti-counterfeiting sub-model, LOSSR is the loss value of the RGB unimodal anti-counterfeiting sub-model, LOSSD is the loss value of the deep unimodal anti-counterfeiting sub-model, LOSSM is the loss value of the multimodal anti-counterfeiting sub-model, and p i is the weighted coefficient of the loss value of the preset IR single-modal anti-counterfeiting sub-model, p r is the weighted coefficient of the loss value of the preset RGB single-modal anti-counterfeiting sub-model, p d is the weighted coefficient of the loss value of the preset deep single-modal anti-counterfeiting sub-model, p m It is the weighting coefficient of the preset multimodal sub-model loss value.
9. A face anti-counterfeiting recognition method, characterized in that: include: Acquire a facial image captured by a binocular camera; wherein the dual-mode module includes an IR lens and an RGB lens; the facial image set includes a plurality of IR images and an RGB image corresponding to each IR image captured at the same time and in the same scene; The facial image is input into a multimodal facial anti-counterfeiting model so that the multimodal facial anti-counterfeiting model performs anti-counterfeiting identification on the facial image to obtain an anti-counterfeiting identification result; wherein the multimodal facial anti-counterfeiting model is trained according to the multimodal facial anti-counterfeiting model training method described in any one of claims 1 to 8.
Citation Information
Patent Citations
Method and device for fusing 2D face detection and 3D face recognition
CN111523398A
Face verification method and device, computer equipment and storage medium
CN112232323A
Face biopsy method, system and equipment based on binocular camera distance measurement and medium
CN115063339A
Face anti-counterfeiting method and device and storage medium
CN116798130A
Photosensitive resin composition, cured product, printed wiring board and method for manufactureing printed wiring board
KR1020220134470A
Cited By
Vehicle-mounted fraud identification method and device based on face detection, equipment and medium
CN121392923A
Cross-modal anti-counterfeiting detection model training method, face anti-counterfeiting detection method and related device
CN121415452A