Method, system and device for generating speckle face recognition model by using infrared face recognition model, and storage medium

By constructing a multimodal teacher model and knowledge distillation technology, the speckle feature knowledge of the infrared face recognition model is transferred to a lightweight student model, which solves the problems of difficulty and high cost in training the speckle face recognition model, and realizes the efficient generation of high-quality speckle face recognition models at low cost.

CN120635957APending Publication Date: 2025-09-12SHENZHEN GUANGJIAN TECH CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510690146.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-27
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

In the existing technology, the training of speckle face recognition models is difficult and costly, and it is difficult to obtain an effective recognition model with a small number of speckle images.

Method used

A multimodal teacher model is constructed using the infrared face recognition model, and the speckle feature knowledge is transferred to a lightweight student model through knowledge distillation technology, including fine-tuning the teacher model with infrared images and speckle images aligned with the training set, designing a multi-task loss function, using a deep separable convolutional structure and batch normalization technology, and performing feature distillation to generate a high-quality speckle face recognition model.

Benefits of technology

With a smaller number of speckle images, a high-quality speckle face recognition model can be quickly obtained, which reduces training cost and time and improves recognition efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120635957A_ABST
    Figure CN120635957A_ABST
Patent Text Reader

Abstract

A method, system and device for generating a speckle face recognition model by using an infrared face recognition model, and a storage medium, the method comprising: step S1, obtaining an infrared face recognition model, and performing fine tuning on the infrared face recognition model through a training set to obtain a teacher face model; wherein the training set comprises an infrared image and a speckle image which are aligned; the output of the teacher face model comprises an infrared feature vector and a speckle feature vector; s2, establishing an initial student model, and aligning the output of the initial student model with the speckle feature vector; and S3, distilling an output result of the initial student model by using the speckle feature vector to obtain a final student model. According to the method, the high-quality speckle face recognition model is quickly obtained by using the high-stability infrared face recognition model, and the method has important practical significance and application value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of face recognition model training, and in particular to a method, system, device and storage medium for generating a speckle face recognition model by using an infrared face recognition model. Background Art

[0002] As one of the core technologies in biometric identification, facial recognition has been widely used in security surveillance, identity authentication, smart terminals, and other fields. Currently, various facial images can be acquired through various techniques, such as infrared images, infrared images, and speckle images. Different images contain varying densities of facial information, and therefore pose varying degrees of difficulty in training corresponding facial recognition models.

[0003] In contrast, speckle images have the lowest information density, making them the most challenging to train for face recognition models, and their success rate is also lower. Training typically requires a very large number of images. However, speckle image data is typically limited, and acquisition requires specialized equipment, making it expensive. Consequently, no existing technology has yet successfully trained an effective speckle face recognition model using a relatively small number of speckle images at a low cost.

[0004] The disclosure of the above background technology content is only used to assist in understanding the inventive concept and technical solution of the present invention. It does not necessarily belong to the prior art of this patent application. In the absence of clear evidence that the above content has been disclosed on the filing date of this patent application, the above background technology should not be used to evaluate the novelty and creativity of this application. Summary of the Invention

[0005] To this end, the present invention proposes a method for generating a speckle face recognition model using an infrared face recognition model. By utilizing the existing infrared face recognition model, a multimodal teacher model is constructed and its speckle feature knowledge is distilled into a lightweight student model. In the case of a small number of speckle images, a high-quality speckle face recognition model can be quickly obtained using a highly stable infrared face recognition model, which has important practical significance and application value.

[0006] In a first aspect, the present invention provides a method for generating a speckle face recognition model using an infrared face recognition model, characterized by comprising:

[0007] Step S1: obtaining an infrared face recognition model, and fine-tuning the infrared face recognition model using a training set to obtain a teacher face model; wherein the training set includes aligned infrared images and speckle images; and the output of the teacher face model includes infrared feature vectors and speckle feature vectors;

[0008] Step S2: establishing an initial student model, and aligning the output of the initial student model with the speckle feature vector;

[0009] Step S3: using the speckle feature vector to distill the output result of the initial student model to obtain a final student model.

[0010] Optionally, the method for generating a speckle face recognition model using an infrared face recognition model is characterized in that, in the training set, the infrared image and the speckle image aligned in step S1 are obtained by:

[0011] Step S11: using a multimodal sensor to synchronously capture an infrared image and a speckle image of the same face;

[0012] Step S12: aligning the spatial coordinates of the speckle image with the infrared image through affine transformation.

[0013] Optionally, the method for generating a speckle face recognition model using an infrared face recognition model is characterized in that the training set also includes a depth image obtained from the speckle image or the infrared image, and in step S3, the depth image is also used as auxiliary supervision for the distillation of the initial student model.

[0014] Optionally, the method for generating a speckle face recognition model using an infrared face recognition model is characterized in that, in step S1, the fine-tuning process includes:

[0015] Step S13: freezing part of the bottom convolutional layers of the infrared face recognition model;

[0016] Step S14: designing a multi-task loss function to simultaneously optimize the extraction of infrared feature vectors and speckle feature vectors;

[0017] Step S15: Batch normalization and regularization techniques are used to prevent overfitting.

[0018] Optionally, the method for generating a speckle face recognition model using an infrared face recognition model is characterized in that the initial student model is a lightweight network, adopts a depthwise separable convolutional structure, and the output dimension of the last fully connected layer remains consistent with the speckle feature vector.

[0019] Optionally, the method for generating a speckle face recognition model using an infrared face recognition model is characterized in that, in step S3, the distillation process includes:

[0020] Step S31: calculating a similarity loss function between the output of the initial student model and the speckle feature vector of the teacher face model;

[0021] Step S32: Optimizing the initial student model using the similarity loss function so that the output of the final student model is closer to the speckle feature vector.

[0022] Optionally, the method for generating a speckle face recognition model using an infrared face recognition model is characterized in that the similarity loss function includes but is not limited to mean square error loss (MSE), cosine similarity loss or KL divergence loss.

[0023] In a second aspect, the present invention provides a system for generating a speckle face recognition model using an infrared face recognition model, which is used to implement any of the aforementioned methods for generating a speckle face recognition model using an infrared face recognition model, and is characterized by comprising:

[0024] a teacher training module for obtaining an infrared face recognition model, and fine-tuning the infrared face recognition model using a training set to obtain a teacher face model; wherein the training set includes aligned infrared images and speckle images; and the output of the teacher face model includes infrared feature vectors and speckle feature vectors;

[0025] a student establishment module, configured to establish an initial student model and align an output of the initial student model with the speckle feature vector;

[0026] A distillation module is used to distill the output result of the initial student model using the speckle feature vector to obtain a final student model.

[0027] In a third aspect, the present invention provides a device for generating a speckle face recognition model using an infrared face recognition model, characterized in that it includes:

[0028] processor;

[0029] a memory storing executable instructions for the processor;

[0030] The processor is configured to execute the executable instructions to perform any of the steps of the aforementioned method for generating a speckle face recognition model using an infrared face recognition model.

[0031] In a fourth aspect, the present invention provides a computer-readable storage medium for storing a program, characterized in that when the program is executed, the steps of any of the aforementioned methods for generating a speckle face recognition model using an infrared face recognition model are implemented.

[0032] Compared with the prior art, the present invention has the following beneficial effects:

[0033] In this method, the training set contains aligned infrared and speckle images, and the teacher face model outputs both infrared and speckle feature vectors. The infrared face recognition model is fine-tuned using the training set to obtain a teacher face model capable of recognizing speckle images. A distillation technique is then used to obtain a final student model capable of recognizing speckle images.

[0034] The present invention utilizes an existing high-quality infrared image face recognition model, enabling the model to quickly recognize speckle images with a relatively small number of training sets. A high-quality speckle image recognition model, namely the final student model, is then obtained through distillation technology. By utilizing the existing model to quickly obtain a high-quality speckle face recognition model, the requirement for the number of speckle images and the training time required for speckle face recognition model training are greatly reduced. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for the embodiments or the description of the prior art. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without inventive work. Other features, purposes and advantages of the present invention will become more apparent by reading the detailed description of the non-limiting embodiments with reference to the following drawings:

[0036] Figure 1 This is a flowchart of a method for generating a speckle face recognition model using an infrared face recognition model according to an embodiment of the present invention;

[0037] Figure 2 This is a flow chart of the steps for obtaining an infrared image and a speckle image in an embodiment of the present invention;

[0038] Figure 3 This is a flow chart of a fine-tuning step in an embodiment of the present invention;

[0039] Figure 4 This is a flow chart of the steps of distillation in an embodiment of the present invention;

[0040] Figure 5 This is a structural diagram of a system for generating a speckle face recognition model using an infrared face recognition model according to an embodiment of the present invention;

[0041] Figure 6 A schematic structural diagram of a device for generating a speckle face recognition model using an infrared face recognition model according to an embodiment of the present invention; and

[0042] Figure 7 Schematic diagram of the structure of a computer-readable storage medium in an embodiment of the present invention. DETAILED DESCRIPTION

[0043] The present invention will be described in detail below with reference to specific embodiments. The following embodiments will help those skilled in the art to further understand the present invention, but are not intended to limit the present invention in any form. It should be noted that, for those skilled in the art, several variations and improvements can be made without departing from the scope of the present invention. These all fall within the scope of protection of the present invention.

[0044] The terms "first," "second," "third," "fourth," and the like (if any) in the description and claims of the present invention and in the accompanying drawings are used to distinguish similar objects and are not necessarily used to describe a particular order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate, so that the embodiments of the invention described herein, for example, can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having," and any variations thereof, are intended to cover non-exclusive inclusions, for example, a process, method, system, product, or apparatus comprising a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to these processes, methods, products, or apparatus.

[0045] The embodiment of the present invention provides a method for generating a speckle face recognition model by using an infrared face recognition model, aiming to solve the problems existing in the prior art.

[0046] The following describes in detail the technical solutions of the present invention and how the technical solutions of this application solve the above-mentioned technical problems using specific embodiments. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described in detail in some embodiments. The following embodiments of the present invention are described in conjunction with the accompanying drawings.

[0047] The present invention utilizes the existing infrared face recognition model, constructs a multimodal teacher model and distills its speckle feature knowledge into a lightweight student model. When the number of speckle images is small, a high-quality speckle face recognition model can be quickly obtained using a highly stable infrared face recognition model, which has important practical significance and application value.

[0048] Figure 1 FIG1 is a flow chart of the steps of a method for generating a speckle face recognition model using an infrared face recognition model according to an embodiment of the present invention. Figure 1 As shown, in an embodiment of the present invention, a method for generating a speckle face recognition model by using an infrared face recognition model includes the following steps:

[0049] Step S1: Obtain an infrared face recognition model, and fine-tune the infrared face recognition model through a training set to obtain a teacher face model; wherein the training set includes aligned infrared images and speckle images; the output of the teacher face model includes infrared feature vectors and speckle feature vectors.

[0050] In this step, the existing infrared face recognition model is fine-tuned through the training set so that it can simultaneously output infrared image feature vectors and speckle image feature vectors, serving as a "teacher" model for subsequent knowledge distillation.

[0051] The infrared face recognition model can be a public pre-trained model (such as ResNet, MobileNet, etc.) or a self-developed model. In the initial state, it is only trained for the face recognition task of infrared images and outputs the feature vector of the infrared image (for identity recognition).

[0052] The training set consists of paired aligned infrared and speckle images—that is, an infrared image and a speckle pattern of the same object from the same viewing angle. (Speckle patterns are typically generated by structured light or laser projectors for depth perception or anti-counterfeiting purposes.) The two images must be strictly aligned at the pixel level (e.g., by calibrating the intrinsic and extrinsic parameters of the camera and projector) to ensure that the same facial region is positioned consistently in both modalities. The training set must contain a sufficient number of samples (covering scenarios with varying lighting, posture, expression, and other characteristics) to ensure model generalization.

[0053] Model fine-tuning includes the following:

[0054] Network structure modification: Based on the original infrared model, some network layers are added or reused to enable it to process speckle image input simultaneously, or the feature extraction layer is shared in a single network to output feature vectors of the two modalities respectively.

[0055] Loss Function Design: For the infrared image branch, traditional face recognition losses (such as ArcFace Loss and Triplet Loss) are used to constrain the intra-class compactness and inter-class separability of the infrared feature vector. For the speckle image branch, face recognition losses are also used to ensure the discriminative power of the speckle feature vector. A cross-modal consistency loss (such as L2 distance constraint) may be introduced to force the infrared and speckle features of the same object to be close in feature space, thereby enhancing their correlation.

[0056] Training process: The network parameters are optimized through back-propagation so that the teacher model can output effective recognition features in both infrared and speckle modalities.

[0057] Step S2: establishing an initial student model, and aligning the output of the initial student model with the speckle feature vector.

[0058] In this step, a lightweight “student” model is constructed so that the feature vector of its output is aligned with the speckle feature vector of the teacher model in the feature space, laying the foundation for knowledge distillation.

[0059] The initial student model adopts a lightweight design, using a network structure with fewer parameters (such as MobileNetV3 and ShuffleNet), or performs channel pruning and layer reduction on the teacher model to meet actual deployment requirements (such as embedded devices and real-time inference). The student model only takes the speckle image as input (because the goal is to generate a speckle face recognition model) and outputs a speckle feature vector.

[0060] Fixed teacher model parameters: After the training in step S1 is completed, the parameters of the teacher model remain unchanged and serve as the "knowledge source" for knowledge distillation.

[0061] Forward propagation to obtain speckle features: The speckle images in the training set are input into the teacher model to extract the corresponding speckle feature vector.

[0062] Initialize the student model output: Through random initialization or transfer learning (such as loading part of the weights of the teacher model speckle branch), make the initial output features of the student model as close as possible to the distribution of the speckle feature vector (such as through Gaussian distribution matching or feature space projection).

[0063] Step S3: using the speckle feature vector to distill the output result of the initial student model to obtain a final student model.

[0064] In this step, the speckle feature discrimination capability contained in the teacher model is transferred to the student model through knowledge distillation technology, so that it can achieve high-precision face recognition when only speckle images are input.

[0065] During distillation, the speckle images from the training set are used as input to the student model, and the speckle feature vectors output by the teacher model are used as “soft labels” to guide the feature learning of the student model.

[0066] Feature distillation loss: measures the difference between the output features of the student model and the speckle features of the teacher model, commonly used mean square error (MSE) or cosine similarity loss L cosine .

[0067] Classification loss: If the teacher model also outputs category labels (such as training with classification loss in step S1), a cross-entropy loss Lc can be introduced to enhance the classification ability of the student model using the teacher model’s category prediction probability (soft label) or true label (hard label).

[0068] Total loss: Ltotal = L cosine +α·Lc, where α is the balance coefficient, which is adjusted according to task requirements.

[0069] Training strategies include:

[0070] End-to-end training: The student model simultaneously optimizes feature distillation loss and classification loss during training, gradually approaching the speckle feature distribution and recognition ability of the teacher model.

[0071] Data augmentation: Data augmentation methods such as random rotation, scaling, and noise addition are applied to the speckle images to improve the robustness of the student model.

[0072] Learning rate scheduling: Use strategies such as cosine annealing and exponential decay to avoid gradient explosion or slow convergence caused by large feature differences in the early stages of training.

[0073] After training, the student model can independently receive speckle image input and output discriminative feature vectors for face recognition tasks (such as feature comparison and identity verification).

[0074] Figure 2 FIG. 1 is a flow chart of steps for obtaining infrared images and speckle images according to an embodiment of the present invention. Figure 2 As shown, in an embodiment of the present invention, a step of obtaining an infrared image and a speckle image includes:

[0075] Step S11: using a multimodal sensor to synchronously capture an infrared image and a speckle image of the same face.

[0076] In this step, the infrared image and speckle image of the same face at the same time are obtained through hardware synchronization technology to eliminate the differences in posture and expression caused by time difference.

[0077] Sensor Type:

[0078] Infrared camera: Typically uses a near-infrared (NIR, wavelength 700-1000nm) or mid-infrared (MIR) camera, which needs to be equipped with an infrared light source (such as an 850nm or 940nm LED array) to provide active illumination and enhance the contrast of facial features.

[0079] Speckle projector + camera: The structured light speckle system (such as Microsoft Kinect and Apple TrueDepth) projects a speckle beam with a specific pattern and uses a camera to capture the reflected light to calculate depth information or generate a speckle image.

[0080] Sensor layout: The infrared camera and speckle camera must be rigidly fixed on the same device to ensure stable relative position and angle and reduce deviation caused by mechanical vibration.

[0081] Synchronization mechanism

[0082] Hardware triggering: Use an external synchronization signal (such as a GPIO pulse) to trigger the shutters of two cameras simultaneously, achieving microsecond-level synchronization (time difference is usually <1ms).

[0083] Timestamp alignment: If hardware synchronization is not feasible, a high-precision clock (such as the PTP protocol) can be used to add a timestamp to each frame of the image. Later, a time matching algorithm can be used to select the image pair with the smallest time difference.

[0084] Collection process

[0085] Data sampling: Multiple sets of facial data are collected under different lighting conditions (such as strong light, weak light, and backlight), posture angles (±30° pitch / yaw), and expressions (neutral, smiling, and frowning) to ensure data diversity.

[0086] Calibration plate acquisition: Synchronously acquire infrared and speckle images of a calibration plate (such as a checkerboard) for subsequent calibration of camera intrinsic and extrinsic parameters.

[0087] Step S12: aligning the spatial coordinates of the speckle image with the infrared image through affine transformation.

[0088] In this step, the pixel coordinates of the speckle image are mapped to the infrared image coordinate system through affine transformation to achieve accurate alignment of facial features in the two modalities.

[0089] Feature point detection and matching

[0090] Feature point extraction:

[0091] In infrared images, a facial key point detector (such as MTCNN, RetinaFace) is used to extract 5 or 68 feature points (such as corners of the eyes, tip of the nose, corners of the mouth).

[0092] In the speckle image, the corresponding key points are extracted through the depth information or the feature point detection algorithm of the speckle pattern (such as SIFT and ORB).

[0093] Feature matching: Use the bidirectional nearest neighbor matching (NNDR) or RANSAC algorithm to filter out mismatched points to ensure the accuracy of feature point pairs.

[0094] Image resampling and alignment

[0095] Transformation application: The estimated affine transformation matrix is ​​applied to all pixels of the speckle image, and the aligned speckle image is generated using bilinear interpolation or bicubic interpolation.

[0096] Boundary processing: The aligned speckle image may have edge holes, which can be processed by zero filling or mirror filling.

[0097] Alignment Verification

[0098] Visual inspection: Overlay the aligned speckle image and the infrared image (e.g., using different levels of transparency) to see if the facial contours and feature points overlap.

[0099] Quantitative evaluation: Calculate the average Euclidean distance (e.g. <1 pixel) or mean square error (MSE) of feature point pairs to evaluate alignment accuracy.

[0100] In some embodiments, the training set also includes a depth image obtained from the speckle image or the infrared image. In step S3, the depth image is also used as auxiliary supervision for the distillation of the initial student model.

[0101] There are three auxiliary supervision methods for depth images in knowledge distillation:

[0102] 1. Feature-level auxiliary supervision

[0103] Deep feature extraction: Add a deep branch (such as a lightweight CNN) to the student model to predict deep features from the speckle image.

[0104] Deep Consistency Loss: Constraining the Deep Features Predicted by the Student Model Features extracted from real depth images Consistency:

[0105]

[0106] Where λ is the balance coefficient.

[0107] 2. Structure-aware feature distillation

[0108] Depth-guided feature alignment:

[0109] Use depth information as weight to enhance the feature distillation of sensitive facial structure areas (such as nose bridge and eye sockets). For example, when calculating the feature distillation loss L distill hour: The weight w(p) is determined by the gradient magnitude of pixel p in the depth map (the edge area with large gradient has a higher weight).

[0110] 3. Multi-task Distillation Framework

[0111] Joint loss function: Combine deep supervision with original speckle feature distillation and classification loss: L total =L distill +α·L cls +β·L depth Among them, α and β are hyperparameters that control the weight of each loss.

[0112] Multi-branch structure: The student model simultaneously outputs speckle feature vectors, depth predictions, and identity classification results, sharing the underlying feature extraction layer to improve parameter utilization efficiency.

[0113] This embodiment has the following advantages:

[0114] Enhanced 3D structural perception: Depth information helps the model distinguish between real faces and photos / masks, improving anti-counterfeiting capabilities.

[0115] Posture robustness: Deep features are more invariant to posture changes, alleviating the recognition degradation caused by profile or pitch angles.

[0116] Feature complementarity: The high-frequency details of the speckle image complement the structural information of the depth image, improving the feature expression capability.

[0117] Figure 3 This is a flow chart of a fine-tuning step in an embodiment of the present invention. Figure 3 As shown, a fine-tuning step in an embodiment of the present invention includes:

[0118] Step S13: Freeze some bottom convolutional layers of the infrared face recognition model.

[0119] In this step, the general visual features (such as edges and textures) extracted by the infrared model at the bottom layer are retained to avoid catastrophic forgetting caused by excessive updates on small datasets; at the same time, the high layer is allowed to adapt to the characteristics of the speckle image.

[0120] Bottom-level features (such as the first 2-3 convolutional blocks): usually responsible for extracting low-level features such as general edges and corners, which are universal for all visual tasks and are frozen.

[0121] Mid-level features (intermediate convolutional blocks): Partially freeze or reduce the learning rate to allow limited adjustment to the characteristics of the speckle image (such as the texture characteristics of the speckle pattern).

[0122] High-level features (fully connected layer or last convolutional block): completely unfrozen, because high-level features are highly relevant to the specific task (face recognition), and discriminative features need to be relearned for speckle images.

[0123] Two-stage training:

[0124] Phase 1: Only train the unfrozen high layers to stabilize the speckle feature extraction capability.

[0125] The second stage: unfreeze some middle layers and perform fine-tuning to balance the feature expressions of the two modalities.

[0126] Learning rate scheduling: Set different learning rates for different layers (e.g. bottom layer lr = 0, middle layer lr = 1e-5, high layer lr = 1e-4).

[0127] Step S14: Design a multi-task loss function to simultaneously optimize the extraction of infrared feature vectors and speckle feature vectors.

[0128] In this step, through joint optimization, the model simultaneously learns the feature extraction of infrared images and speckle images, ensuring that the feature vectors of the two modalities are discriminative and semantically aligned.

[0129] Multi-task loss function design Ltotal=LIRid+Lspecid+λ·Lalign

[0130] Identity loss:

[0131] LIRid: Identity classification loss for infrared images (such as cross entropy loss or ArcFace loss).

[0132] Lspecid: Identity classification loss for speckle images, ensuring that speckle features have identity discrimination capabilities.

[0133] Feature alignment loss:

[0134] Lalign: constrains the distance between the infrared feature vector FIR and the speckle feature vector Fspec of the same object in the feature space.

[0135] Optionally add Center Loss to further constrain the intra-class distance.

[0136] Balance coefficient optimization

[0137] Use adaptive weighting methods (such as GradNorm) to automatically adjust the weights of each loss to prevent one task from dominating the training.

[0138] In experiments, we usually start from λ=0.1 and gradually adjust it to 0.5-1.0, and determine the optimal value based on the performance of the validation set.

[0139] Step S15: Batch normalization and regularization techniques are used to prevent overfitting.

[0140] In this step, the overfitting problem during small sample training is alleviated, the generalization ability of the model is improved, and the training convergence is accelerated.

[0141] Batch Normalization

[0142] Fine-tuning strategy:

[0143] Freeze the BN parameters of the pre-training layer: maintain the original mean and variance statistics to avoid destroying the learned feature distribution.

[0144] Add a trainable BN layer for the newly added layer: adapt to the statistical characteristics of the speckle image.

[0145] Regularization techniques

[0146] L2 regularization: Set weight decay (such as weight_decay = 1e-4) in the optimizer to constrain the norm of model parameters.

[0147] Dropout: Add a Dropout layer (e.g., p = 0.5) before the fully connected layer to randomly drop neurons to reduce feature co-adaptation.

[0148] Label smoothing: Introduce label smoothing (such as smoothing = 0.1) in the cross entropy loss to reduce the model's overconfidence in the training data.

[0149] Data augmentation

[0150] Cross-modal enhancement: synchronously apply transformations such as rotation, scaling, and flipping to infrared and speckle images to maintain semantic consistency between modalities.

[0151] Noise Injection: Add shot noise or Gaussian noise to the speckle image to simulate interference in real applications.

[0152] Through the above fine-tuning strategy, the teacher model can effectively integrate the feature extraction capabilities of the infrared and speckle modalities, providing high-quality supervision signals for subsequent knowledge distillation.

[0153] In some embodiments, the initial student model is a lightweight network that adopts a depthwise separable convolutional structure, and the output dimension of the last fully connected layer remains consistent with the speckle feature vector.

[0154] Depthwise separable convolution decomposes the standard convolution into two steps:

[0155] Depthwise Conv: Perform spatial convolution on each input channel separately, with a computational cost of Cin×K×K×H×W.

[0156] Pointwise Conv: Use 1×1 convolution to integrate channel information, with a computational cost of Cin×Cout×H×W.

[0157] Depthwise separable convolution has a small number of parameters and is suitable for deployment on embedded devices or mobile terminals. It has high computational efficiency, reduces memory access, and accelerates inference speed. It has a regularization effect, and the decomposition operation increases the nonlinear expression ability of the model and alleviates overfitting.

[0158] The architecture of the student model:

[0159] Input: Single-channel or three-channel speckle image (size is usually 112×112 or 224×224).

[0160] Backbone network: A lightweight architecture based on MobileNetV3 or ShuffleNetV2, using depth-wise separable convolution instead of standard convolution.

[0161] Output layer: global average pooling + fully connected layer, the output dimension is consistent with the speckle feature vector of the teacher model (usually 512 or 1024 dimensions).

[0162] Figure 4 FIG. 1 is a flow chart of the steps of distillation in an embodiment of the present invention. Figure 4 As shown, a distillation step in an embodiment of the present invention includes:

[0163] Step S31: Calculating a similarity loss function between the output of the initial student model and the speckle feature vector of the teacher face model.

[0164] In this step, the speckle feature vector F output by the student model is measured student The speckle feature vector F output by the teacher model teacher The similarity between them guides the student model to learn the feature expression pattern of the teacher model.

[0165] The loss function can be based on distance, distribution or mixture.

[0166] Distance-based loss function

[0167] Mean Squared Error Loss (MSE): Directly minimize the Euclidean distance between feature vectors, which is suitable for scenarios with similar feature space distribution.

[0168] Cosine similarity loss: It focuses on the directional consistency of feature vectors, is insensitive to changes in vector scale, and is suitable for scenarios after feature normalization.

[0169] Distribution-based loss functions

[0170] KL divergence loss: After converting the feature vector into a probability distribution (such as through softmax), the KL divergence is calculated: Among them, P teacher and P student is the probability distribution of the eigenvectors after temperature scaling.

[0171] Hybrid loss function: combines the advantages of multiple loss functions to balance vector distance and direction consistency: L similarity =α·L MSE +(1-α)·L cosine α is usually set to 0.5-0.8 and can be tuned using the validation set.

[0172] Step S32: Optimizing the initial student model using the similarity loss function so that the output of the final student model is closer to the speckle feature vector.

[0173] In this step, the similarity loss function is used to iteratively optimize the student model parameters so that its output speckle feature vector gradually approaches the feature expression of the teacher model.

[0174] It is recommended to use the Adam or Adagrad optimizer, with an initial learning rate set to 1e-4 to 1e-3, and a learning rate decay strategy (such as cosine annealing).

[0175] Feature normalization: L2 normalize the output features of the teacher and student models to ensure that the vectors are on the unit hypersphere:

[0176] Temperature scaling: When using KL divergence loss, the output distribution of the teacher model is softened by the temperature parameter T to increase the information entropy of the soft labels:

[0177] Typically, T is set between 2 and 10, with a larger T making the distribution smoother.

[0178] Through a carefully designed similarity loss function and optimization strategy, knowledge distillation in this embodiment can efficiently transfer the speckle feature extraction capability of the teacher model to the lightweight student model, achieving recognition performance close to that of the teacher model while maintaining low computational overhead.

[0179] Figure 5 FIG. 1 is a structural diagram of a system for generating a speckle face recognition model using an infrared face recognition model according to an embodiment of the present invention. Figure 5 As shown, in an embodiment of the present invention, a system for generating a speckle face recognition model by using an infrared face recognition model includes:

[0180] a teacher training module for obtaining an infrared face recognition model, and fine-tuning the infrared face recognition model using a training set to obtain a teacher face model; wherein the training set includes aligned infrared images and speckle images; and the output of the teacher face model includes infrared feature vectors and speckle feature vectors;

[0181] a student establishment module, configured to establish an initial student model and align an output of the initial student model with the speckle feature vector;

[0182] A distillation module is used to distill the output result of the initial student model using the speckle feature vector to obtain a final student model.

[0183] Specifically, the core function of the teacher training module is to build a multimodal teacher model that can process both infrared and speckle images. The specific process is as follows: first, a pre-trained infrared face recognition model is obtained as a basis, and then fine-tuned using a training set containing aligned infrared and speckle images. During the fine-tuning process, a multi-task learning framework (such as simultaneously optimizing the extraction of infrared feature vectors and speckle feature vectors) is used to enable the model to output discriminative features under both modalities. To avoid overfitting, the module freezes some underlying convolutional layers and applies batch normalization and regularization techniques. The resulting teacher model serves as a knowledge source, providing supervisory signals for the subsequent training of the student model.

[0184] The student module is responsible for designing and initializing a lightweight student model, enabling it to efficiently learn the speckle feature extraction capabilities of the teacher model. The student model utilizes lightweight structures such as depthwise separable convolutions. While maintaining low computational overhead, the output dimensions of the final fully connected layer are adjusted to ensure consistency with the teacher model's speckle feature vector dimensions. During the initialization phase, the module uses random initialization or transfer learning to ensure that the student model's initial output closely matches the teacher model's speckle feature distribution, laying the foundation for subsequent knowledge distillation.

[0185] The distillation module is the core of the speckle face recognition model. It uses knowledge distillation to transfer the speckle feature knowledge from the teacher model to the student model. Specifically, the speckle image is input into the teacher and student models, and the similarity loss (such as MSE loss or cosine similarity loss) between their output feature vectors is calculated. Backpropagation is then used to optimize the student model parameters so that its output gradually approaches that of the teacher model. During this process, the distillation module can incorporate auxiliary supervisory information, such as depth images, to further enhance the student model's feature extraction capabilities for speckle images. The resulting student model achieves face recognition performance close to that of the teacher model while maintaining its lightweight.

[0186] The teacher model and speckle feature vector generated by the teacher training module are passed as input to the student setup module and the distillation module. The student model initialized by the student setup module serves as the optimization target of the distillation module. The final student model output by the distillation module can be independently used for speckle face recognition tasks.

[0187] The student setup module relies on the output of the teacher training module (speckle feature vector) to align the output dimensions of the student model. The distillation module relies on the collaborative work of the teacher model and the student model to achieve knowledge transfer by calculating the difference between the two outputs.

[0188] The teacher training module is responsible for encoding knowledge (encoding facial features into speckle feature vectors). The student establishment module is responsible for providing a knowledge decoder (a lightweight student model architecture). The distillation module is responsible for transferring knowledge (transferring knowledge from teacher to student by minimizing feature differences).

[0189] This embodiment utilizes the existing infrared face recognition model. By constructing a multimodal teacher model and distilling its speckle feature knowledge into a lightweight student model, it can quickly obtain a high-quality speckle face recognition model using a highly stable infrared face recognition model when the number of speckle images is small. This has important practical significance and application value.

[0190] An embodiment of the present invention further provides a device for generating a speckle face recognition model using an infrared face recognition model, comprising a processor and a memory storing executable instructions for the processor. The processor is configured to execute the executable instructions to perform the steps of a method for generating a speckle face recognition model using an infrared face recognition model.

[0191] As mentioned above, this embodiment utilizes the existing infrared face recognition model. By constructing a multimodal teacher model and distilling its speckle feature knowledge into a lightweight student model, it can quickly obtain a high-quality speckle face recognition model using a highly stable infrared face recognition model when the number of speckle images is small, which has important practical significance and application value.

[0192] Those skilled in the art will appreciate that various aspects of the present invention may be implemented as systems, methods, or program products. Accordingly, various aspects of the present invention may be implemented in the following forms: entirely in hardware, entirely in software (including firmware, microcode, etc.), or in a combination of hardware and software, collectively referred to herein as "circuits," "modules," or "platforms."

[0193] Figure 6 This is a schematic diagram of the structure of a device for generating a speckle face recognition model using an infrared face recognition model in an embodiment of the present invention. Figure 6 An electronic device 600 according to this embodiment of the present invention will be described. Figure 6 The electronic device 600 shown is merely an example and should not limit the functions and scope of use of the embodiments of the present invention.

[0194] like Figure 6 As shown, electronic device 600 is implemented as a general-purpose computing device. Components of electronic device 600 may include, but are not limited to, at least one processing unit 610, at least one storage unit 620, a bus 630 connecting different platform components (including storage unit 620 and processing unit 610), and a display unit 640.

[0195] The storage unit stores program codes, which can be executed by the processing unit 610, so that the processing unit 610 performs the steps according to various exemplary embodiments of the present invention described in the above-mentioned method of generating a speckle face recognition model using an infrared face recognition model. For example, the processing unit 610 can perform the following steps: Figure 1 Follow the steps shown in .

[0196] The storage unit 620 may include a readable medium in the form of a volatile storage unit, such as a random access memory unit (RAM) 6201 and / or a cache memory unit 6202 , and may further include a read-only memory unit (ROM) 6203 .

[0197] The storage unit 620 may also include a program / utility 6204 having a set (at least one) of program modules 6205, such program modules 6205 including, but not limited to, an operating system, one or more application programs, other program modules, and program data, each of which or some combination may include an implementation of a grid environment.

[0198] Bus 630 may represent one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, a processing unit, or a local bus using any of a variety of bus architectures.

[0199] The electronic device 600 may also communicate with one or more external devices 700 (e.g., keyboards, pointing devices, Bluetooth devices, etc.), one or more devices that enable a user to interact with the electronic device 600, and / or any device that enables the electronic device 600 to communicate with one or more other computing devices (e.g., routers, modems, etc.). Such communication may be performed through an input / output (I / O) interface 650. Furthermore, the electronic device 600 may also communicate with one or more grids (e.g., a local area network (LAN), a wide area network (WAN), and / or a public grid, such as the Internet) through a grid adapter 660. The grid adapter 660 may communicate with other modules of the electronic device 600 through the bus 630. It should be understood that although Figure 6 Not shown, other hardware and / or software modules may be used in conjunction with electronic device 600, including but not limited to microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage platforms.

[0200] An embodiment of the present invention also provides a computer-readable storage medium for storing a program that, when executed, implements the steps of a method for generating a speckle face recognition model using an infrared face recognition model. In some possible implementations, various aspects of the present invention may also be implemented in the form of a program product, which includes program code. When the program product is executed on a terminal device, the program code is configured to cause the terminal device to execute the steps according to various exemplary embodiments of the present invention described in the aforementioned section of this specification regarding a method for generating a speckle face recognition model using an infrared face recognition model.

[0201] As shown above, this embodiment utilizes the existing infrared face recognition model. By constructing a multimodal teacher model and distilling its speckle feature knowledge into a lightweight student model, it can quickly obtain a high-quality speckle face recognition model using a highly stable infrared face recognition model when the number of speckle images is small, which has important practical significance and application value.

[0202] Figure 7 Schematic diagram of the structure of the computer-readable storage medium in an embodiment of the present invention. Figure 7 , a program product 800 for implementing the above method according to an embodiment of the present invention is described. The program product 800 may be a portable compact disc read-only memory (CD-ROM) and include program code, and may be run on a terminal device, such as a personal computer. However, the program product of the present invention is not limited thereto. In this document, a readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.

[0203] The program product may employ any combination of one or more readable media. The readable medium may be a readable signal medium or a readable storage medium. The readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination thereof. More specific examples (a non-exhaustive list) of readable storage media include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof.

[0204] Computer-readable storage media may include a data signal propagated in baseband or as part of a carrier wave, which carries readable program code. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. The readable storage medium may also be any readable medium other than a readable storage medium, which may send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the readable storage medium may be transmitted using any appropriate medium, including but not limited to wireless, wired, optical cable, RF, etc., or any suitable combination thereof.

[0205] The program code for performing the operations of the present invention may be written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Java, C++, and the like, as well as conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the user computing device, partially on the user device, as a stand-alone software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server. In cases involving a remote computing device, the remote computing device may be connected to the user computing device via any type of grid, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., via the Internet using an Internet service provider).

[0206] This embodiment utilizes the existing infrared face recognition model. By constructing a multimodal teacher model and distilling its speckle feature knowledge into a lightweight student model, it can quickly obtain a high-quality speckle face recognition model using a highly stable infrared face recognition model when the number of speckle images is small. This has important practical significance and application value.

[0207] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same and similar parts between the various embodiments can be referred to each other. The above description of the disclosed embodiments enables professionals and technicians in this field to implement or use the present invention. Various modifications to these embodiments will be apparent to professionals and technicians in this field, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but will conform to the widest scope consistent with the principles and novel features disclosed herein.

[0208] The above describes the specific embodiments of the present invention. It should be understood that the present invention is not limited to the above specific embodiments, and those skilled in the art may make various variations or modifications within the scope of the claims, which do not affect the essence of the present invention.

Claims

1. A method for generating a speckle face recognition model using an infrared face recognition model, characterized in that: include: Step S1: obtaining an infrared face recognition model, and fine-tuning the infrared face recognition model using a training set to obtain a teacher face model; wherein the training set includes aligned infrared images and speckle images; and the output of the teacher face model includes infrared feature vectors and speckle feature vectors; Step S2: establishing an initial student model, and aligning the output of the initial student model with the speckle feature vector; Step S3: using the speckle feature vector to distill the output result of the initial student model to obtain a final student model.

2. The method for generating a speckle face recognition model using an infrared face recognition model according to claim 1, characterized in that: In the training set, the infrared image and speckle image aligned in step S1 are obtained by: Step S11: using a multimodal sensor to synchronously capture an infrared image and a speckle image of the same face; Step S12: aligning the spatial coordinates of the speckle image with the infrared image through affine transformation.

3. The method for generating a speckle face recognition model using an infrared face recognition model according to claim 1, characterized in that: The training set also includes a depth image obtained from the speckle image or the infrared image. In step S3, the depth image is also used as auxiliary supervision for the distillation of the initial student model.

4. The method for generating a speckle face recognition model using an infrared face recognition model according to claim 1, characterized in that: In step S1, the fine-tuning process includes: Step S13: freezing part of the bottom convolutional layers of the infrared face recognition model; Step S14: designing a multi-task loss function to simultaneously optimize the extraction of infrared feature vectors and speckle feature vectors; Step S15: Batch normalization and regularization techniques are used to prevent overfitting.

5. The method for generating a speckle face recognition model using an infrared face recognition model according to claim 1, characterized in that: The initial student model is a lightweight network that adopts a depth-separable convolutional structure, and the output dimension of the last fully connected layer remains consistent with the speckle feature vector.

6. The method for generating a speckle face recognition model using an infrared face recognition model according to claim 1, characterized in that: In step S3, the distillation process includes: Step S31: calculating a similarity loss function between the output of the initial student model and the speckle feature vector of the teacher face model; Step S32: Optimizing the initial student model using the similarity loss function so that the output of the final student model is closer to the speckle feature vector.

7. The method for generating a speckle face recognition model using an infrared face recognition model according to claim 6, characterized in that: The similarity loss function includes but is not limited to mean square error loss (MSE), cosine similarity loss or KL divergence loss.

8. A system for generating a speckle face recognition model using an infrared face recognition model, used to implement the method for generating a speckle face recognition model using an infrared face recognition model according to any one of claims 1 to 7, characterized in that: include: a teacher training module for obtaining an infrared face recognition model, and fine-tuning the infrared face recognition model using a training set to obtain a teacher face model; wherein the training set includes aligned infrared images and speckle images; and the output of the teacher face model includes infrared feature vectors and speckle feature vectors; a student establishment module, configured to establish an initial student model and align an output of the initial student model with the speckle feature vector; A distillation module is used to distill the output result of the initial student model using the speckle feature vector to obtain a final student model.

9. A device for generating a speckle face recognition model using an infrared face recognition model, characterized in that: include: processor; a memory storing executable instructions for the processor; The processor is configured to execute the steps of the method for generating a speckle face recognition model using an infrared face recognition model as described in any one of claims 1 to 7 by executing the executable instructions.

10. A computer-readable storage medium for storing a program, characterized in that: When the program is executed, the steps of the method for generating a speckle face recognition model using an infrared face recognition model as described in any one of claims 1 to 7 are implemented.