Method, system and device for generating speckle face recognition model by using infrared face recognition model, and storage medium

By constructing a multimodal teacher model and utilizing knowledge distillation technology, the knowledge of the infrared face recognition model is transferred to a lightweight student model, which solves the problems of difficult and costly training of the speckle face recognition model and enables the rapid generation of high-quality models with a small number of speckle images.

CN120635959APending Publication Date: 2025-09-12SHENZHEN GUANGJIAN TECH CO LTD +1
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510690361.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-27
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

In the existing technology, the training of speckle face recognition models is difficult and costly, and it is difficult to obtain an effective recognition model with a small number of speckle images.

Method used

A multimodal teacher model is constructed using the infrared face recognition model, and the speckle feature knowledge is transferred to the lightweight student model through knowledge distillation technology to generate a high-quality speckle face recognition model.

Benefits of technology

With a smaller number of speckle images, a high-quality speckle face recognition model can be quickly obtained, reducing training cost and time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120635959A_ABST
    Figure CN120635959A_ABST
Patent Text Reader

Abstract

A method, system and device for generating a speckle face recognition model by using an infrared face recognition model, and a storage medium, the method comprising: step S1, obtaining an infrared face recognition model, and performing fine tuning on the infrared face recognition model through a training set to obtain a teacher face model; s2, establishing an initial student model, and aligning the output of the initial student model with the speckle feature vector; and S3, distilling an output result of the initial student model by using the speckle feature vector to obtain a final student model. According to the method, the high-quality speckle face recognition model is quickly obtained by using the high-stability infrared face recognition model, and the method has important practical significance and application value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of face recognition model training, and in particular to a method, system, device and storage medium for generating a speckle face recognition model by using an infrared face recognition model. Background Art

[0002] As one of the core technologies in biometric identification, facial recognition has been widely used in security surveillance, identity authentication, smart terminals, and other fields. Currently, various facial images can be acquired through various techniques, such as infrared images, infrared images, and speckle images. Different images contain varying densities of facial information, and therefore pose varying degrees of difficulty in training corresponding facial recognition models.

[0003] In contrast, speckle images have the lowest information density, making them the most challenging to train for face recognition models, and their success rate is also lower. Training typically requires a very large number of images. However, speckle image data is typically limited, and acquisition requires specialized equipment, making it expensive. Consequently, no existing technology has yet successfully trained an effective speckle face recognition model using a relatively small number of speckle images at a low cost.

[0004] The disclosure of the above background technology content is only used to assist in understanding the inventive concept and technical solution of the present invention. It does not necessarily belong to the prior art of this patent application. In the absence of clear evidence that the above content has been disclosed on the filing date of this patent application, the above background technology should not be used to evaluate the novelty and creativity of this application. Summary of the Invention

[0005] To this end, the present invention proposes a method for generating a speckle face recognition model using an infrared face recognition model. By utilizing the existing infrared face recognition model, a multimodal teacher model is constructed and its speckle feature knowledge is distilled into a lightweight student model. In the case of a small number of speckle images, a high-quality speckle face recognition model can be quickly obtained using a highly stable infrared face recognition model, which has important practical significance and application value.

[0006] In a first aspect, the present invention provides a method for generating a speckle face recognition model using an infrared face recognition model, characterized by comprising:

[0007] Step S1: Obtain an infrared face recognition model, and fine-tune the infrared face recognition model using a training set to obtain a teacher face model; wherein the training set includes an aligned infrared image, a speckle image, and a fused image; the fused image is an image obtained by fusing the aligned infrared image and the aligned speckle image; and the output of the teacher face model includes an infrared feature vector, a speckle feature vector, and a fused feature vector.

[0008] Step S2: establishing an initial student model, and aligning the output of the initial student model with the speckle feature vector;

[0009] Step S3: using the speckle feature vector to distill the output result of the initial student model to obtain a final student model.

[0010] Optionally, the method for generating a speckle face recognition model using an infrared face recognition model is characterized in that the training process of the teacher face model includes: fine-tuning the infrared face recognition model using a deep learning framework, the fine-tuning process using the aligned infrared images, speckle images and fused images in the training set as input data, and optimizing the parameters of the teacher face model through a loss function so that the teacher face model can simultaneously output accurate infrared feature vectors and speckle feature vectors.

[0011] Optionally, the method for generating a speckle face recognition model using an infrared face recognition model is characterized in that, when generating the fused image, the aligned infrared image and the aligned speckle image are fused by weighted summation, and a set of fused images is generated by adjusting the weight coefficient.

[0012] Optionally, the method for generating a speckle face recognition model using an infrared face recognition model is characterized in that the process of establishing the initial student model includes: selecting a pre-trained lightweight convolutional neural network as the initial student model, and adjusting the network structure of the initial student model so that its output matches the dimension of the speckle feature vector.

[0013] Optionally, the method for generating a speckle face recognition model using an infrared face recognition model is characterized in that the step of using the speckle feature vector to distill the output result of the initial student model includes: calculating the similarity between the output of the initial student model and the speckle feature vector, where the similarity is measured by cosine similarity or Euclidean distance, and adjusting the parameters of the initial student model based on the similarity so that the output of the final student model is as close as possible to the speckle feature vector.

[0014] Optionally, the method for generating a speckle face recognition model using an infrared face recognition model is characterized in that the aligned infrared images and speckle images in the training set are aligned using an image registration algorithm, and the image registration algorithm includes a feature point-based registration method or a mutual information-based registration method.

[0015] Optionally, the method for generating a speckle face recognition model using an infrared face recognition model is characterized in that the training process of the final student model also includes a regularization step, and the regularization step includes weight attenuation, Dropout or BatchNormalization to prevent overfitting of the final student model.

[0016] In a second aspect, the present invention provides a system for generating a speckle face recognition model using an infrared face recognition model, which is used to implement any of the aforementioned methods for generating a speckle face recognition model using an infrared face recognition model, and is characterized by comprising:

[0017] a teacher training module, configured to obtain an infrared face recognition model by fine-tuning the infrared face recognition model using a training set to obtain a teacher face model; wherein the training set includes an aligned infrared image, a speckle image, and a fused image; the fused image is a fusion of the aligned infrared image and the aligned speckle image; and the output of the teacher face model includes an infrared feature vector, a speckle feature vector, and a fused feature vector;

[0018] a student establishment module, configured to establish an initial student model and align an output of the initial student model with the speckle feature vector;

[0019] A distillation module is used to distill the output result of the initial student model using the speckle feature vector to obtain a final student model.

[0020] In a third aspect, the present invention provides a device for generating a speckle face recognition model using an infrared face recognition model, characterized in that it includes:

[0021] processor;

[0022] a memory storing executable instructions for the processor;

[0023] The processor is configured to execute the executable instructions to perform any of the steps of the aforementioned method for generating a speckle face recognition model using an infrared face recognition model.

[0024] In a fourth aspect, the present invention provides a computer-readable storage medium for storing a program, characterized in that when the program is executed, the steps of any of the aforementioned methods for generating a speckle face recognition model using an infrared face recognition model are implemented.

[0025] Compared with the prior art, the present invention has the following beneficial effects:

[0026] In this method, the training set contains aligned infrared images, speckle images, and fused images. The teacher face model outputs infrared feature vectors, speckle feature vectors, and fused feature vectors. The infrared face recognition model is fine-tuned using the training set to obtain a teacher face model capable of recognizing speckle images. A distillation technique is then used to obtain a final student model capable of recognizing speckle images.

[0027] The present invention utilizes an existing high-quality infrared image face recognition model, enabling the model to quickly recognize speckle images with a relatively small number of training sets. A high-quality speckle image recognition model, namely the final student model, is then obtained through distillation technology. By utilizing the existing model to quickly obtain a high-quality speckle face recognition model, the requirement for the number of speckle images and the training time required for speckle face recognition model training are greatly reduced. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for the embodiments or the description of the prior art. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without inventive work. Other features, purposes and advantages of the present invention will become more apparent by reading the detailed description of the non-limiting embodiments with reference to the following drawings:

[0029] Figure 1 This is a flowchart of a method for generating a speckle face recognition model using an infrared face recognition model according to an embodiment of the present invention;

[0030] Figure 2 This is a structural diagram of a system for generating a speckle face recognition model using an infrared face recognition model according to an embodiment of the present invention;

[0031] Figure 3 A schematic structural diagram of a device for generating a speckle face recognition model using an infrared face recognition model according to an embodiment of the present invention; and

[0032] Figure 4 Schematic diagram of the structure of a computer-readable storage medium in an embodiment of the present invention. DETAILED DESCRIPTION

[0033] The present invention will be described in detail below with reference to specific embodiments. The following embodiments will help those skilled in the art to further understand the present invention, but are not intended to limit the present invention in any form. It should be noted that, for those skilled in the art, several variations and improvements can be made without departing from the scope of the present invention. These all fall within the scope of protection of the present invention.

[0034] The terms "first," "second," "third," "fourth," and the like (if any) in the description and claims of the present invention and in the accompanying drawings are used to distinguish similar objects and are not necessarily used to describe a particular order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate, so that the embodiments of the invention described herein, for example, can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having," and any variations thereof, are intended to cover non-exclusive inclusions, for example, a process, method, system, product, or apparatus comprising a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to these processes, methods, products, or apparatus.

[0035] The embodiment of the present invention provides a method for generating a speckle face recognition model by using an infrared face recognition model, aiming to solve the problems existing in the prior art.

[0036] The following describes in detail the technical solutions of the present invention and how the technical solutions of this application solve the above-mentioned technical problems using specific embodiments. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described in detail in some embodiments. The following embodiments of the present invention are described in conjunction with the accompanying drawings.

[0037] The present invention utilizes the existing infrared face recognition model, constructs a multimodal teacher model and distills its speckle feature knowledge into a lightweight student model. When the number of speckle images is small, a high-quality speckle face recognition model can be quickly obtained using a highly stable infrared face recognition model, which has important practical significance and application value.

[0038] Figure 1 FIG1 is a flow chart of the steps of a method for generating a speckle face recognition model using an infrared face recognition model according to an embodiment of the present invention. Figure 1 As shown, in an embodiment of the present invention, a method for generating a speckle face recognition model by using an infrared face recognition model includes the following steps:

[0039] Step S1: Obtain an infrared face recognition model, and fine-tune the infrared face recognition model using a training set to obtain a teacher face model; wherein the training set includes aligned infrared images, speckle images, and fused images; the fused image is an image obtained by fusing the aligned infrared images and the aligned speckle images; and the output of the teacher face model includes an infrared feature vector, a speckle feature vector, and a fused feature vector.

[0040] In this step, you first need to obtain a pre-trained infrared face recognition model. This model is usually trained on a large-scale infrared face image dataset and can extract facial features from infrared images.

[0041] Prepare a training set containing aligned infrared images, speckle images, and fused images. These images need to be spatially aligned, that is, their pixel positions correspond to the same part of the same face.

[0042] The fused image is obtained by fusing the aligned infrared image and the aligned speckle image through an algorithm. The purpose of fusion is to combine the information of the two modalities so that the model can learn richer features.

[0043] The infrared face recognition model is fine-tuned using the above training set. The purpose of fine-tuning is to make the model better adaptable to data containing speckle images and fused images, thereby generating a more comprehensive feature vector.

[0044] During fine-tuning, the model learns the associations and differences between infrared images, speckle images, and fused images.

[0045] After fine-tuning, the resulting teacher face model can simultaneously output three feature vectors: infrared feature vector, speckle feature vector, and fusion feature vector. These feature vectors correspond to the input infrared image, speckle image, and fusion image, respectively.

[0046] Step S2: establishing an initial student model, and aligning the output of the initial student model with the speckle feature vector.

[0047] In this step, an initial student model is designed, whose architecture is usually simpler than the teacher model. The goal of the student model is to learn knowledge from the teacher model and ultimately generate outputs that are aligned with the speckle feature vector.

[0048] The output of the initial student model is aligned with the speckle feature vector of the teacher model. This means that the output of the student model should be as close as possible to the speckle feature vector of the teacher model so that the feature representation of the speckle image can be learned.

[0049] Step S3: using the speckle feature vector to distill the output result of the initial student model to obtain a final student model.

[0050] In this step, the output of the initial student model is distilled using the teacher model’s speckle feature vector as a guide. Knowledge distillation is a model compression technique that improves the performance of a smaller student model by having it learn the knowledge of a larger teacher model.

[0051] During the distillation process, the student model will try to imitate the speckle eigenvector of the teacher model and adjust its own parameters by optimizing the loss function so that its output result is as close as possible to the speckle eigenvector of the teacher model.

[0052] After distillation training, the final student model is able to generate outputs that are aligned with the speckle feature vector, which means that the student model has successfully learned the feature representation of the speckle image and can be used for speckle face recognition tasks.

[0053] In some embodiments, the training process of the teacher face model includes: fine-tuning the infrared face recognition model using a deep learning framework, the fine-tuning process using the aligned infrared images, speckle images and fused images in the training set as input data, and optimizing the parameters of the teacher face model through a loss function so that the teacher face model can simultaneously output accurate infrared feature vectors and speckle feature vectors.

[0054] The starting point for the teacher face model is a pre-trained infrared face recognition model. This model has been trained on a large-scale infrared face image dataset and can effectively extract facial features from infrared images. Pre-trained models typically have good generalization and feature extraction capabilities, providing a good foundation for subsequent fine-tuning.

[0055] The input data used in the fine-tuning process is the aligned infrared image, speckle image, and fused image from the training set. The infrared image, speckle image, and fused image are strictly spatially aligned, meaning their pixel locations correspond to the same part of the same face. This alignment can be achieved through an image registration algorithm, ensuring that the model can learn the associations between images from different modalities. The fused image is obtained by fusing the aligned infrared image and speckle image using an algorithm. The purpose of fusion is to combine information from both modalities so that the model can learn richer feature representations. The goal of fine-tuning is to enable the teacher face model to simultaneously output accurate infrared feature vectors and speckle feature vectors. This means that the model needs to learn feature representations for both the infrared image and the speckle image and be able to effectively fuse the features of these two modalities.

[0056] In order to achieve the goal of simultaneously outputting infrared feature vectors and speckle feature vectors, a suitable loss function needs to be designed. The loss function usually includes the following parts:

[0057] Infrared feature loss: used to measure the difference between the infrared feature vector output by the teacher model and the true infrared feature. The loss can be calculated using mean squared error (MSE) or other similarity metrics.

[0058] Speckle feature loss: This is used to measure the difference between the speckle feature vector output by the teacher model and the true speckle feature. The loss can also be calculated using mean squared error or other similarity metrics.

[0059] The final loss function is the weighted sum of the losses of the above parts to balance the learning of different modal features.

[0060] An optimization algorithm (such as gradient descent, Adam optimizer, etc.) is used to update the parameters of the teacher model according to the comprehensive loss function. The goal of the optimization process is to minimize the comprehensive loss function so that the teacher model can simultaneously output accurate infrared feature vectors and speckle feature vectors.

[0061] The fine-tuning process typically requires multiple iterations until the model's performance converges. In each iteration, the model adjusts parameters based on the current loss function value, gradually improving its ability to extract features from infrared and speckle images.

[0062] Through the above fine-tuning process and loss function design, the teacher face model can simultaneously learn the feature representations of infrared and speckle images and generate accurate infrared and speckle feature vectors. This process provides a high-quality teacher model for subsequent knowledge distillation and lays the foundation for generating a speckle face recognition model.

[0063] In some embodiments, when generating the fused image, the aligned infrared image and the aligned speckle image are fused by weighted summation, and a set of fused images is generated by adjusting weight coefficients.

[0064] Before fusion, the infrared image and the speckle image must be aligned. The purpose of alignment is to ensure that the two modal images are spatially matched, that is, their pixel locations correspond to the same part of the same face.

[0065] Weighted summation is a simple yet effective method for fusing aligned infrared and speckle images into a single fused image. First, two weight coefficients are set, representing the weights of the infrared and speckle images in the fusion process. The sum of these weight coefficients is 1 to ensure that the intensity of the fused image remains within a reasonable range. All pixels in an image have the same weight coefficient.

[0066] A set of weight coefficients can produce a fused image. By adjusting the weight coefficients, a set of different fused images can be generated. These images differ in the weighting of infrared and speckle information. This diversity helps the model learn the fusion characteristics of different modal information.

[0067] By fusing the aligned infrared and speckle images through weighted summation and adjusting the weight coefficients to generate a set of fused images, the model can be fed with diverse input data. This approach not only preserves the characteristic information of the two modal images but also helps the model learn their complementarity, thereby improving its performance and robustness.

[0068] In some embodiments, the process of establishing the initial student model includes: selecting a pre-trained lightweight convolutional neural network as the initial student model, and adjusting the network structure of the initial student model so that its output matches the dimension of the speckle feature vector.

[0069] Select a pre-trained lightweight convolutional neural network as the initial student model. Lightweight networks typically have fewer parameters and lower computational complexity, making them suitable for resource-constrained environments, such as mobile devices or embedded systems. Common lightweight convolutional neural networks include MobileNet, ShuffleNet, and SqueezeNet. These networks are designed to balance efficiency and performance, reducing computational and storage requirements while maintaining high accuracy.

[0070] Pretrained lightweight networks have been trained on large-scale image datasets (such as ImageNet) and are capable of extracting common image features. These pretrained models have excellent feature extraction and generalization capabilities, providing a good foundation for subsequent fine-tuning. Using pretrained models can significantly reduce training time and computing resources while improving model performance.

[0071] The network structure of the initial student model is adjusted so that its output matches the dimension of the speckle feature vector of the teacher model. This means that the output feature vector of the student model should have the same dimension as the speckle feature vector for subsequent knowledge distillation.

[0072] Check the output layer structure of the pre-trained lightweight network to ensure that the dimension of its output feature vector is consistent with the dimension of the speckle feature vector. If the output dimension of the pre-trained model does not match the target dimension, you can adjust the dimension by modifying the output layer structure. For example, you can adjust the number of output units in the fully connected layer to match the dimension of the speckle feature vector.

[0073] After adjusting the network structure, you need to verify that the dimensions of the output feature vector of the student model are consistent with the dimensions of the speckle feature vector. This can be done by inputting some sample images and checking the dimensions of the output feature vector. If the dimensions do not match, you need to readjust the network structure until the output dimensions match the target dimensions.

[0074] An initial student model is established by selecting a pre-trained lightweight convolutional neural network and adjusting its network structure. The key to this adjustment is to ensure that the dimensions of the student model's output feature vector match those of the teacher model's speckle feature vector, facilitating subsequent knowledge distillation. This process not only leverages the advantages of the pre-trained model but also adapts the network structure to specific task requirements.

[0075] In some embodiments, the step of distilling the output result of the initial student model using the speckle feature vector includes: calculating the similarity between the output of the initial student model and the speckle feature vector, where the similarity is measured by cosine similarity or Euclidean distance, and adjusting the parameters of the initial student model based on the similarity so that the output of the final student model is as close as possible to the speckle feature vector.

[0076] In knowledge distillation, it is necessary to measure the similarity between the output feature vector of the initial student model and the speckle feature vector of the teacher model. This similarity can be calculated using cosine similarity and Euclidean distance. The appropriate similarity metric should be selected based on the specific task and data characteristics. Cosine similarity focuses more on the direction of the vector and is suitable for situations where the feature vector lengths vary significantly. Euclidean distance focuses more on the absolute distance between vectors and is suitable for situations where the feature vector lengths are relatively consistent.

[0077] Based on the selected similarity metric, a distillation loss function is constructed. The goal of the loss function is to minimize the difference between the output of the student model and the speckle feature vector of the teacher model.

[0078] An optimization algorithm (such as gradient descent or the Adam optimizer) is used to update the parameters of the student model based on the distillation loss function. The goal of the optimization process is to minimize the distillation loss function so that the output feature vector of the student model is as close as possible to the speckle feature vector of the teacher model. In each iteration, the similarity between the output of the current student model and the speckle feature vector of the teacher model is calculated, and the parameters of the student model are adjusted based on the similarity.

[0079] The distillation process typically requires multiple iterations until the similarity between the output of the student model and the speckle feature vector of the teacher model reaches a satisfactory level. In each iteration, the student model adjusts its parameters based on the current loss function value, gradually improving its feature extraction ability.

[0080] After the distillation process is complete, the performance of the final student model is evaluated. Evaluation metrics can include feature vector similarity, face recognition accuracy, etc. Ensure that the output feature vector of the final student model is highly consistent with the speckle feature vector of the teacher model in terms of similarity.

[0081] This embodiment achieves knowledge distillation by calculating the similarity between the output of the initial student model and the speckle feature vector of the teacher model and adjusting the student model parameters based on the similarity. This process not only utilizes the knowledge of the teacher model but also optimizes the student model parameters to make its output as close as possible to the teacher model's speckle feature vector, thereby generating an efficient and high-performance final student model.

[0082] In some embodiments, the aligned infrared images and speckle images in the training set are aligned using an image registration algorithm, wherein the image registration algorithm includes a feature point-based registration method or a mutual information-based registration method.

[0083] In facial image processing, aligning infrared and speckle images is a key step to ensure that the two modal images match spatially. Image registration algorithms are an important tool to achieve this goal.

[0084] The goal of image registration is to spatially align two or more images so that their pixel locations correspond to the same locations in the same scene. In the alignment of infrared and speckle images, this means that the facial features (such as eyes, nose, mouth, etc.) in the two images are completely matched in spatial location.

[0085] Feature point-based registration is a common image registration technique. Its core idea is to achieve alignment by detecting and matching feature points in the image. The following are the specific steps of this method:

[0086] (1) Feature point detection

[0087] -Detection algorithm: Use feature point detection algorithms (such as SIFT, SURF, ORB, etc.) to extract feature points in infrared images and speckle images. These feature points are usually key points in the image, such as corner points, edge points, etc.

[0088] -Feature description: Generate a descriptor for each detected feature point. The descriptor is a vector that describes the local image information around the feature point.

[0089] (2) Feature point matching

[0090] - Matching algorithm: Use a feature point matching algorithm (such as nearest neighbor matching, FLANN matching, etc.) to match the feature points in the infrared image with the feature points in the speckle image. The basis for matching is the similarity between the feature point descriptors.

[0091] -Filter matching points: Use certain strategies (such as Lowe's ratio test) to filter out reliable matching point pairs and remove incorrectly matched point pairs.

[0092] (3) Spatial transformation estimation

[0093] -Transformation model: Estimate the spatial transformation relationship between two images based on matching point pairs. Common transformation models include rigid transformation (translation and rotation), affine transformation (translation, rotation, scaling, and shearing), and homography transformation (applicable to planar scenes).

[0094] - Parameter estimation: Estimate the parameters of the transformation model using least squares or other optimization algorithms such as RANSAC.

[0095] (4) Image alignment

[0096] - Apply transformation: Apply the estimated transformation parameters to the infrared image or speckle image to generate the aligned image through methods such as interpolation.

[0097] -Result Verification: Verify the alignment effect by visualizing the aligned images or calculating the alignment error.

[0098] Mutual information-based registration is a statistical information-based registration technique that is applicable to images with different modalities (such as infrared images and speckle images). Mutual information measures the degree of information sharing between two images. The greater the mutual information, the higher the correlation between the two images. The following are the specific steps of this method:

[0099] Mutual information measures the degree of information sharing between random variables.

[0100] The optimization goal is to maximize mutual information: The goal of mutual information-based registration methods is to maximize the mutual information between two images. The mutual information is maximized by adjusting the spatial transformation parameters between the images.

[0101] Optimization algorithms (such as gradient descent and Powell's method) are used to search for optimal spatial transformation parameters that maximize mutual information. To improve optimization efficiency, a multi-scale optimization strategy is often employed, starting with optimization at a coarse scale and then gradually refining it to finer scales. The optimized transformation parameters are applied to the infrared or speckle image to generate the aligned image. The alignment is verified by calculating mutual information or other similarity metrics.

[0102] When aligning infrared and speckle images, you can choose between feature-based or mutual information-based registration. Feature-based registration is suitable for situations with distinct feature points and offers high computational efficiency, while mutual information-based registration is suitable for images with significant modal differences and requires higher registration accuracy. Choosing the appropriate registration method, depending on the specific application scenario and image characteristics, can effectively improve alignment and provide high-quality input data for subsequent fusion and analysis.

[0103] In some embodiments, the training process of the final student model further includes a regularization step, and the regularization step includes weight decay, Dropout or BatchNormalization to prevent the final student model from overfitting.

[0104] Regularization is a method to improve the generalization ability of a model by limiting its complexity or introducing additional constraints. The following is a detailed description of the regularization steps (including weight decay, dropout, and batch normalization) used in the final student model training process:

[0105] 1. Weight Decay

[0106] Weight decay is a common regularization method that limits the size of model weights by adding a regularization term to the loss function. Weight decay adds a penalty term proportional to the sum of the squared weights to the loss function, which helps the model maintain a low weight value during training. Weight decay prevents model weights from becoming excessively large, thereby reducing model complexity and mitigating the risk of overfitting. The strength of the regularization can be controlled by adjusting the regularization coefficient.

[0107] 2. Dropout

[0108] Dropout is a technique that prevents overfitting by randomly dropping neurons. During training, dropout randomly sets the output of a subset of neurons to zero, so that the model only uses a subset of neurons for training at each iteration. This randomness prevents over-dependence between neurons, thereby improving the model's generalization ability. The dropout probability is a hyperparameter that indicates the proportion of neurons that are dropped.

[0109] Dropout can simulate multiple different sub-networks to enhance the robustness of the model. By randomly dropping neurons, it can prevent the model from overfitting to noise or outliers in the training data.

[0110] 3. Batch Normalization

[0111] BatchNormalization is a technique that stabilizes the training process by normalizing the input to each layer.

[0112] BatchNormalization reduces internal covariate shift (InternalCovariateShift) by normalizing the input of each layer to a distribution with mean 0 and variance 1.

[0113] BatchNormalization can accelerate the training process and enable faster model convergence. It can also reduce sensitivity to initialization parameters and improve model stability. By normalizing the input, it can indirectly provide a certain degree of regularization, reducing the risk of overfitting.

[0114] 4. Comprehensive application of regularization steps

[0115] During the training of the final student model, the above regularization techniques are usually used in combination to achieve the best regularization effect:

[0116] Weight decay: By setting the weight decay parameter in the optimizer, the size of the weight is limited.

[0117] Dropout: Insert the Dropout layer after certain layers of the model (such as the fully connected layer) to randomly discard some neurons.

[0118] BatchNormalization: Insert the BatchNormalization layer after the convolutional layer or the fully connected layer to normalize the input data.

[0119] These regularization techniques can complement each other and work together in the model training process, thereby effectively preventing overfitting and improving the generalization ability of the model.

[0120] This example effectively prevents model overfitting by introducing regularization techniques such as weight decay, dropout, and batch normalization during the training of the final student model. These techniques improve the model's generalization ability by limiting model complexity, randomly dropping neurons, and normalizing input data, ensuring that the model's performance on the training set is better generalized to the test set.

[0121] Figure 2 FIG. 1 is a structural diagram of a system for generating a speckle face recognition model using an infrared face recognition model according to an embodiment of the present invention. Figure 2 As shown, in an embodiment of the present invention, a system for generating a speckle face recognition model using an infrared face recognition model includes:

[0122] a teacher training module, configured to obtain an infrared face recognition model by fine-tuning the infrared face recognition model using a training set to obtain a teacher face model; wherein the training set includes an aligned infrared image, a speckle image, and a fused image; the fused image is a fusion of the aligned infrared image and the aligned speckle image; and the output of the teacher face model includes an infrared feature vector, a speckle feature vector, and a fused feature vector;

[0123] a student establishment module, configured to establish an initial student model and align an output of the initial student model with the speckle feature vector;

[0124] A distillation module is used to distill the output result of the initial student model using the speckle feature vector to obtain a final student model.

[0125] Specifically, the teacher training module obtains a pre-trained infrared face recognition model. The training set includes aligned infrared images, speckle images, and fused images. The pre-trained infrared face recognition model is fine-tuned using the training set. During the fine-tuning process, the model learns the associations between infrared images, speckle images, and fused images. By optimizing the loss function (usually including infrared feature loss, speckle feature loss, and fusion feature loss), the model parameters are adjusted so that the model can simultaneously output accurate infrared feature vectors, speckle feature vectors, and fusion feature vectors. The fine-tuned teacher face model is output, which can output infrared feature vectors, speckle feature vectors, and fusion feature vectors.

[0126] The core function of the teacher training module is to generate a teacher model that can process both infrared and speckle images. This model not only retains the feature extraction capabilities of the infrared face recognition model but also learns the feature representation of speckle images through fine-tuning, providing a high-quality "teacher" model for subsequent knowledge distillation.

[0127] The student setup module selects a pretrained lightweight convolutional neural network as the initial student model. The student model's network structure is adjusted so that the dimensions of its output feature vector match those of the teacher model's speckle feature vector. The student model's parameters are initialized to prepare for the subsequent distillation process. The initial student model's output feature vector is aligned with the teacher model's speckle feature vector.

[0128] The core role of the student setup module is to create a lightweight student model and align its output with the speckle feature vector of the teacher model. This module provides the basis for knowledge distillation, enabling the student model to learn the knowledge of the teacher model.

[0129] The input to the distillation module includes the speckle feature vector of the teacher face model and the output feature vector of the initial student model. The similarity between the output feature vector of the initial student model and the speckle feature vector of the teacher model is calculated (cosine similarity or Euclidean distance can be used). A distillation loss function is constructed based on the similarity, and the parameters of the student model are adjusted by optimizing this loss function. Through multiple iterative training, the output feature vector of the student model is made as close as possible to the speckle feature vector of the teacher model. Ultimately, the output feature vector of the student model is highly similar to the speckle feature vector of the teacher model, making it suitable for speckle face recognition tasks.

[0130] The core function of the distillation module is to transfer the knowledge of the teacher model to the student model through knowledge distillation. By optimizing the distillation loss function, the student model can learn the teacher model's feature representation of speckle images, thereby generating a lightweight and high-performance speckle face recognition model.

[0131] This embodiment utilizes the existing infrared face recognition model. By constructing a multimodal teacher model and distilling its speckle feature knowledge into a lightweight student model, it can quickly obtain a high-quality speckle face recognition model using a highly stable infrared face recognition model when the number of speckle images is small. This has important practical significance and application value.

[0132] An embodiment of the present invention further provides a device for generating a speckle face recognition model using an infrared face recognition model, comprising a processor and a memory storing executable instructions for the processor. The processor is configured to execute the executable instructions to perform the steps of a method for generating a speckle face recognition model using an infrared face recognition model.

[0133] As mentioned above, this embodiment utilizes the existing infrared face recognition model. By constructing a multimodal teacher model and distilling its speckle feature knowledge into a lightweight student model, it can quickly obtain a high-quality speckle face recognition model using a highly stable infrared face recognition model when the number of speckle images is small, which has important practical significance and application value.

[0134] Those skilled in the art will appreciate that various aspects of the present invention may be implemented as systems, methods, or program products. Accordingly, various aspects of the present invention may be implemented in the following forms: entirely in hardware, entirely in software (including firmware, microcode, etc.), or in a combination of hardware and software, collectively referred to herein as "circuits," "modules," or "platforms."

[0135] Figure 3 This is a schematic diagram of the structure of a device for generating a speckle face recognition model using an infrared face recognition model in an embodiment of the present invention. Figure 3 An electronic device 600 according to this embodiment of the present invention will be described. Figure 3 The electronic device 600 shown is merely an example and should not limit the functions and scope of use of the embodiments of the present invention.

[0136] like Figure 3 As shown, electronic device 600 is implemented as a general-purpose computing device. Components of electronic device 600 may include, but are not limited to, at least one processing unit 610, at least one storage unit 620, a bus 630 connecting different platform components (including storage unit 620 and processing unit 610), and a display unit 640.

[0137] The storage unit stores program codes, which can be executed by the processing unit 610, so that the processing unit 610 performs the steps according to various exemplary embodiments of the present invention described in the above-mentioned method of generating a speckle face recognition model using an infrared face recognition model. For example, the processing unit 610 can perform the following steps: Figure 1 Follow the steps shown in .

[0138] The storage unit 620 may include a readable medium in the form of a volatile storage unit, such as a random access memory unit (RAM) 6201 and / or a cache memory unit 6202 , and may further include a read-only memory unit (ROM) 6203 .

[0139] The storage unit 620 may also include a program / utility 6204 having a set (at least one) of program modules 6205, such program modules 6205 including, but not limited to, an operating system, one or more application programs, other program modules, and program data, each of which or some combination may include an implementation of a grid environment.

[0140] Bus 630 may represent one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, a processing unit, or a local bus using any of a variety of bus architectures.

[0141] The electronic device 600 may also communicate with one or more external devices 700 (e.g., keyboards, pointing devices, Bluetooth devices, etc.), one or more devices that enable a user to interact with the electronic device 600, and / or any device that enables the electronic device 600 to communicate with one or more other computing devices (e.g., routers, modems, etc.). Such communication may be performed through an input / output (I / O) interface 650. Furthermore, the electronic device 600 may also communicate with one or more grids (e.g., a local area network (LAN), a wide area network (WAN), and / or a public grid, such as the Internet) through a grid adapter 660. The grid adapter 660 may communicate with other modules of the electronic device 600 through the bus 630. It should be understood that although Figure 3 Not shown, other hardware and / or software modules may be used in conjunction with electronic device 600, including but not limited to microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage platforms.

[0142] An embodiment of the present invention also provides a computer-readable storage medium for storing a program that, when executed, implements the steps of a method for generating a speckle face recognition model using an infrared face recognition model. In some possible implementations, various aspects of the present invention may also be implemented in the form of a program product, which includes program code. When the program product is executed on a terminal device, the program code is configured to cause the terminal device to execute the steps according to various exemplary embodiments of the present invention described in the aforementioned section of this specification regarding a method for generating a speckle face recognition model using an infrared face recognition model.

[0143] As shown above, this embodiment utilizes the existing infrared face recognition model. By constructing a multimodal teacher model and distilling its speckle feature knowledge into a lightweight student model, it can quickly obtain a high-quality speckle face recognition model using a highly stable infrared face recognition model when the number of speckle images is small, which has important practical significance and application value.

[0144] Figure 4 Schematic diagram of the structure of the computer-readable storage medium in an embodiment of the present invention. Figure 4 , a program product 800 for implementing the above method according to an embodiment of the present invention is described. The program product 800 may be a portable compact disc read-only memory (CD-ROM) and include program code, and may be run on a terminal device, such as a personal computer. However, the program product of the present invention is not limited thereto. In this document, a readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.

[0145] The program product may employ any combination of one or more readable media. The readable medium may be a readable signal medium or a readable storage medium. The readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination thereof. More specific examples (a non-exhaustive list) of readable storage media include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof.

[0146] Computer-readable storage media may include a data signal propagated in baseband or as part of a carrier wave, which carries readable program code. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. The readable storage medium may also be any readable medium other than a readable storage medium, which may send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the readable storage medium may be transmitted using any appropriate medium, including but not limited to wireless, wired, optical cable, RF, etc., or any suitable combination thereof.

[0147] The program code for performing the operations of the present invention may be written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Java, C++, and the like, as well as conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the user computing device, partially on the user device, as a stand-alone software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server. In cases involving a remote computing device, the remote computing device may be connected to the user computing device via any type of grid, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., via the Internet using an Internet service provider).

[0148] This embodiment utilizes the existing infrared face recognition model. By constructing a multimodal teacher model and distilling its speckle feature knowledge into a lightweight student model, it can quickly obtain a high-quality speckle face recognition model using a highly stable infrared face recognition model when the number of speckle images is small. This has important practical significance and application value.

[0149] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same and similar parts between the various embodiments can be referred to each other. The above description of the disclosed embodiments enables professionals and technicians in this field to implement or use the present invention. Various modifications to these embodiments will be apparent to professionals and technicians in this field, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but will conform to the widest scope consistent with the principles and novel features disclosed herein.

[0150] The above describes the specific embodiments of the present invention. It should be understood that the present invention is not limited to the above specific embodiments, and those skilled in the art may make various variations or modifications within the scope of the claims, which do not affect the essence of the present invention.

Claims

1. A method for generating a speckle face recognition model using an infrared face recognition model, characterized in that: include: Step S1: Obtain an infrared face recognition model, and fine-tune the infrared face recognition model using a training set to obtain a teacher face model; wherein the training set includes an aligned infrared image, a speckle image, and a fused image; the fused image is an image obtained by fusing the aligned infrared image and the aligned speckle image; and the output of the teacher face model includes an infrared feature vector, a speckle feature vector, and a fused feature vector. Step S2: establishing an initial student model, and aligning the output of the initial student model with the speckle feature vector; Step S3: using the speckle feature vector to distill the output result of the initial student model to obtain a final student model.

2. The method for generating a speckle face recognition model using an infrared face recognition model according to claim 1, characterized in that: The training process of the teacher face model includes: fine-tuning the infrared face recognition model using a deep learning framework, the fine-tuning process uses the aligned infrared images, speckle images and fused images in the training set as input data, and optimizes the parameters of the teacher face model through a loss function so that the teacher face model can simultaneously output accurate infrared feature vectors and speckle feature vectors.

3. The method for generating a speckle face recognition model using an infrared face recognition model according to claim 1, characterized in that: When generating the fused image, the aligned infrared image and the aligned speckle image are fused by weighted summation, and a set of fused images is generated by adjusting weight coefficients.

4. The method for generating a speckle face recognition model using an infrared face recognition model according to claim 1, characterized in that: The process of establishing the initial student model includes: selecting a pre-trained lightweight convolutional neural network as the initial student model, and adjusting the network structure of the initial student model so that its output matches the dimension of the speckle feature vector.

5. The method for generating a speckle face recognition model using an infrared face recognition model according to claim 1, characterized in that: The step of distilling the output result of the initial student model using the speckle feature vector includes: calculating the similarity between the output of the initial student model and the speckle feature vector, where the similarity is measured by cosine similarity or Euclidean distance, and adjusting the parameters of the initial student model according to the similarity so that the output of the final student model is as close as possible to the speckle feature vector.

6. The method for generating a speckle face recognition model using an infrared face recognition model according to claim 1, characterized in that: The aligned infrared images and speckle images in the training set are aligned using an image registration algorithm, wherein the image registration algorithm includes a feature point-based registration method or a mutual information-based registration method.

7. The method for generating a speckle face recognition model using an infrared face recognition model according to claim 1, characterized in that: The training process of the final student model also includes a regularization step, which includes weight decay, Dropout or BatchNormalization to prevent the final student model from overfitting.

8. A system for generating a speckle face recognition model using an infrared face recognition model, used to implement the method for generating a speckle face recognition model using an infrared face recognition model according to any one of claims 1 to 7, characterized in that: include: a teacher training module, configured to obtain an infrared face recognition model by fine-tuning the infrared face recognition model using a training set to obtain a teacher face model; wherein the training set includes an aligned infrared image, a speckle image, and a fused image; the fused image is a fusion of the aligned infrared image and the aligned speckle image; and the output of the teacher face model includes an infrared feature vector, a speckle feature vector, and a fused feature vector; a student establishment module, configured to establish an initial student model and align an output of the initial student model with the speckle feature vector; A distillation module is used to distill the output result of the initial student model using the speckle feature vector to obtain a final student model.

9. A device for generating a speckle face recognition model using an infrared face recognition model, characterized in that: include: processor; a memory storing executable instructions for the processor; The processor is configured to execute the steps of the method for generating a speckle face recognition model using an infrared face recognition model as described in any one of claims 1 to 7 by executing the executable instructions.

10. A computer-readable storage medium for storing a program, characterized in that: When the program is executed, the steps of the method for generating a speckle face recognition model using an infrared face recognition model as described in any one of claims 1 to 7 are implemented.

Citation Information

Cited By

  • Training method of medical ultrasound large model

    CN121212270A

  • A training method for large medical ultrasound models

    CN121212270B