Face speckle recognition model training method, system and device and storage medium

By distilling the speckle feature knowledge in the RGBD face recognition model into a lightweight student model, the problem of difficulty in training an effective speckle face recognition model under low cost and a small number of speckle images is solved, and efficient model training and recognition performance are achieved.

CN120635958APending Publication Date: 2025-09-12SHENZHEN GUANGJIAN TECH CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510690228.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-27
Publication Date
2025-09-12

Smart Images

  • Figure CN120635958A_ABST
    Figure CN120635958A_ABST
Patent Text Reader

Abstract

The invention discloses a face speckle recognition model training method, system and device and a storage medium, and the method comprises the steps: S1, obtaining an RGBD face recognition model, and carrying out the fine adjustment of the RGBD face recognition model through a training set, and obtaining a teacher face model; wherein the training set comprises an RGB image, a depth image and a speckle image which are aligned; the output of the teacher face model comprises an RGB feature vector, a depth feature vector and a speckle feature vector; s2, establishing an initial student model, and aligning the output of the initial student model with the speckle feature vector; and S3, distilling an output result of the initial student model by using the speckle feature vector to obtain a final student model. According to the method, the high-quality speckle face recognition model is quickly obtained by using the high-stability RGBD face recognition model, and the method has important practical significance and application value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of face recognition model training, and in particular to a face speckle recognition model training method, system, device and storage medium. Background Art

[0002] As one of the core technologies in biometrics, facial recognition has been widely used in security surveillance, identity authentication, smart terminals, and other fields. Currently, a variety of facial recognition models have been developed with excellent performance. However, because speckle images are discontinuous and carry relatively little information, using them for facial recognition requires a vast amount of data.

[0003] However, the acquisition of speckle images requires a structured light camera, which is expensive. When there are only a few speckle images, it is difficult to train a speckle face recognition model that meets the application requirements. Existing technologies do not yet have an effective speckle face recognition model that can be trained with a small number of speckle images and at a low cost.

[0004] The disclosure of the above background technology content is only used to assist in understanding the inventive concept and technical solution of the present invention. It does not necessarily belong to the prior art of this patent application. In the absence of clear evidence that the above content has been disclosed on the filing date of this patent application, the above background technology should not be used to evaluate the novelty and creativity of this application. Summary of the Invention

[0005] To this end, the facial speckle recognition model training method proposed in the present invention utilizes the existing RGBD face recognition model. By constructing a multimodal teacher model and distilling its speckle feature knowledge into a lightweight student model, it can quickly obtain a high-quality speckle face recognition model using a highly stable RGBD face recognition model when the number of speckle images is small, which has important practical significance and application value.

[0006] In a first aspect, the present invention provides a facial speckle recognition model training method, characterized by comprising:

[0007] Step S1: obtaining an RGBD face recognition model, and fine-tuning the RGBD face recognition model using a training set to obtain a teacher face model; wherein the training set includes aligned RGB images, depth images, and speckle images; and the output of the teacher face model includes an RGB feature vector, a depth feature vector, and a speckle feature vector;

[0008] Step S2: establishing an initial student model, and aligning the output of the initial student model with the speckle feature vector;

[0009] Step S3: using the speckle feature vector to distill the output result of the initial student model to obtain a final student model.

[0010] Optionally, the facial speckle recognition model training method is characterized in that the RGBD face recognition model is constructed based on a deep learning framework, and the deep learning framework is selected from TensorFlow, PyTorch or Caffe.

[0011] Optionally, the facial speckle recognition model training method is characterized in that the training set also includes an infrared image aligned with the RGB image, and in step S3, the infrared image is also used as auxiliary supervision for the distillation of the initial student model.

[0012] Optionally, the facial speckle recognition model training method is characterized in that the training set is obtained by preprocessing the collected RGB images, depth images and speckle images, and the preprocessing includes image alignment, normalization and data enhancement.

[0013] Optionally, the facial speckle recognition model training method is characterized in that the fine-tuning process of the teacher face model includes adjusting the learning rate, optimizer and loss function, and the loss function includes cross entropy loss and feature consistency loss.

[0014] Optionally, the facial speckle recognition model training method is characterized in that, when performing the feature alignment operation in step S2, a learnable projection matrix is ​​added after the initial student model output layer; and the Euclidean distance between the student model output and the speckle feature vector is minimized using an L2 regularization loss.

[0015] Optionally, the facial speckle recognition model training method is characterized in that the distillation process in step S3 adopts a temperature-regulated KL divergence loss function; a dynamically adjusted distillation weight coefficient, and the distillation weight coefficient increases linearly from 0.1 to 0.5 with the training rounds.

[0016] In a second aspect, the present invention provides a facial speckle recognition model training system for implementing any of the aforementioned facial speckle recognition model training methods, characterized by comprising:

[0017] a teacher training module for obtaining an RGBD face recognition model, and fine-tuning the RGBD face recognition model using a training set to obtain a teacher face model; wherein the training set includes aligned RGB images, depth images, and speckle images; and the output of the teacher face model includes an RGB feature vector, a depth feature vector, and a speckle feature vector;

[0018] a student establishment module, configured to establish an initial student model and align an output of the initial student model with the speckle feature vector;

[0019] A distillation module is used to distill the output result of the initial student model using the speckle feature vector to obtain a final student model.

[0020] In a third aspect, the present invention provides a facial speckle recognition model training device, characterized by comprising:

[0021] processor;

[0022] a memory storing executable instructions for the processor;

[0023] The processor is configured to execute the steps of any of the aforementioned facial speckle recognition model training methods by executing the executable instructions.

[0024] In a fourth aspect, the present invention provides a computer-readable storage medium for storing a program, characterized in that when the program is executed, the steps of any of the aforementioned facial speckle recognition model training methods are implemented.

[0025] Compared with the prior art, the present invention has the following beneficial effects:

[0026] In this paper, the training set contains aligned RGB images, depth images, and speckle images. The teacher face model outputs RGB, depth, and speckle feature vectors. The RGBD face recognition model is fine-tuned using the training set to obtain a teacher face model capable of recognizing speckle images. Distillation techniques are then used to obtain the final student model capable of recognizing speckle images.

[0027] The present invention utilizes an existing high-quality RGBD image face recognition model, enabling the model to quickly recognize speckle images with a relatively small number of training sets. A high-quality speckle image recognition model, namely the final student model, is then obtained through distillation technology. Using the existing model, a high-quality speckle face recognition model is quickly obtained, significantly reducing the requirement for the number of speckle images and the training time required for speckle face recognition model training. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for the embodiments or the description of the prior art. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without inventive work. Other features, purposes and advantages of the present invention will become more apparent by reading the detailed description of the non-limiting embodiments with reference to the following drawings:

[0029] Figure 1 This is a flowchart of the steps of a facial speckle recognition model training method according to an embodiment of the present invention;

[0030] Figure 2 Schematic diagram of the structure of an RGBD camera according to an embodiment of the present invention;

[0031] Figure 3 Schematic diagram of the structure of a facial speckle recognition model training system according to an embodiment of the present invention;

[0032] Figure 4 is a structural diagram of a facial speckle recognition model training device according to an embodiment of the present invention; and

[0033] Figure 5 Schematic diagram of the structure of a computer-readable storage medium in an embodiment of the present invention.

[0034] 1-Structured light projector;

[0035] 2-Visible light projector;

[0036] 3- speckle receiver;

[0037] 4-RGB receiver; DETAILED DESCRIPTION

[0038] The present invention will be described in detail below with reference to specific embodiments. The following embodiments will help those skilled in the art to further understand the present invention, but are not intended to limit the present invention in any form. It should be noted that, for those skilled in the art, several variations and improvements can be made without departing from the scope of the present invention. These all fall within the scope of protection of the present invention.

[0039] The terms "first," "second," "third," "fourth," and the like (if any) in the description and claims of the present invention and in the accompanying drawings are used to distinguish similar objects and are not necessarily used to describe a particular order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate, so that the embodiments of the invention described herein, for example, can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having," and any variations thereof, are intended to cover non-exclusive inclusions, for example, a process, method, system, product, or apparatus comprising a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to these processes, methods, products, or apparatus.

[0040] The embodiment of the present invention provides a facial speckle recognition model training method, which aims to solve the problems existing in the prior art.

[0041] The following describes in detail the technical solutions of the present invention and how the technical solutions of this application solve the above-mentioned technical problems using specific embodiments. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described in detail in some embodiments. The following embodiments of the present invention are described in conjunction with the accompanying drawings.

[0042] The present invention utilizes the existing RGBD face recognition model, constructs a multimodal teacher model and distills its speckle feature knowledge into a lightweight student model. When the number of speckle images is small, the highly stable RGBD face recognition model can be used to quickly obtain a high-quality speckle face recognition model, which has important practical significance and application value.

[0043] Figure 1 FIG. 1 is a flowchart of a facial speckle recognition model training method according to an embodiment of the present invention. Figure 1 As shown, the steps of a facial speckle recognition model training method in an embodiment of the present invention include:

[0044] Step S1: Obtain an RGBD face recognition model, and fine-tune the RGBD face recognition model using a training set to obtain a teacher face model; wherein the training set includes aligned RGB images, depth images, and speckle images; the output of the teacher face model includes an RGB feature vector, a depth feature vector, and a speckle feature vector.

[0045] In this step, a training set containing aligned RGB images, depth images, and speckle images is used. These images capture facial information from different angles. RGB images provide color information, depth images provide 3D structural information, and speckle images can be used to capture texture or other detailed features of the face.

[0046] The initial model is a multimodal face recognition model that can process both RGB and depth images. This model is usually pre-trained and has certain feature extraction capabilities.

[0047] During fine-tuning, further training (fine-tuning) on ​​a specific training set allows the model to better adapt to the characteristics of speckle images and simultaneously optimize its processing capabilities for RGB and depth images. The RGB, depth, and speckle images from the training set are input into the model, and the model's weights and parameters are adjusted so that the model can simultaneously learn the feature representations of all three image modalities. The fine-tuned teacher face model can output three feature vectors: RGB feature vectors, depth feature vectors, and speckle feature vectors. These feature vectors represent the feature representations of the face in different modalities, providing a foundation for subsequent model distillation.

[0048] The teacher model is a complex and high-performance model whose purpose is to learn rich feature representations, especially those of speckle images, through fine-tuning. These features are then passed on to the student model as "knowledge," helping it to learn more effective feature extraction capabilities.

[0049] Step S2: establishing an initial student model, and aligning the output of the initial student model with the speckle feature vector.

[0050] In this step, the student model is typically a relatively simple and computationally efficient model. Its purpose is to reduce model complexity and computational cost while maintaining high recognition performance. The input to the student model is the speckle image.

[0051] The output of the student model is aligned with the speckle feature vector of the teacher model. This means that the student model needs to learn a speckle feature representation similar to that of the teacher model. Alignment can be achieved through a loss function, such as using the mean squared error (MSE) or other similarity metrics to measure the difference between the student model output and the teacher model speckle feature vector. Backpropagation is then used to optimize the parameters of the student model to make the two as close as possible.

[0052] The alignment process is one of the core steps of knowledge distillation, which enables the student model to learn key feature representations from the teacher model. In this way, the student model can remain simple in structure but approach the teacher model in performance.

[0053] Step S3: using the speckle feature vector to distill the output result of the initial student model to obtain a final student model.

[0054] In this step, the student model is distilled using the speckle feature vector of the teacher model to further optimize the performance of the student model. The purpose of distillation is to effectively transfer the knowledge of the teacher model (i.e., the speckle feature vector) to the student model.

[0055] During the training process, the speckle feature vector of the teacher model is used as a "soft target" to compare with the output of the student model. By optimizing the loss function (such as the distillation loss function), the output of the student model is made as close as possible to the speckle feature vector of the teacher model.

[0056] The distillation loss function usually consists of two parts: one is the classification loss of the student model itself (such as cross entropy loss), and the other is the distillation loss between the student model output and the teacher model speckle feature vector (such as KL divergence or MSE).

[0057] During the distillation process, the parameters of the student model are continuously updated to better fit the speckle feature vector of the teacher model. This process may require multiple iterations until the performance of the student model reaches a satisfactory level.

[0058] The final student model is lightweight, has performance close to that of the teacher model, and has generalization capabilities.

[0059] In some embodiments, the RGBD face recognition model is built based on a deep learning framework selected from TensorFlow, PyTorch, or Caffe. The RGBD face recognition model is built based on a deep learning framework selected from TensorFlow, PyTorch, or Caffe. These three frameworks are currently widely used deep learning tools, each with its own characteristics and advantages. TensorFlow is known for its powerful computational graphs and distributed computing capabilities, making it suitable for training and deploying large-scale models; PyTorch is known for its dynamic computational graphs and flexible programming interfaces, making it easy for researchers to quickly implement and debug models; and Caffe excels in computer vision, especially image processing, with efficient memory management and fast forward propagation capabilities. Choosing any of these frameworks can provide strong technical support for the construction of the RGBD face recognition model, ensuring that the model can efficiently process multiple modal data such as RGB images, depth images, and speckle images, thereby achieving high-quality face recognition capabilities.

[0060] In some embodiments, the training set also includes infrared images aligned with the RGB images. In step S3, the infrared images are also used as auxiliary supervision for the distillation of the initial student model.

[0061] In the facial speckle recognition model training method, the training set not only includes aligned RGB images, depth images, and speckle images, but also includes infrared images aligned with the RGB images. Infrared images provide unique facial features in the infrared spectrum, such as the thermal radiation properties of the skin or its reflection characteristics for specific wavelengths of light. This information can serve as supplementary features to improve the model's robustness and recognition performance.

[0062] In step S3, in addition to using the teacher model's speckle feature vector to distill the student model, infrared images are also introduced as auxiliary supervisory signals. Specifically, the feature representation of the infrared image or related supervisory information (such as feature vectors or classification labels extracted from the infrared image) is incorporated into the distillation process. In this way, the student model not only learns the teacher model's speckle feature representation, but also obtains additional feature information from the infrared image, further optimizing its feature extraction and recognition performance.

[0063] This distillation strategy of multimodal data fusion enables the student model to comprehensively utilize information from multiple modalities such as RGB, depth, speckle and infrared, enhancing its recognition and generalization capabilities in complex environments while maintaining the model's lightweight and high efficiency.

[0064] In some embodiments, the training set is obtained by preprocessing the collected RGB images, depth images, and speckle images, and the preprocessing includes image alignment, normalization, and data enhancement.

[0065] The training set is obtained by preprocessing the collected RGB images, depth images and speckle images. The preprocessing process mainly includes image alignment, normalization and data enhancement. The following is a detailed description:

[0066] Image alignment is one of the key steps in preprocessing. Since RGB images, depth images, and speckle images are usually acquired through different sensors or at different time points, there may be spatial inconsistencies between them. The purpose of image alignment is to spatially align these images so that they correspond to the same facial region. Alignment methods generally include:

[0067] - Feature point detection and matching: By detecting key feature points in the image (such as the outline of the face, eyes, nose, etc.), and using these feature points to spatially align images of different modalities.

[0068] -Geometric transformation: Based on the matching results of feature points, geometric transformations such as translation, rotation or scaling are performed on the images to align them in space.

[0069] -Depth information-assisted alignment: Utilize the 3D information provided by the depth image to assist in aligning the RGB image and the speckle image, ensuring their consistency in 3D space.

[0070] Normalization is to adjust image data to a uniform range or distribution to reduce the differences between different images and improve the training effect of the model. Normalization methods include:

[0071] -Pixel value normalization: Normalize the pixel values ​​of an image from the original range (such as 0-255) to [0,1] or [-1,1] so that the model can handle it better.

[0072] -Histogram Normalization: Adjust the histogram distribution of the image to make the brightness and contrast of different images more consistent.

[0073] - Feature normalization: Normalize the extracted feature vectors, such as using L2 norm normalization, to make the length of the feature vectors consistent, which facilitates subsequent feature comparison and learning.

[0074] Data augmentation is to increase the generalization ability of the model by generating more diverse training samples. Data augmentation methods include:

[0075] -Geometric transformation: Randomly rotate, translate, scale, or flip an image to generate more varied image samples.

[0076] - Color Adjustment: Change the brightness, contrast, saturation, or hue of an image to simulate different lighting conditions.

[0077] -Noise addition: Add random noise, such as Gaussian noise, to the image to enhance the model's robustness to noise.

[0078] -Multimodal fusion enhancement: Combining the characteristics of RGB, depth, and speckle images, we design specific data augmentation strategies, such as performing the same geometric transformation on multimodal images simultaneously to maintain the alignment between them.

[0079] Preferably, the RGB image, the depth image and the speckle image are generated by the same RGBD camera. Figure 2 The RGBD camera shown includes a structured light projector 1, a visible light projector 2, a speckle receiver 3, and an RGB receiver 4. The RGBD camera can generate pixel-level aligned RGB images, infrared images, and speckle images, which can make the results of model training more accurate.

[0080] In some embodiments, the fine-tuning process of the teacher face model includes adjusting the learning rate, optimizer and loss function, and the loss function includes cross entropy loss and feature consistency loss.

[0081] In the facial speckle recognition model training method, the fine-tuning process of the teacher face model is achieved by fine-tuning the learning rate, optimizer, and loss function. These adjustments ensure that the model can better adapt to the RGB images, depth images, and speckle images in the training set and learn effective feature representations.

[0082] The learning rate is a key hyperparameter in deep learning, which determines the step size of the model parameter update during training. During the fine-tuning of the teacher face model, the adjustment of the learning rate is crucial:

[0083] - Initial learning rate selection: Usually a small initial learning rate (such as 1e-4 or 1e-5) is chosen to avoid excessive adjustments to the weights of the pre-trained model during the fine-tuning phase, thereby retaining the useful information learned during the pre-training phase.

[0084] -Learning rate scheduling: Use learning rate scheduling strategies such as Step Decay, Cosine Annealing, or adaptive learning rate adjustment (such as the adaptive learning rate in the Adam optimizer). These strategies can dynamically adjust the learning rate during training, allowing the model to converge quickly in the early stages and fine-tune in the later stages to avoid overfitting.

[0085] The optimizer is responsible for updating the model's parameters based on the gradient of the loss function. During the fine-tuning of the teacher face model, the choice of optimizer has a significant impact on the model's convergence speed and final performance:

[0086] Commonly used optimizers: Commonly used optimizers include SGD (stochastic gradient descent), Adam, and RMSprop. SGD is suitable for large datasets and provides stable convergence. The Adam optimizer combines the advantages of momentum and adaptive learning rates, making it suitable for complex models and multimodal data. RMSprop excels at handling non-stationary targets.

[0087] -Optimizer parameter adjustment: During the fine-tuning stage, the optimizer parameters (such as momentum coefficient, weight decay, etc.) also need to be adjusted according to the specific situation to achieve the best training effect.

[0088] The loss function is a function that measures the difference between the model prediction and the true label. Its design directly affects the learning objectives and performance of the model. During the fine-tuning process of the teacher face model, the loss function includes cross entropy loss and feature consistency loss:

[0089] Cross-entropy loss: A commonly used loss function in classification tasks, cross-entropy loss measures the difference between the probability distribution of the model output and the probability distribution of the true label. In the face speckle recognition task, cross-entropy loss ensures that the model can correctly classify different facial identities.

[0090] Feature consistency loss: To ensure that the model learns consistent feature representations, especially for multimodal inputs (RGB, depth, and speckle images), a feature consistency loss is introduced. This loss function compares the similarity between feature vectors extracted from images of different modalities, ensuring that the model learns robust feature representations. For example, the mean squared error (MSE) or cosine similarity can be used to measure the consistency between feature vectors.

[0091] By adjusting the learning rate, selecting a suitable optimizer, and designing a composite loss function that includes cross-entropy loss and feature consistency loss, the teacher face model can better adapt to the multimodal data in the training set during the fine-tuning phase, learn robust and effective feature representations, and provide a high-quality knowledge foundation for subsequent student model distillation.

[0092] In some embodiments, when performing the feature alignment operation in step S2, a learnable projection matrix is ​​added after the initial student model output layer; and the Euclidean distance between the student model output and the speckle feature vector is minimized by L2 regularization loss.

[0093] In step S2 of the facial speckle recognition model training method, in order to achieve feature alignment between the output of the initial student model and the speckle feature vector of the teacher model, an alignment mechanism based on a learnable projection matrix and L2 regularization loss is specially designed.

[0094] A learnable projection matrix is ​​added after the output layer of the initial student model. Its function is to map the output feature vector of the student model to a feature space that is identical or similar to the speckle feature vector of the teacher model. The projection matrix is ​​a trainable parameter matrix. Through matrix multiplication, the output feature vector of the student model is mapped to the new feature space to obtain the projected feature vector. The projection matrix is ​​updated through backpropagation during training, with the goal of making the projected feature vector as close as possible to the speckle feature vector of the teacher model.

[0095] In order to minimize the difference between the output of the student model and the speckle feature vector of the teacher model, the L2 regularization loss (also known as the mean square error loss) is used. The L2 regularization loss is used to measure the Euclidean distance between two feature vectors. For the projection feature vector of the student model and the speckle feature vector of the teacher model. By minimizing the above L2 regularization loss, the output feature vector of the student model can be as close as possible to the speckle feature vector of the teacher model after projection. This alignment mechanism ensures that the student model can learn the key feature representations of the teacher model. During the training process, the gradient of the loss function is calculated by backpropagation, and the parameters and projection matrix of the student model are updated. The L2 regularization loss not only aligns the feature vectors, but also prevents the model from overfitting through regularization.

[0096] This embodiment achieves feature alignment by adding a learnable projection matrix after the initial student model output layer and using L2 regularization loss to minimize the Euclidean distance between the student model output and the speckle feature vector. This mechanism not only enables the student model to learn the feature representation of the teacher model, but also improves the performance and generalization ability of the student model through the flexibility of the projection matrix and the constraints of the L2 regularization loss.

[0097] In some embodiments, the distillation process in step S3 adopts a temperature-regulated KL divergence loss function; a dynamically adjusted distillation weight coefficient, and the distillation weight coefficient increases linearly from 0.1 to 0.5 with the training rounds.

[0098] In knowledge distillation, the KL divergence (Kullback-Leibler Divergence) loss function is a commonly used method to measure the difference between two probability distributions. Temperature scaling can be used to control the "softness" of the probability distribution, thereby better transferring the knowledge of the teacher model to the student model. Specifically:

[0099] The temperature parameter is used to adjust the "softness" of the softmax distribution output by the teacher model and the student model. A higher temperature will make the softmax distribution smoother and contain more information, while a lower temperature will make the distribution closer to the hard classification result.

[0100] During the distillation process, the distillation weight coefficient is used to balance the distillation loss and the student model's own classification loss. The dynamically adjusted distillation weight coefficient can be adjusted according to the training progress to better guide the student model's learning process. The distillation weight coefficient starts at 0.1 and increases linearly to 0.5. A low initial value allows the student model to rely more on its own classification loss for learning in the early stages of training, avoiding premature influence from the teacher model. As training progresses, the distillation weight coefficient gradually increases, allowing the student model to gradually learn more from the teacher model. The distillation weight coefficient increases linearly with each training round. This linear adjustment strategy can smoothly guide the student model's transition from relying on its own learning to relying on knowledge transfer from the teacher model.

[0101] The final loss function is the weighted sum of the student model’s own classification loss (such as cross entropy loss) and the distillation loss.

[0102] This embodiment uses a temperature-adjusted KL divergence loss function to better transfer knowledge from the teacher model to the student model. Dynamically adjusting the distillation weight coefficients also balances the student model's own learning with knowledge transfer from the teacher model. This strategy not only improves the performance of the student model but also ensures the stability and effectiveness of the training process.

[0103] Figure 3 FIG. 1 is a structural diagram of a face speckle recognition model training system according to an embodiment of the present invention. Figure 3 As shown, a facial speckle recognition model training system in an embodiment of the present invention includes:

[0104] a teacher training module for obtaining an RGBD face recognition model, and fine-tuning the RGBD face recognition model using a training set to obtain a teacher face model; wherein the training set includes aligned RGB images, depth images, and speckle images; and the output of the teacher face model includes an RGB feature vector, a depth feature vector, and a speckle feature vector;

[0105] a student establishment module, configured to establish an initial student model and align an output of the initial student model with the speckle feature vector;

[0106] A distillation module is used to distill the output result of the initial student model using the speckle feature vector to obtain a final student model.

[0107] Specifically, the teacher training module receives a training set containing aligned RGB images, depth images, and speckle images. A pre-trained RGBD face recognition model is used as the base model. The RGBD face recognition model is fine-tuned using the training set so that it can simultaneously learn the feature representations of RGB images, depth images, and speckle images. During the fine-tuning process, the learning rate, optimizer, and loss function (such as cross entropy loss and feature consistency loss) are adjusted to ensure that the model can learn robust feature representations. The teacher face model is obtained, whose output includes RGB feature vectors, depth feature vectors, and speckle feature vectors.

[0108] The teacher training module is the foundation of the entire system and is responsible for generating a teacher model with strong performance and rich feature representation. The speckle feature vector of the teacher model is passed to the student model as "knowledge" to provide guidance for the subsequent distillation process.

[0109] The student setup module receives the speckle feature vectors of the teacher face model generated by the teacher training module. It then establishes an initial student model with a relatively simple structure and improved computational efficiency. A learnable projection matrix is ​​added after the output layer of the initial student model to map the student model's output feature vectors to a feature space identical to or similar to the teacher model's speckle feature vectors. Feature alignment is achieved by minimizing the Euclidean distance between the student model output and the teacher model's speckle feature vectors using an L2 regularization loss. This results in an initial student model aligned with the teacher model's speckle feature vectors.

[0110] The goal of the student setup module is to build a lightweight initial student model and align its output feature vector with the speckle feature vector of the teacher model. This module lays the foundation for the subsequent distillation process, ensuring that the student model can learn the key feature representations of the teacher model.

[0111] The distillation module receives the initial student model generated by the student setup module and the speckle feature vector generated by the teacher training module. Using a temperature-adjusted KL divergence loss function, the teacher model's speckle feature vector is used as a "soft target" to distill the student model's output. The distillation weight coefficient is dynamically adjusted, linearly increasing from 0.1 to 0.5, to balance the student model's classification loss and distillation loss. The student model's performance is further optimized by optimizing the composite loss function (student model classification loss + distillation loss). The resulting student model is simple in structure, computationally efficient, and performs close to the teacher model.

[0112] The distillation module is the final step in the system, further optimizing the student model's performance through knowledge distillation. It relies on the speckle feature vectors generated by the teacher training module and the initial student model generated by the student setup module, ultimately outputting a lightweight and high-performance student model.

[0113] This embodiment utilizes the existing RGBD face recognition model. By constructing a multimodal teacher model and distilling its speckle feature knowledge into a lightweight student model, it can quickly obtain a high-quality speckle face recognition model using a highly stable RGBD face recognition model when the number of speckle images is small. This has important practical significance and application value.

[0114] An embodiment of the present invention further provides a facial speckle recognition model training device, comprising a processor and a memory storing executable instructions for the processor. The processor is configured to execute the executable instructions to perform the steps of a facial speckle recognition model training method.

[0115] As mentioned above, this embodiment utilizes the existing RGBD face recognition model. By constructing a multimodal teacher model and distilling its speckle feature knowledge into a lightweight student model, it can quickly obtain a high-quality speckle face recognition model using a highly stable RGBD face recognition model when the number of speckle images is small, which has important practical significance and application value.

[0116] Those skilled in the art will appreciate that various aspects of the present invention may be implemented as systems, methods, or program products. Accordingly, various aspects of the present invention may be implemented in the following forms: entirely in hardware, entirely in software (including firmware, microcode, etc.), or in a combination of hardware and software, collectively referred to herein as "circuits," "modules," or "platforms."

[0117] Figure 4 This is a structural diagram of a facial speckle recognition model training device in an embodiment of the present invention. Figure 4 An electronic device 600 according to this embodiment of the present invention will be described. Figure 4 The electronic device 600 shown is merely an example and should not limit the functions and scope of use of the embodiments of the present invention.

[0118] like Figure 4 As shown, electronic device 600 is implemented as a general-purpose computing device. Components of electronic device 600 may include, but are not limited to, at least one processing unit 610, at least one storage unit 620, a bus 630 connecting different platform components (including storage unit 620 and processing unit 610), and a display unit 640.

[0119] The storage unit stores program codes, which can be executed by the processing unit 610, so that the processing unit 610 performs the steps according to various exemplary embodiments of the present invention described in the above-mentioned method for training a facial speckle recognition model. For example, the processing unit 610 can perform the following steps: Figure 1 Follow the steps shown in .

[0120] The storage unit 620 may include a readable medium in the form of a volatile storage unit, such as a random access memory unit (RAM) 6201 and / or a cache memory unit 6202 , and may further include a read-only memory unit (ROM) 6203 .

[0121] The storage unit 620 may also include a program / utility 6204 having a set (at least one) of program modules 6205, such program modules 6205 including, but not limited to, an operating system, one or more application programs, other program modules, and program data, each of which or some combination may include an implementation of a grid environment.

[0122] Bus 630 may represent one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, a processing unit, or a local bus using any of a variety of bus architectures.

[0123] The electronic device 600 may also communicate with one or more external devices 700 (e.g., keyboards, pointing devices, Bluetooth devices, etc.), one or more devices that enable a user to interact with the electronic device 600, and / or any device that enables the electronic device 600 to communicate with one or more other computing devices (e.g., routers, modems, etc.). Such communication may be performed through an input / output (I / O) interface 650. Furthermore, the electronic device 600 may also communicate with one or more grids (e.g., a local area network (LAN), a wide area network (WAN), and / or a public grid, such as the Internet) through a grid adapter 660. The grid adapter 660 may communicate with other modules of the electronic device 600 through the bus 630. It should be understood that although Figure 4 Not shown, other hardware and / or software modules may be used in conjunction with electronic device 600, including but not limited to microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage platforms.

[0124] An embodiment of the present invention also provides a computer-readable storage medium for storing a program that, when executed, implements the steps of a facial speckle recognition model training method. In some possible implementations, various aspects of the present invention may also be implemented in the form of a program product, which includes program code. When the program product is executed on a terminal device, the program code is configured to cause the terminal device to execute the steps according to various exemplary embodiments of the present invention described in the aforementioned section of this specification regarding a facial speckle recognition model training method.

[0125] As shown above, this embodiment utilizes the existing RGBD face recognition model. By constructing a multimodal teacher model and distilling its speckle feature knowledge into a lightweight student model, it can quickly obtain a high-quality speckle face recognition model using a highly stable RGBD face recognition model when the number of speckle images is small, which has important practical significance and application value.

[0126] Figure 5 Schematic diagram of the structure of the computer-readable storage medium in an embodiment of the present invention. Figure 5 , a program product 800 for implementing the above method according to an embodiment of the present invention is described. The program product 800 may be a portable compact disc read-only memory (CD-ROM) and include program code, and may be run on a terminal device, such as a personal computer. However, the program product of the present invention is not limited thereto. In this document, a readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.

[0127] The program product may employ any combination of one or more readable media. The readable medium may be a readable signal medium or a readable storage medium. The readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination thereof. More specific examples (a non-exhaustive list) of readable storage media include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof.

[0128] Computer-readable storage media may include a data signal propagated in baseband or as part of a carrier wave, which carries readable program code. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. The readable storage medium may also be any readable medium other than a readable storage medium, which may send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the readable storage medium may be transmitted using any appropriate medium, including but not limited to wireless, wired, optical cable, RF, etc., or any suitable combination thereof.

[0129] The program code for performing the operations of the present invention may be written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Java, C++, and the like, as well as conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the user computing device, partially on the user device, as a stand-alone software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server. In cases involving a remote computing device, the remote computing device may be connected to the user computing device via any type of grid, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., via the Internet using an Internet service provider).

[0130] This embodiment utilizes the existing RGBD face recognition model. By constructing a multimodal teacher model and distilling its speckle feature knowledge into a lightweight student model, it can quickly obtain a high-quality speckle face recognition model using a highly stable RGBD face recognition model when the number of speckle images is small. This has important practical significance and application value.

[0131] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same and similar parts between the various embodiments can be referred to each other. The above description of the disclosed embodiments enables professionals and technicians in this field to implement or use the present invention. Various modifications to these embodiments will be apparent to professionals and technicians in this field, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but will conform to the widest scope consistent with the principles and novel features disclosed herein.

[0132] The above describes the specific embodiments of the present invention. It should be understood that the present invention is not limited to the above specific embodiments, and those skilled in the art may make various variations or modifications within the scope of the claims, which do not affect the essence of the present invention.

Claims

1. A facial speckle recognition model training method, characterized in that: include: Step S1: obtaining an RGBD face recognition model, and fine-tuning the RGBD face recognition model using a training set to obtain a teacher face model; wherein the training set includes aligned RGB images, depth images, and speckle images; and the output of the teacher face model includes an RGB feature vector, a depth feature vector, and a speckle feature vector; Step S2: establishing an initial student model, and aligning the output of the initial student model with the speckle feature vector; Step S3: using the speckle feature vector to distill the output result of the initial student model to obtain a final student model.

2. The facial speckle recognition model training method according to claim 1, characterized in that: The RGBD face recognition model is built based on a deep learning framework, and the deep learning framework is selected from TensorFlow, PyTorch or Caffe.

3. The facial speckle recognition model training method according to claim 1, characterized in that: The training set also includes an infrared image aligned with the RGB image. In step S3, the infrared image is also used as auxiliary supervision for the distillation of the initial student model.

4. The facial speckle recognition model training method according to claim 1, characterized in that: The training set is obtained by preprocessing the collected RGB images, depth images and speckle images, and the preprocessing includes image alignment, normalization and data enhancement.

5. The facial speckle recognition model training method according to claim 1, characterized in that: The fine-tuning process of the teacher face model includes adjusting the learning rate, optimizer and loss function, and the loss function includes cross entropy loss and feature consistency loss.

6. The facial speckle recognition model training method according to claim 1, characterized in that: When performing the feature alignment operation in step S2, a learnable projection matrix is ​​added after the initial student model output layer; and the Euclidean distance between the student model output and the speckle feature vector is minimized by L2 regularization loss.

7. The facial speckle recognition model training method according to claim 1, characterized in that: The distillation process in step S3 adopts a temperature-regulated KL divergence loss function; a dynamically adjusted distillation weight coefficient, and the distillation weight coefficient increases linearly from 0.1 to 0.5 with the training rounds.

8. A facial speckle recognition model training system, used to implement the facial speckle recognition model training method according to any one of claims 1 to 7, characterized in that: include: a teacher training module for obtaining an RGBD face recognition model, and fine-tuning the RGBD face recognition model using a training set to obtain a teacher face model; wherein the training set includes aligned RGB images, depth images, and speckle images; and the output of the teacher face model includes an RGB feature vector, a depth feature vector, and a speckle feature vector; a student establishment module, configured to establish an initial student model and align an output of the initial student model with the speckle feature vector; A distillation module is used to distill the output result of the initial student model using the speckle feature vector to obtain a final student model.

9. A facial speckle recognition model training device, characterized in that: include: processor; a memory storing executable instructions for the processor; The processor is configured to execute the steps of the facial speckle recognition model training method according to any one of claims 1 to 7 by executing the executable instructions.

10. A computer-readable storage medium for storing a program, characterized in that: When the program is executed, the steps of the facial speckle recognition model training method according to any one of claims 1 to 7 are implemented.