Model training method, fingerprint classification method, device, medium and program product

By introducing collaborative losses and confrontation losses in online knowledge distillation technology, the noise problem of student models in feature representation is solved, the accuracy of fingerprint classification is improved, and the feature diversity and prediction ability of teacher models are enhanced.

CN120220193APending Publication Date: 2025-06-27JIHAO TECHNOLOGY (TIANJIN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411413046.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-10-10
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

In online knowledge distillation technology, since there is a lot of noise information in the characteristics of the student model, directly imitating the global feature representation of the teacher model will lead to strict constraints on the teacher model, which will affect the learning effect of the student model, resulting in a low accuracy of fingerprint classification.

Method used

By setting collaboration loss and adversarial loss in model training, the collaboration loss makes the student model focus on areas with high similarity to the middle layer feature map of the teacher model, and learns accurate and effective feature representation; adversarial loss makes the teacher model focus on areas with low similarity to the feature map, avoiding rapid convergence to the local optimal solution, and enhancing feature diversity and prediction capabilities.

Benefits of technology

Through this mechanism, the student model can better learn feature representations similar to the teacher model, improve the accuracy of fingerprint classification, and at the same time, the feature diversity and prediction ability of the teacher model have also been improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120220193A_ABST
    Figure CN120220193A_ABST
Patent Text Reader

Abstract

The invention provides a model training method, a fingerprint classification method, equipment, a medium and a program product, and the method comprises the steps: inputting training samples into a teacher model and a student model; calculating training loss; wherein the training loss comprises cooperation loss for the student model; and according to the training loss, updating model parameters of the teacher model and the student model, and obtaining a trained student model. According to the scheme, the online knowledge distillation technology is utilized to carry out mutual learning on the teacher model and the student model, and the cooperation loss is set for the student model during training, so that the student model pays more attention to an area with relatively high feature similarity between the interlayer feature maps extracted by the student model and the teacher model; therefore, the student model is helped to pay more attention to the features which contribute to the classification of the prediction classification result, so that more accurate and effective feature representation can be learned, and the image classification capability of the student model can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular, to a model training method, a fingerprint classification method, a device, a medium, and a program product. Background Art

[0002] Online Knowledge Distillation (OKD) is a technology for deep learning model compression and optimization. In the OKD technology, the teacher model and the student model perform synchronous learning and provide guidance to each other. In fingerprint classification technology, due to the complexity and diversity of fingerprint image data, it is usually necessary to train a large deep learning model to achieve a high fingerprint classification accuracy. However, large models may be restricted by computing resources, storage space, and response speed during deployment. By using the online knowledge distillation technology, mutual learning between a large fingerprint classification model and a small fingerprint classification model can be realized, so as to obtain a small fingerprint classification model with better performance, and the small fingerprint classification model can be deployed on actual devices.

[0003] When the related technology performs mutual learning between the teacher model and the student model by using the online knowledge distillation technology, due to the existence of a large amount of noise information in the student model features, the student model directly imitating the global feature representation of the teacher model will cause the noise information to impose relatively strict constraints on the teacher model, resulting in a decline in the performance of the teacher model, and further affecting the learning effect of the student model. The classification accuracy is relatively low when using such a student model for fingerprint classification. Summary of the Invention

[0004] The purpose of the embodiments of the present application is to provide a model training method, a fingerprint classification method, a device, a medium, and a program product to improve the classification accuracy of the fingerprint classification model.

[0005] In a first aspect, the embodiments of the present application provide a model training method, and the method includes:

[0006] Inputting training samples into a teacher model and a student model respectively; wherein, the training samples are training images with real classification results, and the teacher model and the student model are used to perform classification prediction on the training images;

[0007] Calculate the training loss; wherein, the training loss includes a collaborative loss for the student model; the collaborative loss is used to characterize the difference between the predicted classification result obtained by the student model based on the first attention feature and the true classification result; the first attention feature is the result of the first weighted calculation performed by the student model on the first intermediate layer feature map extracted from the training samples; the first weight used in the first weighted calculation is determined according to the similarity between the first intermediate layer feature map and the second intermediate layer feature map extracted by the teacher model from the training samples, and for regions with higher similarity between the first intermediate layer feature map and the second intermediate layer feature map, the corresponding first weight is larger;

[0008] Update the model parameters of the teacher model and the student model according to the training loss, and obtain the trained student model.

[0009] In the implementation process of the above solution, the online knowledge distillation technology is used to perform mutual learning between the teacher model and the student model. When training, a collaborative loss is set for the student model, so that the student model pays more attention to the regions with higher feature similarity between the intermediate layer feature maps extracted by the student model and the teacher model. This helps the student model pay more attention to the features that contribute more to the classification of the predicted classification result, thereby learning more accurate and effective feature representations, which is beneficial to improving the image classification ability of the student model.

[0010] In an implementation manner of the first aspect, the training loss further includes an adversarial loss for the teacher model; wherein, the adversarial loss is used to characterize the difference between the predicted classification result obtained by the teacher model based on the second attention feature and the true classification result; the second attention feature is the result of the second weighted calculation performed by the teacher model on the second intermediate layer feature map extracted from the training samples; the second weight used in the second weighted calculation is determined according to the similarity between the first intermediate layer feature map and the second intermediate layer feature map, and for regions with lower similarity between the first intermediate layer feature map and the second intermediate layer feature map, the corresponding second weight is larger.

[0011] In the implementation process of the above solution, an adversarial loss is set during the training of the teacher model and the student model. On the one hand, the adversarial loss enables the teacher model to pay more attention to the regions with lower similarity in the intermediate layer feature maps, which helps the teacher model avoid quickly converging to a local optimal solution similar to the student model and is beneficial to improving the diversity of the features learned by the teacher model. On the other hand, the adversarial loss encourages the teacher model to strengthen the learning of the regions with lower similarity between the feature maps, which helps improve the overall prediction ability of the teacher model. On the other hand, the adversarial loss and the collaborative loss work together. The student model uses the collaborative loss to continuously improve its feature extraction ability in the regions with higher similarity between the feature maps and the teacher model, while the teacher model uses the adversarial loss to continuously strengthen its feature extraction ability in the regions with lower similarity between the feature maps. This mutually promoting mechanism helps the student model better catch up with the feature learning ability of the teacher model.

[0012] In one implementation of the first aspect, the training loss further includes a feature learning loss for the student model; wherein, the feature learning loss is used to characterize the difference between the first intermediate layer feature map and the second intermediate layer feature map.

[0013] In the implementation process of the above solution, by setting a feature learning loss only for the student model, on the one hand, the student model can learn the intermediate layer feature expression ability of the teacher model, which is beneficial to improving the classification ability of the student model; on the other hand, the gradient backpropagation process of the feature learning loss to the teacher model is cut off, thereby avoiding the teacher model being interfered by the noise information of the student model, which is beneficial to improving the classification performance of the teacher model and further improving the classification performance of the student model.

[0014] In one implementation of the first aspect, calculating the training loss includes: obtaining the first intermediate layer feature map and the second intermediate layer feature map; wherein, the first intermediate layer feature map is the input feature map of the global average pooling layer in the student model, and the second intermediate layer feature map is the input feature map of the global average pooling layer in the teacher model; calculating the training loss based on the first intermediate layer feature map and the second intermediate layer feature map.

[0015] In the implementation process of the above solution, the input feature map of the global average pooling layer in the model is set as the intermediate layer feature map. Since the global average pooling layer is usually located at the end of the model architecture and its input feature map has been processed by multiple convolutional layers in front of the global average pooling layer, it contains relatively rich semantic information. Using these features to calculate the collaborative loss, adversarial loss, and feature learning loss will enable the student model to learn the deep feature representation ability of the teacher model, which is beneficial to improving the classification performance of the student model.

[0016] In an implementation of the first aspect, the training loss further includes a first supervision loss for the student model and a second supervision loss for the teacher model; wherein, the first supervision loss is used to characterize the difference between the first prediction result and the true classification result; the second supervision loss is used to characterize the difference between the second prediction result and the true classification result; the first prediction result and the second prediction result are the prediction results obtained by the student model and the teacher model based on the training samples, respectively.

[0017] In the implementation process of the above solution, a first supervision loss for the student model and a second supervision loss for the teacher model are respectively set for the model training process, so as to use the supervision information to guide the model learning process, thereby helping the student model to obtain better classification performance.

[0018] In an implementation of the first aspect, the training loss further includes a first matching loss for the student model and a second matching loss for the teacher model; wherein, both the first matching loss and the second matching loss are used to characterize the difference between the first prediction result probability distribution and the second prediction result probability distribution; the first prediction result probability distribution and the second prediction result probability distribution are the prediction result probability distributions obtained by the student model and the teacher model based on the training samples, respectively.

[0019] In the implementation process of the above solution, by respectively setting a first matching loss for the student model and a second matching loss for the teacher model in the training process, the model can also learn the prediction result probability distribution knowledge between models, thereby improving the generalization ability of the student model.

[0020] In a second aspect, an embodiment of the present application provides a fingerprint classification method, the method includes: obtaining a fingerprint image to be classified; inputting the fingerprint image to be classified into a fingerprint classification model to obtain a fingerprint classification result; wherein, the fingerprint classification model is a trained student model obtained by using the method provided in the first aspect or any possible implementation manner of the first aspect.

[0021] In a third aspect, an embodiment of the present application provides an electronic device, including: a processor, a memory and a communication bus, wherein the processor and the memory complete communication with each other through the communication bus; computer program instructions executable by the processor are stored in the memory, and when the computer program instructions are read and run by the processor, the methods provided in the first aspect or any possible implementation manner of the first aspect or the second aspect or any possible implementation manner of the second aspect are executed.

[0022] Fourthly, an embodiment of the present application provides a computer-readable storage medium, on which computer program instructions are stored. When the computer program instructions are read and run by a processor, the methods provided in the first aspect or any possible implementation manner of the first aspect or the second aspect or any possible implementation manner of the second aspect are executed.

[0023] Fifthly, an embodiment of the present application provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the methods provided in the first aspect or any possible implementation manner of the first aspect or the second aspect or any possible implementation manner of the second aspect are implemented.

[0024] Other features and advantages of the present application will be described in the subsequent description. And, partly, they will become obvious from the description, or can be understood by implementing the embodiments of the present application. The objectives and other advantages of the present application can be achieved and obtained by the structures specifically pointed out in the written description, claims, and drawings. Description of the Drawings

[0025] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required to be used in the embodiments of the present application will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present application, and thus should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.

[0026] Figure 1 It is a schematic flowchart of the model training method provided by the embodiment of the present application;

[0027] Figure 2 It is a training strategy framework diagram of the teacher model and the student model in a certain scenario provided by the embodiment of the present application;

[0028] Figure 3 It is a schematic structural diagram of the electronic device provided by the embodiment of the present application. Detailed Embodiments

[0029] Next, the technical solutions in the embodiments of the present application will be described with reference to the drawings in the embodiments of the present application. The following embodiments are only used to more clearly illustrate the technical solutions of the present application, and thus are only examples and cannot be used to limit the protection scope of the present application.

[0030] Knowledge Distillation (KD) is a deep learning model compression method that extracts the knowledge of a model with a large number of parameters (i.e., the teacher model) into a small and efficient model (i.e., the student model). The KD technology requires the teacher model to be a pre-trained model, and the knowledge transfer direction is also one-way, that is, from the teacher model to the student model.

[0031] Different from the knowledge distillation KD technology, the online knowledge distillation OKD technology does not require the teacher model to be a pre-trained model. The training processes of the teacher model and the student model are carried out synchronously, and the teacher model and the student model can provide guidance to each other during the training process. Since the teacher model and the student model in the OKD technology can adapt to and learn from each other, the probability of information loss is reduced, thereby obtaining a student model with better performance.

[0032] When using the OKD technology for knowledge distillation in related technologies, the student model is encouraged to directly learn the global feature representation of the teacher model, which makes the student model overly dependent on the teacher model, thus forming a "lazy" learning method, which makes the teacher model and the student model quickly converge to a local optimal solution. The classification accuracy of the fingerprint classification using the student model that converges to the local optimum is relatively low.

[0033] Based on this, the embodiments of the present application provide a model training method. This method uses the online knowledge distillation technology to perform mutual learning on the teacher model and the student model. When training, a collaborative loss is set for the student model, so that the student model pays more attention to the regions with high feature similarity between the intermediate layer feature maps extracted by the student model and the teacher model. This helps the student model pay more attention to the features that contribute more to the prediction classification result, thereby learning more accurate and effective feature representations, which is beneficial to improving the image classification ability of the student model.

[0034] The above model training method will be introduced in detail below:

[0035] Please refer to Figure 1 The embodiments of the present application provide a model training method, which includes:

[0036] Step S110: Input the training samples into the teacher model and the student model respectively;

[0037] Among them, the training samples are training images with real classification results, and the teacher model and the student model are used to perform classification predictions on the training images; the real classification results carried by the training samples refer to the correct class labels corresponding to each training image, and these labels can be used to guide the learning process of the model;

[0038] Step S120: Calculate the training loss;

[0039] Step S130: Update the model parameters of the teacher model and the student model according to the training loss to obtain a trained student model.

[0040] The above teacher model can be a large and complex deep learning model with a large number of parameters, and the student model can be a small and lightweight deep learning model with a small number of parameters. Both the teacher model and the student model can classify the input images of the model.

[0041] It can be understood that the application scenarios of the above model training method include but are not limited to fingerprint classification, face classification, gesture classification, plant and animal classification, etc. In different application scenarios, the input images of the above teacher model and student model are also different. For example, in the fingerprint classification scenario, the input images of the above teacher model and student model can be fingerprint images; in the face classification scenario, the input images of the above teacher model and student model can be face images, etc.

[0042] In addition, it can be understood that the above fingerprint images can be the fingerprint images collected in biometric technology. In the application scenario of fingerprint classification, the fingerprint images can be classified to accurately identify the individual identity.

[0043] The trained student model obtained in the above step S130 can also be applied to the biometric scenario to identify the information of the fingerprint in the fingerprint image.

[0044] The training loss adopted in the training process is introduced below:

[0045] (1) The collaborative loss for the student model;

[0046] The collaborative loss is used to characterize the difference between the predicted classification result obtained by the student model based on the first attention feature and the true classification result.

[0047] The above first attention feature is the result of the first weighted calculation of the first intermediate layer feature map extracted by the student model based on the training samples. The first weight used in the first weighted calculation is determined according to the similarity between the first intermediate layer feature map and the second intermediate layer feature map extracted by the teacher model based on the training samples. For the regions with higher similarity between the first intermediate layer feature map and the second intermediate layer feature map, the corresponding first weight is larger.

[0048] The above first intermediate layer feature map is the feature map output by the intermediate layer of the student model. The intermediate layer refers to the layer located between the input layer and the output layer of the model. In the embodiments of the present application, the intermediate layer feature Figure 1 generally refers to the feature map processed by the convolutional layer. Of course, the intermediate layer feature map may also be processed by the activation layer.

[0049] Before introducing the specific implementation of the above collaborative loss, the reasons for setting the collaborative loss in the embodiments of this application are introduced first:

[0050] The inventors of this application have found through research that:

[0051] 1. There is a certain degree of similarity between the first intermediate layer feature map extracted by the student model based on the training samples and the second intermediate layer feature map extracted by the teacher model based on the training samples. The similar features are mainly concentrated on the foreground objects of the training samples, and in the fingerprint classification scenario, the foreground object is the fingerprint object in the fingerprint image.

[0052] 2. Compared with the student model, the second intermediate layer feature map extracted by the teacher model pays more attention to the foreground object.

[0053] It can be understood that in a neural network model, each convolutional layer will try to extract meaningful features from the input image, that is, the features that are most helpful for the task. If the teacher model and the student model show a high degree of similarity in some regions of the feature map, it means that these regions in the feature map are the regions that the model focuses on during prediction, and the relevant features in these regions may also be the parts most relevant to class discrimination. From this, it can be inferred that the features in the regions with a high degree of similarity in the intermediate layer feature map are the features that make a greater contribution to prediction classification.

[0054] Based on the above findings, the embodiments of this application set a collaborative loss during the training process of the teacher model and the student model. On the one hand, the collaborative loss can enable the student model to pay more attention to the regions with a high degree of similarity in the intermediate layer feature map, thereby helping the student model learn more accurate and effective feature representations; on the other hand, the collaborative loss also encourages the student model to catch up with the feature learning process of the teacher model, so that the student model obtains better classification performance.

[0055] Next, an implementation of the collaborative loss is introduced:

[0056] The expression of the collaborative loss in this implementation can be:

[0057]

[0058] where L co is the collaborative loss; L ce (·) is the cross-entropy operation; W s is the parameter of the last linear classifier in the student model; ψ avg (·) is the global average pooling GAP operation; S is the similarity matrix of the first intermediate layer feature map and the second intermediate layer feature map; is the average similarity calculated along the height H and width W directions of the similarity matrix; Fs is the first intermediate layer feature map extracted by the student model based on the training samples; y is the true classification result;

[0059] It can be understood that the above is the first weight, is the first attention feature obtained by weighting the first intermediate layer feature map based on the first weight, is the predicted classification result obtained by the student model based on the first attention feature.

[0060] The above first weight is a monotonically increasing function of S, that is, for regions with higher similarity between feature maps, the corresponding first weight is larger. Of course, in addition to this form, the first weight can also be a monotonically increasing exponential function or a monotonically increasing power function, etc., and the specific form is set according to the specific application scenario.

[0061] In addition, it should be noted that the above collaborative loss uses cross-entropy to measure the difference between the two classification results, but this is not a specific limitation on the collaborative loss. The collaborative loss in the embodiments of this application can adopt a calculation method that can measure the difference between the two classification results, and is not limited to the cross-entropy calculation method.

[0062] In addition, it should be noted that the above formula uses the parameter W of the last linear classifier in the student model s and the global average pooling GAP operation ψ avg (·), which is set based on the architecture of the student model. Specifically: if the student model finally uses the global average pooling operation and the linear classifier to obtain the classification result, the above formula can be used to calculate the collaborative loss; if the student model required in some scenarios uses other architectures, then the above collaborative loss expression can be adaptively modified according to the specific architecture of the student model.

[0063] The above similarity matrix can be calculated using cosine similarity, and the calculation method can be:

[0064]

[0065] where, is the feature vector at the i-th row and j-th column in the first intermediate layer feature map extracted by the student model based on the training samples; is the feature vector at the i-th row and j-th column in the second intermediate layer feature map extracted by the teacher model based on the training samples; D(·) is the operation for calculating cosine similarity.

[0066] Of course, in addition to cosine similarity, other metrics can also be used for calculation, such as Euclidean distance, Hamming distance, etc., as long as the similarity measurement between the first intermediate layer feature map and the second intermediate layer feature map can be achieved.

[0067] (2) The adversarial loss for the teacher model;

[0068] The adversarial loss is used to characterize the difference between the predicted classification result obtained by the teacher model based on the second attention feature and the true classification result.

[0069] The above-mentioned second attention feature is the result of the second weighted calculation performed by the teacher model on the second intermediate layer feature map extracted based on the training samples; the second weight used in the second weighted calculation is determined according to the similarity between the first intermediate layer feature map and the second intermediate layer feature map. For regions with lower similarity between the first intermediate layer feature map and the second intermediate layer feature map, the corresponding second weight is larger.

[0070] The above-mentioned second intermediate layer feature map is the feature map output by the intermediate layer of the teacher model, and the intermediate layer refers to the layer located between the input layer and the output layer of the model. In the embodiments of the present application, the intermediate layer feature Figure 1 generally refers to the feature map processed by the convolutional layer. Of course, the intermediate layer feature map may also be processed by the activation layer.

[0071] Before introducing the specific implementation manner of the above-mentioned adversarial loss, the reason for setting the adversarial loss in the embodiments of the present application will be introduced first:

[0072] In the above introduction of the collaborative loss, the content discovered by the inventors of the present application through research is described. Based on these discoveries, the embodiments of the present application set the adversarial loss during the training process of the teacher model and the student model. On the one hand, the adversarial loss enables the teacher model to pay more attention to regions with lower similarity in the intermediate layer feature map, which helps the teacher model avoid quickly converging to a local optimal solution similar to the student model and is beneficial to improving the diversity of the features learned by the teacher model; on the other hand, the adversarial loss encourages the teacher model to strengthen the learning of regions with lower similarity between feature maps, which helps to improve the overall prediction ability of the teacher model; on the third hand, the adversarial loss and the collaborative loss cooperate together. The student model uses the collaborative loss to continuously improve the feature extraction ability in regions with higher similarity between feature maps and the teacher model, while the teacher model uses the adversarial loss to continuously strengthen the feature extraction ability in regions with lower similarity between feature maps. This mutually promoting mechanism helps the student model better catch up with the feature learning ability of the teacher model.

[0073] Next, an implementation manner of the above-mentioned adversarial loss will be introduced:

[0074] The expression of the adversarial loss in this implementation manner can be:

[0075]

[0076] Among them, L ad is the adversarial loss; L ce (·) is the cross-entropy operation; W t is the parameter of the last linear classifier in the teacher model; ψ avg (·) is the global average pooling GAP operation; S is the similarity matrix between the first intermediate layer feature map and the second intermediate layer feature map; is the average similarity calculated along the height H and width W directions of the similarity matrix; F t is the second intermediate layer feature map extracted by the teacher model based on the training samples; y is the true classification result.

[0077] It can be understood that the above is the second weight, is the second attention feature obtained by weighting the second intermediate layer feature map based on the second weight, is the predicted classification result obtained by the teacher model based on the second attention feature.

[0078] The above second weight is a monotonically decreasing function of S, that is, for regions with lower similarity between feature maps, the corresponding second weight is larger. Of course, in addition to this form, the second weight can also be other forms of monotonically decreasing functions of S, such as negative exponential functions, negative power functions, etc.

[0079] In addition, it should be noted that the above adversarial loss uses cross-entropy to measure the difference between two classification results, but this is not a specific limitation on the adversarial loss. The adversarial loss in the embodiments of the present application can adopt a calculation method that can measure the difference between two classification results, and is not limited to the cross-entropy calculation method.

[0080] In addition, it should be noted that the above formula uses the parameter W t of the last linear classifier in the teacher model and the global average pooling GAP operation ψ avg (·). This is set based on the architecture of the teacher model. Specifically: if the teacher model finally uses the global average pooling operation and the linear classifier to obtain the classification result, the above formula can be used to calculate the adversarial loss; if the teacher model set in some scenarios uses other architectures, then the above adversarial loss expression can be adaptively modified according to the specific architecture of the teacher model.

[0081] (3) Feature learning loss for the student model;

[0082] The feature learning loss is used to characterize the difference between the first intermediate layer feature map and the second intermediate layer feature map.

[0083] It should be noted that the feature learning loss in the embodiments of this application is only for the student model, and the reasons are as follows:

[0084] Currently, most of the OKD techniques in the related art are online prediction probability distillation. However, online prediction probability distillation only focuses on the difference between the student model and the teacher model in the prediction results, and does not pay attention to the process of feature extraction. The model obtained by online prediction probability distillation is not suitable for tasks that require accurate feature representation. In order to obtain a model with accurate feature representation ability in the embodiments of this application, an online feature distillation scheme is adopted. The online feature distillation scheme focuses on improving the performance of the student model by matching the intermediate layer feature maps between the teacher model and the student model. It attempts to let the student model learn the internal feature representation ability of the teacher model when processing tasks.

[0085] However, when performing online feature distillation, due to the small number of parameters of the student model, compared with the teacher model, the feature representation ability of the student model is limited. Therefore, during the training process, the intermediate layer feature map extraction process of the student model is easily affected by noise, so that the intermediate layer feature map extracted by the student model may contain a large amount of noise information. If, during the training process, the teacher model directly copies or imitates the intermediate layer feature map extracted by the student model, it will impose too strict constraints on the teacher model and more interference on the intermediate layer feature map extraction process of the teacher model, so that the teacher model can only obtain a sub-optimal solution.

[0086] For the above reasons, the feature learning loss in the embodiments of this application is only for the student model. That is, during the training process, the gradient of the feature learning loss only backpropagates to the student model, and the backpropagation process of the gradient to the teacher model is cut off.

[0087] Next, an implementation manner of the feature learning loss will be introduced:

[0088] The expression of the feature learning loss in this implementation manner can be:

[0089]

[0090] where L feat is the feature learning loss; B is the model training batch size; H, W, and C are the height, width, and number of channels of the intermediate layer feature map respectively; is the feature in the c-th channel at the i-th row and j-th column in the first intermediate layer feature map extracted by the student model based on the training samples during the b-th training; At the b-th training, it is the feature in the c-th channel of the i-th row and j-th column in the first intermediate feature map extracted by the teacher model based on the training samples; φ(·) is an alignment operation that uses a 1×1 convolutional layer to align and dimensions.

[0091] It can be understood that the above feature learning loss uses the mean square error to measure the difference between two features, but this is not a specific limitation on the feature learning loss. The feature learning loss in the embodiments of the present application can adopt a calculation method that can measure the difference between two features, and is not limited to the mean square error calculation method.

[0092] In addition, it can be understood that the above and are both normalization coefficients, which are only one implementation form of the normalization operation, and are not a specific limitation on the feature learning loss.

[0093] The above solution sets the feature learning loss only for the student model. On the one hand, it enables the student model to learn the expression ability of the intermediate feature map of the teacher model, which is beneficial to improving the classification ability of the student model; on the other hand, it cuts off the gradient backpropagation process of the feature learning loss to the teacher model, thereby avoiding the teacher model being interfered by the noise information of the student model, which is beneficial to improving the classification performance of the teacher model, and further improving the classification performance of the student model.

[0094] (4) Supervision loss;

[0095] It includes a first supervision loss for the student model and a second supervision loss for the teacher model; among them, the first supervision loss is used to represent the difference between the first prediction result and the true classification result, and the second supervision loss is used to represent the difference between the second prediction result and the true classification result. The first prediction result and the second prediction result are the prediction results obtained by the student model and the teacher model based on the training samples, respectively.

[0096] It can be understood that the above feature learning loss is used to guide the learning process of the model in the feature space, while the supervision loss is used to guide the learning process of the model under the supervision information.

[0097] In addition, it can be understood that the supervision loss can be calculated by using calculation methods such as cross entropy, which will not be elaborated in the embodiments of the present application.

[0098] The above solution sets a first supervision loss for the student model and a second supervision loss for the teacher model respectively in the model training process, so as to use the supervision information to guide the model learning process, thereby helping the student model to obtain better classification performance.

[0099] (5) Matching loss;

[0100] It includes a first matching loss for the student model and a second matching loss for the teacher model; wherein, both the first matching loss and the second matching loss are used to characterize the difference between the first prediction result probability distribution and the second prediction result probability distribution; the first prediction result probability distribution and the second prediction result probability distribution are the prediction result probability distributions obtained by the student model and the teacher model based on the training samples respectively.

[0101] It can be understood that the above matching loss can be calculated using the KL (Kullback-Leibler Divergence) divergence. The KL divergence is an index used to measure the matching degree of two probability distributions. The greater the difference between the two probability distributions, the greater the KL divergence value. The difference between the first prediction result probability distribution and the second prediction result probability distribution can be measured through the KL divergence.

[0102] The above solution sets a first matching loss for the student model and a second matching loss for the teacher model respectively in the training process, so that the models can also perform soft label learning, that is, learning of the prediction result probability distribution knowledge, thereby improving the generalization ability of the student model.

[0103] The following introduces the setting scheme of the intermediate layer feature maps involved in the above collaborative loss, adversarial loss, and feature learning loss:

[0104] The first setting method: Set the features output by all intermediate layers of the model as the intermediate layer feature maps for calculating the above collaborative loss, adversarial loss, and feature learning loss.

[0105] It can be understood that when calculating the above collaborative loss and adversarial loss using the intermediate layer feature maps output by different intermediate layers, the parameters of the linear classifier used are also different. For example:

[0106] For the input feature map of the last global average pooling layer, that is, the output feature of the last intermediate layer, the parameters of the linear classifier used when calculating the corresponding collaborative loss and adversarial loss can be the parameters of the last linear classifier in the model;

[0107] For the output features of other intermediate layers, the parameters of the linear classifier used when calculating the corresponding collaborative loss and adversarial loss can be the parameters of the linear classifier separately trained for this intermediate layer in advance.

[0108] The second setting method: Set the features output by some intermediate layers of the model as the intermediate layer feature maps for calculating the above collaborative loss, adversarial loss, and feature learning loss;

[0109] The set intermediate layer feature maps can include shallow features and / or deep features. Among them, shallow features refer to the features extracted by the convolutional layers at the front of the model (for example, convolutional layers with a layer number less than a certain set threshold), while deep features are the features extracted by the convolutional layers at the back of the model (for example, convolutional layers with a layer number not less than a certain set threshold). Generally, shallow features contain relatively rich information such as edges, textures, and colors, while deep features contain more high-level semantic information such as shapes and structures.

[0110] It can be understood that since the number of intermediate layers in the teacher model and the student model may not be the same, when selecting the features output by some intermediate layers as the intermediate layer feature maps, the correspondence of the intermediate layers in the student model and the teacher model can be considered. When selecting intermediate layers, intermediate layers with similar functions or similar positions in the two models can be selected, or representative intermediate layers in the models can be selected for correspondence. For example, the last convolutional layer in the two models can be selected.

[0111] The following introduces a scheme for setting deep features as intermediate layer feature maps:

[0112] As an optional implementation manner of the above model training method, the above step S120 includes:

[0113] Obtain a first intermediate layer feature map and a second intermediate layer feature map; among them, the first intermediate layer feature map is the input feature map of the global average pooling layer in the student model; the second intermediate layer feature map is the input feature map of the global average pooling layer in the teacher model; calculate the training loss based on the first intermediate layer feature map and the second intermediate layer feature map.

[0114] It can be understood that the above global average pooling layer refers to the last global average pooling layer in the model, and the input feature map of the last global average pooling layer is also the output feature of the last intermediate layer in the model. This feature is the deepest feature that the model can extract and contains relatively rich high-level semantic information.

[0115] The above scheme sets the input feature map of the global average pooling layer in the model as the intermediate layer feature map, and uses these features to calculate the collaborative loss, adversarial loss, and feature learning loss, which will enable the student model to learn the deepest feature representation ability of the teacher model and is beneficial to improving the classification performance of the student model.

[0116] After calculating the training loss, the above step S130 can use the gradient backpropagation strategy to update the model parameters of the teacher model and the student model.

[0117] It should be pointed out that, under the premise of being able to complete model training, the above-mentioned collaborative loss can be freely combined with adversarial loss, feature learning loss, supervision loss and matching loss to realize the training of teacher model and student model. The following introduces the training strategy of online feature distillation using the above-mentioned collaborative loss, adversarial loss, feature learning loss, supervision loss and matching loss in a certain scenario:

[0118] In this scenario, both the teacher model and the student model can be used for fingerprint classification tasks. The teacher model before training is a large model, and the student model before training is a small model. The overall training strategy is as follows: Figure 2 As shown, the training strategies of the teacher model and the student model are introduced below:

[0119] 1. The overall loss function for the student model is:

[0120] L s =L C1 +L KD1 +L feat +L co

[0121] (1)L C1 is the first supervision loss, which is used to characterize the first prediction result p obtained by the student model based on the training sample s The difference between the true classification result y and the first supervised loss can be calculated using cross entropy.

[0122] It is understandable that the prediction result p s is the predicted probability, which is used to indicate the probability that a sample belongs to a certain category.

[0123] (2)L KD1 It is the first matching loss, which is used to characterize the difference between the first prediction result probability distribution obtained by the student model based on the training samples and the second prediction result probability distribution obtained by the teacher model based on the training samples. The first matching loss can be calculated using KL divergence.

[0124] (3)L feat is the feature learning loss, used to characterize the first intermediate layer feature map F s and the second intermediate layer feature map F t The difference between.

[0125] (4)L co is the collaborative loss, used to characterize the student model based on the first attention feature A1F s The predicted classification result o s The difference between ′ and the true classification result y.

[0126] 2. Overall loss function for the teacher model:

[0127] L t = L C2 + L KD2 + L ad

[0128] (1)L C2 is the second supervision loss, which is used to characterize the difference between the second prediction result p t obtained by the teacher model based on the training samples and the true classification result y. The second supervision loss can be calculated using cross entropy.

[0129] (2)L KD2 is the second matching loss, which is used to characterize the difference between the probability distribution of the second prediction result obtained by the teacher model based on the training samples and the probability distribution of the first prediction result obtained by the student model based on the training samples. The second matching loss can be calculated using KL divergence.

[0130] (3)L ad is the adversarial loss, which is used to characterize the difference between the predicted classification result p t ' obtained by the teacher model based on the second attention feature A2F t and the true classification result y.

[0131] Based on the same inventive concept, an embodiment of the present application also provides a fingerprint classification method, which includes:

[0132] Step S210: Obtain a fingerprint image to be classified;

[0133] Step S220: Input the fingerprint image to be classified into the fingerprint classification model to obtain a fingerprint classification result.

[0134] The above fingerprint classification model is a trained student model obtained by using any one of the above model training methods. The fingerprint classification model can classify according to the pattern of the fingerprint lines, such as circular, arch, and spiral, etc. It can also classify according to specific features of the fingerprint, such as ridge lines, bifurcation points, and endpoints, etc. It can also classify according to the fingerprint state, such as dry, normal, and wet, etc. It can also be a classification of the fingerprint quality, such as poor fingerprint quality, normal fingerprint quality, and good fingerprint quality, etc. It can also be a classification of the fingerprint integrity, such as incomplete fingerprint and complete fingerprint, etc.

[0135] The above solution uses the trained student model obtained by the above model training method to classify the fingerprint images to be classified. On the one hand, the above student model is a student model obtained by using the online distillation technology, which can learn more effective and accurate feature expressions from the teacher model, facilitating the improvement of the fingerprint classification accuracy. On the other hand, the student model is a lightweight model, enabling the above fingerprint classification method to be deployed on devices with limited computing resources, which is conducive to improving the adaptability of the above fingerprint classification method. On the further hand, since the number of parameters of the student model obtained by using the online distillation technology is small, compared with large fingerprint classification models, the above student model has a faster inference speed and can provide a more rapid response.

[0136] The following introduces the above fingerprint classification method in combination with a specific application scenario. The application scenario is the fingerprint authentication scenario of a certain APP in a mobile terminal. In this scenario, the above fingerprint classification method includes:

[0137] Step 1: Obtain the fingerprint images to be classified;

[0138] The mobile terminal is provided with a fingerprint scanning area. When the user places a finger in the fingerprint scanning area, the mobile terminal will collect the fingerprint image.

[0139] Step 2: Input the fingerprint images to be classified into the fingerprint classification model to obtain the fingerprint classification result;

[0140] The above fingerprint classification model is a trained student model obtained by using any of the above model training methods and pre-deployed in the above mobile terminal.

[0141] It can be understood that after the mobile terminal obtains the fingerprint classification result, it can perform user authentication according to the fingerprint classification result. If the fingerprint image matches the fingerprint at the time of user registration, the verification passes.

[0142] Figure 3 This is a schematic diagram of an electronic device provided by an embodiment of the present application. Refer to Figure 3 , the electronic device 300 includes: a processor 310, a memory 320, and a communication interface 330. These components are interconnected and communicate with each other through a communication bus 340 and / or other forms of connection mechanisms (not shown).

[0143] Among them, the memory 320 includes one or more (only one is shown in the figure), which may be, but is not limited to, random access memory (RAM), read only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), etc. The processor 310 and other possible components can access, read, and / or write data in the memory 320.

[0144] The processor 310 includes one or more (only one is shown in the figure), which may be an integrated circuit chip with signal processing capabilities. The above-mentioned processor 310 may be a general-purpose processor, including a central processing unit (CPU), a microcontroller unit (MCU), a network processor (NP), or other conventional processors; it may also be a dedicated processor, including a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0145] The communication interface 330 includes one or more (only one is shown in the figure), which can be used to communicate directly or indirectly with other devices for data interaction. For example, the communication interface 330 may be an Ethernet interface; it may be a mobile communication network interface, such as an interface for 3G, 4G, or 5G networks; or it may be other types of interfaces with data transceiver functions.

[0146] One or more computer program instructions can be stored in the memory 320, and the processor 310 can read and run these computer program instructions to implement the model training method or fingerprint classification method provided in the embodiments of the present application and other desired functions.

[0147] It can be understood thatFigure 3 The structure shown is only schematic, and the electronic device 300 may also include more or fewer components than those shown in Figure 3 or have a different configuration from that shown in Figure 3 . Figure 3 Each component shown in can be implemented using hardware, software, or a combination thereof. For example, the electronic device 300 can be a single server (or other device with computing and processing capabilities), a combination of multiple servers, a cluster of a large number of servers, etc., and can be either a physical device or a virtual device.

[0148] It can be understood that the embodiments of the present application relate to a model training method and a fingerprint classification method. The electronic devices involved in these two methods can be different. The following introduces the electronic devices involved in the above two methods from the training stage and the application stage:

[0149] (1) Training stage:

[0150] In the model training stage, the electronic device usually requires high computing power to process a large amount of sample data and deep learning models. Therefore, the electronic device for executing the above model training method can be: a server, a cloud computing platform, etc.

[0151] (2) Application stage:

[0152] In the application stage of the student model, the electronic device needs to have a certain computing power to run the trained student model for fingerprint classification. Therefore, the electronic device for executing the above fingerprint classification method can be a smart mobile terminal, a smart home device, an embedded device, etc.

[0153] The electronic device deployed with the above fingerprint classification method can be used in application scenarios such as personal authentication, secure payment, access control, and attendance systems.

[0154] The embodiments of the present application also provide a computer-readable storage medium. When the computer program instructions stored on the computer-readable storage medium are read and run by the processor of the computer, they execute the model training method or the fingerprint classification method provided by the embodiments of the present application. For example, the computer-readable storage medium can be implemented as Figure 3 the memory 320 in the electronic device 300 in.

[0155] In the embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For another example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some communication interfaces. The indirect coupling or communication connection of the devices or units can be in electrical, mechanical or other forms.

[0156] In addition, the units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0157] Furthermore, in each embodiment of this application, the various functional modules can be integrated together to form an independent part, or each module can exist alone, or two or more modules can be integrated to form an independent part.

[0158] Based on the same inventive concept, an embodiment of this application also provides a computer program product. The product includes a computer program. When the computer program is executed by a processor, it implements any one of the above model training methods or any one of the fingerprint classification methods.

[0159] The above is only the embodiment of this application and is not used to limit the protection scope of this application. For those skilled in the art, this application can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of this application shall be included in the protection scope of this application.

Claims

1. A model training method, characterized in that: The method comprises: Inputting training samples into the teacher model and the student model respectively; wherein the training samples are training images with real classification results, and the teacher model and the student model are used to perform classification prediction on the training images; Calculate the training loss; wherein the training loss includes the collaborative loss for the student model; the collaborative loss is used to characterize the difference between the predicted classification result obtained by the student model based on the first attention feature and the true classification result; the first attention feature is the result of a first weighted calculation performed by the student model based on the first intermediate layer feature map extracted from the training sample; the first weight used in the first weighted calculation is determined according to the similarity between the first intermediate layer feature map and the second intermediate layer feature map extracted from the training sample by the teacher model, and the higher the similarity between the first intermediate layer feature map and the second intermediate layer feature map, the larger the first weight corresponding to the region; According to the training loss, the model parameters of the teacher model and the student model are updated to obtain the trained student model.

2. The model training method according to claim 1, characterized in that: The training loss also includes an adversarial loss for the teacher model; wherein the adversarial loss is used to characterize the difference between the predicted classification result obtained by the teacher model based on the second attention feature and the true classification result; the second attention feature is the result of a second weighted calculation performed by the teacher model based on the second intermediate layer feature map extracted by the training sample; the second weight used in the second weighted calculation is determined according to the similarity between the first intermediate layer feature map and the second intermediate layer feature map, and the lower the similarity between the first intermediate layer feature map and the second intermediate layer feature map, the larger the corresponding second weight.

3. The model training method according to claim 1, characterized in that: The training loss also includes a feature learning loss for the student model; wherein the feature learning loss is used to characterize the difference between the first intermediate layer feature map and the second intermediate layer feature map.

4. The model training method according to any one of claims 1 to 3, characterized in that: The calculation of training loss includes: Obtaining the first intermediate layer feature map and the second intermediate layer feature map; wherein the first intermediate layer feature map is an input feature map of the global average pooling layer in the student model, and the second intermediate layer feature map is an input feature map of the global average pooling layer in the teacher model; Based on the first intermediate layer feature map and the second intermediate layer feature map, a training loss is calculated.

5. The model training method according to any one of claims 1 to 3, characterized in that: The training loss also includes a first supervised loss for the student model and a second supervised loss for the teacher model; wherein the first supervised loss is used to characterize the difference between the first prediction result and the true classification result; the second supervised loss is used to characterize the difference between the second prediction result and the true classification result; the first prediction result and the second prediction result are respectively the prediction results obtained by the student model and the teacher model based on the training sample.

6. The model training method according to any one of claims 1 to 3, characterized in that: The training loss also includes a first matching loss for the student model and a second matching loss for the teacher model; wherein the first matching loss and the second matching loss are both used to characterize the difference between the first prediction result probability distribution and the second prediction result probability distribution; the first prediction result probability distribution and the second prediction result probability distribution are respectively the prediction result probability distributions obtained by the student model and the teacher model based on the training samples.

7. A fingerprint classification method, characterized in that: The method comprises: Obtaining a fingerprint image to be classified; The fingerprint image to be classified is input into a fingerprint classification model to obtain a fingerprint classification result; wherein the fingerprint classification model is a trained student model obtained by using the model training method according to any one of claims 1 to 6.

8. An electronic device, characterized in that: include: A processor, a memory and a communication bus, wherein the processor and the memory communicate with each other via the communication bus; The memory stores program instructions executable by the processor, and the processor can execute the method according to any one of claims 1 to 7 by calling the program instructions.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and when the computer instructions are executed by a computer, the computer is caused to perform the method according to any one of claims 1 to 7.

10. A computer program product, characterized in that The computer program product comprises a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.