Image recognition model training method, program product, electronic device and storage medium

Through the target distillation algorithm and dynamically updating the distillation center method, the problem of poor generalization ability in the training of IR image face recognition model is solved, and the generalization ability and training efficiency of the model are improved.

CN120012870APending Publication Date: 2025-05-16DIANYUN TECH (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510201917.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-24
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

In the prior art, in the face recognition model training based on IR images, due to the small data set size and inconsistent image types, the model generalization ability is poor, making it difficult to achieve the expected training effect.

Method used

The target distillation algorithm is used to transfer the knowledge of the teacher model obtained based on RGB image training to the student model, and the degree of dependence of the student model on the teacher model is reduced by dynamically updating the distillation center.

Benefits of technology

The generalization ability of image recognition model is improved, the problem of slow training convergence caused by structural differences between teacher model and student model is alleviated, and the convergence speed and training efficiency of the model are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120012870A_ABST
    Figure CN120012870A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of image recognition, and particularly provides an image recognition model training method, a program product, electronic equipment and a storage medium. The image recognition model training method comprises the following steps: taking a model obtained by training based on a first type of training image as a teacher model; migrating knowledge of the teacher model to a student model by adopting a target distillation algorithm; and based on a second type of target image and a target recognition result of the target image, performing adjustment training on the student model to obtain a trained image recognition model. According to the method, the target distillation algorithm is adopted, and the distillation center is dynamically updated based on the similarity between the teacher feature representation and the student feature representation in the distillation process, so that the degree of dependence of the student model on the teacher model is reduced, and the generalization ability of the student model obtained through pre-training is improved; and the generalization ability of the image recognition model obtained by adjusting and training the student model is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image recognition technology, and more specifically, to an image recognition model training method, a program product, an electronic device, and a storage medium. Background Art

[0002] Image recognition models refer to machine learning models that are specially designed to process image data. They play an important role in security monitoring, autonomous driving, medical diagnosis and other fields. A dataset refers to a set of data used to train and evaluate a model, and is the main source of data for model learning. For some datasets that are difficult to collect or have a small amount of data, the generalization ability of the model will be poor, affecting the model's image recognition accuracy and precision.

[0003] For example, face recognition technology based on IR images (infrared images output by infrared image sensors) does not have a large-scale public IR face dataset available, and face data collection is difficult. Currently, a feasible solution is to use a transfer learning method to pre-train the model based on a large-scale RGB face dataset, and then fine-tune the model using IR face images. However, the number of IR face images used for fine-tuning is small, and the image types used for fine-tuning (IR face images) and pre-training (RGB face dataset) are inconsistent. Usually, the pre-trained model needs to be significantly adjusted to complete the training. Such a fine-tuning process will also cause the model to lose its generalization ability, making it difficult to achieve the expected training effect. Summary of the invention

[0004] In view of this, the purpose of the embodiments of the present application is to provide an image recognition model training method, program product, electronic device and storage medium to solve the above technical problems.

[0005] In a first aspect, an embodiment of the present application provides an image recognition model training method, the method comprising:

[0006] Using a model obtained by training based on the first type of training images as a teacher model;

[0007] A target distillation algorithm is used to transfer the knowledge of the teacher model to the student model; wherein the target distillation algorithm includes: in the distillation process, based on the similarity between the teacher feature representation and the student feature representation, updating the distillation center; the teacher feature representation includes the feature representation of the training image by the teacher model, and the student feature representation includes the feature representation of the training image by the student model;

[0008] Based on the second type of target image and the target recognition result of the target image, the student model is adjusted and trained to obtain a trained image recognition model.

[0009] In the above implementation process, the image recognition model training method uses the model obtained by training the first type of training images as the teacher model; the target distillation algorithm is used to transfer the knowledge of the teacher model to the student model; wherein the target distillation algorithm includes: in the distillation process, based on the similarity between the teacher feature representation and the student feature representation, updating the distillation center; the teacher feature representation includes the feature representation of the training image by the teacher model, and the student feature representation includes the feature representation of the training image by the student model; based on the second type of target image and the target recognition result of the target image, the student model is adjusted and trained to obtain a trained image recognition model. The image recognition model training method can reduce the dependence of the student model on the teacher model by dynamically updating the distillation center based on the similarity between the teacher feature representation and the student feature representation during the distillation process, improve the generalization ability of the student model obtained by pre-training, and thus improve the generalization ability of the image recognition model obtained by adjusting and training the student model. In addition, by dynamically updating the distillation center using the similarity between the teacher feature representation and the student feature representation, the problem of slow convergence during training caused by the structural difference between the teacher model and the student model can also be alleviated, thereby improving the convergence speed of the model and the efficiency of model training.

[0010] Optionally, in an embodiment of the present application, the use of a target distillation algorithm to transfer the knowledge of the teacher model to the student model includes: dividing the training images into multiple image sets; wherein each of the image sets includes category images of multiple categories; based on the similarity between the teacher category feature representation of the category image by the teacher model and the student category feature representation of the category image by the student model, updating the distillation class center of the category image; based on the updated distillation class center, calculating the distillation loss value of the student model; and adjusting the model parameters of the student model according to the distillation loss value to transfer the knowledge of the teacher model to the student model.

[0011] In the above implementation process, by dividing the training images into multiple image sets, the multiple image sets can be assigned to different GPUs for processing in combination with the actual number of available GPUs, so that each GPU only needs to process part of the training images, thereby alleviating the GPU pressure during the training process; it can also improve the speed and efficiency of model distillation. The distillation class center of the category image is updated through the similarity between the teacher category feature representation of the category image by the teacher model and the student category feature representation of the category image by the student model, and the distillation loss value used to adjust the model parameters is calculated based on the updated distillation class center; this reduces the dependence of the student model on the teacher model and improves the generalization ability of the student model obtained by pre-training.

[0012] Optionally, in an embodiment of the present application, calculating the distillation loss value of the student model based on the updated distillation class center includes: sampling the updated distillation class center to obtain a sampled class center; calculating the distillation loss value of the student model according to the sampled class center and the student category feature representation of the category image corresponding to the sampled class center.

[0013] In the above implementation process, by sampling the updated distillation class center to obtain the sampling class center, and calculating the distillation loss value of the student model based on the sampling class center and the student category feature representation of the category image corresponding to the sampling class center, the computing resources required in the calculation process of the distillation loss value can be reduced, thereby further improving the deployability of the image recognition model training method.

[0014] Optionally, in an embodiment of the present application, the student model is adjusted and trained based on the second type of target image and the target recognition result of the target image to obtain a trained image recognition model, including: based on the student model, obtaining a target student feature representation of the target image; constructing a negative sample pair according to the target student feature representation of the target image and the target recognition result of the target image; wherein the negative sample pair includes two target student feature representations with different target recognition results; calculating a first adjusted loss value of the student model according to the negative similarity corresponding to the negative sample pair; and adjusting the model parameters of the student model based on the first adjusted loss value to obtain the trained image recognition model.

[0015] In the above implementation process, a negative sample pair is constructed based on the target student feature representation of the target image and the target recognition result of the target image in the fine-tuning stage; since the negative similarity between the negative sample pairs can reflect the probability of the image being misrecognized, the model parameters of the student model are adjusted based on the first adjustment loss value calculated based on the negative similarity corresponding to the negative sample pairs, which can reduce the probability of the image being misrecognized. For a biometric system, such as a face recognition system, the model parameters of the student model are adjusted based on the first adjustment loss value calculated based on the negative similarity corresponding to the negative sample pairs, and the image recognition model obtained can reduce the system's misrecognition rate, that is, improve the system's recognition accuracy for unauthorized users, thereby improving the security of the face recognition system.

[0016] Optionally, in an embodiment of the present application, the target recognition result includes an image category of the target image;

[0017] The method further comprises: calculating a second adjustment loss value based on a distillation center corresponding to an image category of the target image and a target student feature representation of the target image by the student model;

[0018] The step of adjusting the model parameters of the student model based on the first adjustment loss value to obtain the trained image recognition model includes: calculating an adjustment loss value based on the first adjustment loss value and the second adjustment loss value; and adjusting the model parameters of the student model based on the adjustment loss value to obtain the trained image recognition model.

[0019] In the above implementation process, the adjusted loss value calculated based on the first adjusted loss value and the second adjusted loss value can more comprehensively reflect the degree of fit between the model and the second type of target image; so as to further improve the recognition performance of the image recognition model for the second type of image.

[0020] Optionally, in the embodiment of the present application, the calculating the adjusted loss value based on the first adjusted loss value and the second adjusted loss value includes: calculating the adjusted loss value based on the first adjusted loss value, the second adjusted loss value and L T =L cls +β*L far , calculate the adjusted loss value L T Among them, L far represents the first adjusted loss value, L cls represents the second adjusted loss value, and β represents the weight coefficient of the first adjusted loss value.

[0021] In the above implementation process, through L T =L cls +β*L far , calculate the adjusted loss value L t ; For biometric systems, such as face recognition systems, the weight coefficient β of the first adjustment loss value can be increased to more strictly control the model's error rate, so as to further reduce the system error rate of the face recognition system and ensure the security of the face recognition system; for non-biometric systems, the weight coefficient β of the first adjustment loss value can be reduced to more comprehensively consider the model's error rate, recognition accuracy and recognition precision, so as to improve the overall recognition performance of the non-biometric system.

[0022] Optionally, in an embodiment of the present application, the calculating the first adjusted loss value of the student model according to the negative similarity corresponding to the negative sample pair includes: calculating the first adjusted loss value of the student model according to the negative similarity corresponding to the negative sample pair and Calculate the first adjusted loss value L far ; Wherein, N represents the number of negative sample pairs, si represents the negative similarity corresponding to the i-th negative sample pair, and T represents the preset similarity threshold.

[0023] In a second aspect, an embodiment of the present application provides a computer program product, including a computer program / instruction, which, when executed by a processor, implements the image recognition model training method as described in any one of the first aspects above.

[0024] In a third aspect, an embodiment of the present application further provides an electronic device; the electronic device includes:

[0025] Memory;

[0026] processor;

[0027] The memory stores a computer program executable by the processor, and when the computer program is executed by the processor, the image recognition model training method described in any one of the first aspects is executed.

[0028] In a fourth aspect, an embodiment of the present application further provides a computer-readable storage medium, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the image recognition model training method as described in any one of the first aspects is executed.

[0029] The beneficial effects of the present application include at least: the image recognition model training method can reduce the dependence of the student model on the teacher model and improve the generalization ability of the student model obtained by pre-training by dynamically updating the distillation center based on the similarity between the teacher feature representation and the student feature representation during the distillation process, thereby improving the generalization ability of the image recognition model obtained by adjusting the student model.

[0030] In addition, by utilizing the similarity between the teacher feature representation and the student feature representation and dynamically updating the distillation center, the problem of slow convergence during training caused by the structural differences between the teacher model and the student model can be alleviated, thereby improving the model's convergence speed and model training efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for use in the embodiments of the present application will be briefly introduced below. It should be understood that the following drawings only show certain embodiments of the present application and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying creative work.

[0032] Figure 1A flowchart of an image recognition model training method provided in an embodiment of the present application;

[0033] Figure 2 A flowchart of a knowledge transfer method provided in an embodiment of the present application;

[0034] Figure 3 A schematic diagram of a flow chart of a model pre-training process provided in an embodiment of the present application;

[0035] Figure 4 A flow chart of a method for adjusting and training a student model provided in an embodiment of the present application;

[0036] Figure 5 A schematic diagram of a process flow of a model fine-tuning process provided in an embodiment of the present application;

[0037] Figure 6 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0038] The following embodiments of the technical solution of the present application will be described in detail in conjunction with the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solution of the present application, and are therefore only used as examples, and cannot be used to limit the scope of protection of the present application.

[0039] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by technicians in the technical field to which this application belongs; the terms used herein are only for the purpose of describing specific embodiments and are not intended to limit this application.

[0040] In the description of the embodiments of the present application, the technical terms "first", "second", etc. are only used to distinguish different objects, and cannot be understood as indicating or implying relative importance or implicitly indicating the number, specific order or primary and secondary relationship of the indicated technical features. In the description of the embodiments of the present application, the meaning of "multiple" is two or more, unless otherwise clearly and specifically defined.

[0041] See also Figure 1 The flowchart of an image recognition model training method provided by an embodiment of the present application is shown. The image recognition model training method may include the following steps:

[0042] S101, using a model obtained by training based on the first type of training images as a teacher model;

[0043] S102, using a target distillation algorithm to transfer the knowledge of the teacher model to the student model; wherein the target distillation algorithm includes: in the distillation process, based on the similarity between the teacher feature representation and the student feature representation, updating the distillation center; the teacher feature representation includes the feature representation of the training image by the teacher model, and the student feature representation includes the feature representation of the training image by the student model;

[0044] S103: Based on the second type of target image and the target recognition result of the target image, the student model is adjusted and trained to obtain a trained image recognition model.

[0045] In step S101, the first type of training image may be an RGB image, a depth image, a point cloud image, or an IR image, etc. The training image may specifically be a public RGB face image dataset, a bridge crack depth dataset, etc. The training image may be adjusted according to the actual use of the model, and this application does not specifically limit this.

[0046] Among them, in step S102, the traditional distillation algorithm usually takes the teacher model as the distillation center, that is, the class center predicted by the teacher model is taken as the true value of the class center learned by the student model. Based on the traditional distillation algorithm, the distillation center is fixed during the pre-training process. Based on the target distillation algorithm, during the distillation process, the distillation center will be updated based on the similarity between the teacher feature representation and the student feature representation. Specifically, in the early stage of pre-training, the representation ability of the student model is poor, and the distillation center can be closer to the class center of the teacher model; with the iteration of the training process, the student model will learn more complex feature representations, and the distillation center can slowly approach the class center of the student model. The teacher features and the student features can be vectorized, and based on the cosine similarity between the teacher feature representation and the student feature representation Calculate the similarity λ between the teacher feature representation and the student feature representation, and based on D t,k =λ*D t,k-1 +(λ-1)*ω t,k Update the distillation center; among them, f t represents the teacher feature representation, f s represents the student characteristics, Indicates that the cosine similarity λ is truncated in the range [0,1], D t,k-1 represents the updated value of the distillation center at step (k-1), ω t,k represents the class center predicted by the teacher model at step k, D t,k Represents the updated value of the distillation center at step k.

[0047] Wherein, in step S103, the second type is different from the first type, and the second type of image may be an RGB image, a depth image, a point cloud image, or an IR image. Taking the first type of image as an RGB image as an example, the second type of image may be an IR image or a depth image. Taking the second type of image as an IR image as an example, the target image may be an IR image including facial features, fingerprint features, bridge crack features, or animal body size features. The target image may be adjusted according to the actual use of the model, and this application does not specifically limit this.

[0048] It can be seen that the image recognition model training method provided by the embodiment of the present application can reduce the dependence of the student model on the teacher model by dynamically updating the distillation center based on the similarity between the teacher feature representation and the student feature representation during the distillation process, improve the generalization ability of the student model obtained by pre-training, and thus improve the generalization ability of the image recognition model obtained by adjusting the student model. In addition, by utilizing the similarity between the teacher feature representation and the student feature representation to dynamically update the distillation center, the problem of slow convergence during the training process caused by the structural differences between the teacher model and the student model can also be alleviated, thereby improving the convergence speed of the model and the efficiency of model training.

[0049] Please refer to Figure 2 , Figure 2 A flow chart of a knowledge transfer method provided in an embodiment of the present application. Figure 2 As shown, in some optional embodiments, S102, using a target distillation algorithm to transfer the knowledge of the teacher model to the student model, including: S1021, dividing the training images into multiple image sets; wherein each of the image sets includes category images of multiple categories; S1022, based on the similarity between the teacher category feature representation of the category image by the teacher model and the student category feature representation of the category image by the student model, updating the distillation class center of the category image; S1023, based on the updated distillation class center, calculating the distillation loss value of the student model; S1024, adjusting the model parameters of the student model according to the distillation loss value to transfer the knowledge of the teacher model to the student model.

[0050] The training images can be evenly divided into multiple image sets according to the total number of image categories. The number of image sets can be adjusted based on the number of available graphics processing units (GPUs), and the number of image sets must be less than or equal to the number of available GPUs. The cosine similarity between the teacher category feature representation and the student category feature representation can be used to calculate the number of image sets. Calculate the similarity λ between the teacher category feature representation and the student category feature representation corresponding to the i-th category imagei , and based on Update distillation center; among them, represents the teacher category feature representation corresponding to the i-th category image, represents the student category feature representation corresponding to the i-th category image, Represents the cosine similarity λ i is truncated in the range [0,1]. represents the updated value of the (k-1)th step distillation class center corresponding to the i-th category image, represents the class center predicted by the k-th step teacher model corresponding to the i-th category image, Represents the updated value of the k-th step distillation class center corresponding to the i-th category image. By dividing the training images into multiple image sets, the multiple image sets can be assigned to different GPUs for processing in combination with the actual number of available GPUs, so that each GPU only needs to process part of the training images, thereby alleviating the GPU pressure during the training process; it can also improve the speed and efficiency of model distillation. The distillation class center of the category image is updated through the similarity between the teacher category feature representation of the category image by the teacher model and the student category feature representation of the category image by the student model, and the distillation loss value used to adjust the model parameters is calculated based on the updated distillation class center; this reduces the dependence of the student model on the teacher model and improves the generalization ability of the student model obtained by pre-training.

[0051] In some optional embodiments, S1023, based on the updated distillation class center, calculating the distillation loss value of the student model, including: sampling the updated distillation class center to obtain a sampling class center; calculating the distillation loss value of the student model according to the sampling class center and the student category feature representation of the category image corresponding to the sampling class center.

[0052] The updated distilled class center can be sampled based on a preset sampling rate to obtain the sampled class center. The preset sampling rate can be 0.1, 0.2, or other reasonable values. The data on all available GPUs and the feature representations output by the teacher model and the student model can be strung together, that is, based on The data corresponding to the 1st to the Mth GPUs are grouped, and then the feature representations and data labels of the corresponding categories are assigned to them according to the predetermined class center range of each GPU. Then, the updated distilled class centers are sampled. For example, the types of category images included in the image set are [0, 45K), and the preset sampling rate is 0.1. The number of sampled class centers corresponding to the image set is 0.1*45K=4.5K; among them, the data whose labels fall within the range of [0, 45K) are taken as positive samples. If the number of positive samples is less than 4.5K, samples whose labels are not in the range of [0, 45K) can be randomly selected as negative samples to ensure that the number of sampled distilled class centers is 4.5K. The sampled distilled class centers The student category feature representation f of the assigned category image with the student model s′ It will participate in the subsequent calculation of the distillation loss value. Taking the trained image recognition model as a face recognition model as an example, the distillation loss value can be calculated based on ArcFace Loss, that is, based on Calculate the distillation loss value L D ; Where C represents the total number of image categories corresponding to the training image, i represents the i-th category image, α represents the vector scaling parameter, m represents the hyperparameter in ArcFace Loss, m is used to increase the distance between classes, yi represents the class center prediction value corresponding to the i-th category image, k represents other category images different from the i-th category image, Represents the student category feature representation of the category image corresponding to the i-th category image in the sampling category center, It represents the distillation class center corresponding to the i-th class image in the sampled class center. The distillation loss value can also be calculated based on Softmax Loss. Since the principle of Softmax Loss is to push the training samples to the positive class center and away from the negative class center, it is not necessary to use all negative samples. The selection of negative samples can be approximated by random partial sampling. In a large number of iterations, all positive samples will be sampled. Therefore, there is no need to consider the related issues of positive sample sampling.

[0053] Please refer to Figure 3 , Figure 3 A flowchart of a model pre-training process provided in an embodiment of the present application. Figure 3 Specifically, the case where the training image is an RGB face image dataset is shown. Figure 3As shown, based on the image recognition model training method provided by the present application, during the pre-training process, a large-scale public RGB face image dataset can be evenly divided into M image sets according to image categories (taking the total number of categories C and the number of available GPUs greater than or equal to M as an example), and correspondingly allocated to the M available GPUs to achieve model-based parallel data allocation, so that each GPU only needs to process part of the training images, thereby alleviating the GPU pressure during the training process. Accordingly, the jth GPU (corresponding to Figure 3 The range of the image set allocated on GPU-j in Based on the processing results of the assigned image data by M GPUs, the distillation class centers of different categories of images are dynamically updated, and the sampled class centers are obtained by sampling the updated distillation class centers; the distillation loss value of the student model is calculated based on the sampled class centers and the student category feature representation of the category image corresponding to the sampled class centers; the model parameters of the student model are adjusted according to the distillation loss value to transfer the knowledge of the teacher model to the student model, completing the model pre-training process. Based on the above model pre-training method, the performance of the lightweight model can be improved without using technical means such as model quantization, compression and pruning that reduce the accuracy of the model.

[0054] Please refer to Figure 4 , Figure 4 The following is a flow chart of a method for adjusting and training a student model provided in an embodiment of the present application. Figure 4 As shown, in some optional embodiments, S103, based on the second type of target image and the target recognition result of the target image, the student model is adjusted and trained to obtain a trained image recognition model, including: S1031, based on the student model, obtaining the target student feature representation of the target image; S1032, constructing a negative sample pair according to the target student feature representation of the target image and the target recognition result of the target image; wherein the negative sample pair includes two target student feature representations with different target recognition results; S1033, calculating a first adjusted loss value of the student model according to the negative similarity corresponding to the negative sample pair; S1034, adjusting the model parameters of the student model based on the first adjusted loss value to obtain the trained image recognition model.

[0055] Among them, taking the image recognition model as a face recognition model as an example, a negative sample pair can be constructed by simulating the situation of face verification, and by showing the difference relationship between the negative sample pairs, the model can take into account the problem of false recognition rate. In the fine-tuning stage, a negative sample pair is constructed based on the target student feature representation of the target image and the target recognition result of the target image; since the negative similarity between the negative sample pairs can reflect the probability of the image being misrecognized, the model parameters of the student model are adjusted based on the first adjustment loss value calculated based on the negative similarity corresponding to the negative sample pair, which can reduce the probability of the image being misrecognized. The first adjustment loss value may include the false recognition probability that the negative similarity corresponding to the negative sample pair is higher than the negative similarity threshold. The trained image recognition model can be obtained when the false recognition probability is less than or equal to the preset false recognition rate. For a biometric recognition system, such as a face recognition system, the model parameters of the student model are adjusted based on the first adjustment loss value calculated based on the negative similarity corresponding to the negative sample pair, and the image recognition model obtained can reduce the system false recognition rate, that is, improve the system's recognition accuracy for unauthorized users, thereby improving the security of the face recognition system.

[0056] In some optional embodiments, the target recognition result includes the image category of the target image; S103, based on the second type of target image and the target recognition result of the target image, adjusting and training the student model to obtain a trained image recognition model, also includes: calculating a second adjusted loss value based on the distillation center corresponding to the image category of the target image and the target student feature representation of the target image by the student model; S1034, adjusting the model parameters of the student model based on the first adjusted loss value to obtain the trained image recognition model, including: calculating an adjusted loss value based on the first adjusted loss value and the second adjusted loss value; adjusting the model parameters of the student model based on the adjusted loss value to obtain the trained image recognition model.

[0057] Among them, the second adjusted loss value can be calculated based on the recognition accuracy or recall rate of the model. Taking the image recognition model as a face recognition model as an example, the second adjusted loss value can be calculated based on methods such as ArcFace Loss or Softmax Loss. The adjusted loss value calculated based on the first adjusted loss value and the second adjusted loss value can more comprehensively reflect the degree of fit between the model and the second type of target image; so as to further improve the recognition performance of the image recognition model for the second type of image.

[0058] In some optional embodiments, the calculating the adjusted loss value based on the first adjusted loss value and the second adjusted loss value includes: calculating the adjusted loss value based on the first adjusted loss value, the second adjusted loss value and L T =Lcls +β*L far , calculate the adjusted loss value L T Among them, L far represents the first adjusted loss value, L cls represents the second adjusted loss value, and β represents the weight coefficient of the first adjusted loss value.

[0059] Among them, when the image recognition model is a biometric model, such as a face recognition model or a fingerprint recognition model, the value of β can be 50, 100 or 200, etc., to more strictly control the model's error rate, so as to further reduce the model's error rate and improve the application security of the biometric model. When the image recognition model is a non-biometric model, such as a bridge crack recognition model or a garbage classification recognition model, the value of β can be 1, 2 or other smaller values, so as to more comprehensively consider the model's error rate, recognition accuracy and recognition precision, etc., and improve the overall recognition performance of the non-biometric system. Therefore, through L T =L cls +β*L far , calculate the adjusted loss value L T ; For biometric systems, such as face recognition systems, the weight coefficient β of the first adjustment loss value can be increased to more strictly control the model's error rate, so as to further reduce the system error rate of the face recognition system and ensure the security of the face recognition system; for non-biometric systems, the weight coefficient β of the first adjustment loss value can be reduced to more comprehensively consider the model's error rate, recognition accuracy and recognition precision, so as to improve the overall recognition performance of the non-biometric system.

[0060] In some optional embodiments, the calculating the first adjusted loss value of the student model according to the negative similarity corresponding to the negative sample pair includes: calculating the first adjusted loss value of the student model according to the negative similarity corresponding to the negative sample pair and Calculate the first adjusted loss value L far ; Wherein, N represents the number of negative sample pairs, si represents the negative similarity corresponding to the i-th negative sample pair, and T represents the preset similarity threshold.

[0061] Among them, the error rate N FA Indicates the number of times the negative similarity corresponding to the negative sample pair is greater than the preset similarity threshold T. The preset similarity threshold T required for model training can be calculated under a fixed false positive rate FAR (which can be 1e-6, 1e-5 or other reasonable values). Since the calculation method of FAR is not differentiable, this application is based on the sigmoid function to determine To calculate the first adjusted loss value L far .

[0062] Please refer to Figure 5 , Figure 5 A flowchart of a model fine-tuning process provided in an embodiment of the present application. Figure 5 Specifically, the case where the target image is an IR face image dataset is shown. Figure 5 As shown, based on the image recognition model training method provided by the present application, in the fine-tuning process, a small-scale IR face image dataset can be first input into the student model obtained by pre-training, and the target student feature representation f output by the student model can be used as the basis for the fine-tuning process. sm And the target recognition result of the target image, construct a negative sample pair; the negative sample pair may include two target student feature representations with different target recognition results. Based on the negative sample pair, the first adjusted loss value L can be calculated far . And based on the first adjusted loss value L far and the second adjusted loss value L cls , calculate the adjusted loss value L T ; Based on the adjusted loss value L T Adjust the model parameters of the student model to obtain a trained image recognition model. Figure 5 As shown, the construction process of negative sample pairs can be realized by maintaining the sample matrix Q, and the size of the matrix Q is C q *K q *D q , C q , K q , D q They represent the number of categories of training samples, including the storage capacity of the memory and the size of the feature f output by the model. q , K q , D q Taking 360232, 10 and 512 as examples, we can regard them as 360232 queues, each with a capacity of 10*512, which means that each category of images in the matrix Q can cache up to 10 feature samples, and the size of each feature sample is 512 dimensions (model output dimension). In this way, in each update process, a negative sample pair can be constructed by using the feature f corresponding to the current category image and the feature sample cached in Q. The queue update of the sample matrix Q can be implemented according to the first-in-first-out principle.

[0063] Taking the teacher model as Adaface-101 pre-trained in WebFace42M (Adaface model with backbone ResNet101) and the student model as the classic lightweight model MobileNetv2 as examples, the model fine-tuning method provided in this application was verified based on the two RGB public datasets IJBB and IJBC, and the infrared face dataset (PSF-A1) collected by the IR sensor. The model training effect was evaluated using two indicators, accuracy (ACC) and acceptance rate at a specified error rate (TAR@FAR=1e-6). The evaluation results are shown in Table 1.

[0064] Table 1

[0065]

[0066] It can be seen from Table 1 that the accuracy (ACC) and the acceptance rate at a specified error rate (TAR@FAR=1e-6) of the image recognition model trained based on the fine-tuning method provided in the present application are both higher; that is, the recognition performance of the image recognition model trained based on the fine-tuning method provided in the present application is better than the recognition performance of the image recognition model trained based on the traditional fine-tuning method.

[0067] An embodiment of the present application also provides a computer program product, including a computer program / instruction, which, when executed by a processor, implements the image recognition model training method as described in any one of the first aspects above.

[0068] Please refer to Figure 6 , Figure 6 The present invention provides a schematic diagram of the structure of an electronic device 300 provided in an embodiment of the present application. The electronic device 300 includes: a memory 302 and a processor 301; the memory 302 stores a computer program executable by the processor 301, and when the computer program is executed by the processor 301, the image recognition model training method described in any one of the first aspects is executed.

[0069] The memory 302 and the processor 301 may be interconnected and communicate with each other via a communication bus 303 and / or other forms of connection mechanisms (not shown). The memory 302 stores a computer program executable by the processor 301, and when the computer program is executed by the processor 301, the image recognition model training method described in the first aspect above is executed.

[0070] An embodiment of the present application also provides a computer-readable storage medium, on which computer program instructions are stored. When the computer program instructions are executed by the processor 301, the image recognition model training method described in the first aspect above is executed.

[0071] Among them, the storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable red-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk or optical disk.

[0072] In several embodiments provided by the embodiments of the present application, it should be understood that the disclosed device / system and method can also be implemented in other ways. The device embodiments described above are merely schematic. For example, the flowcharts and block diagrams in the accompanying drawings show the possible architecture, functions and operations of the devices, methods and computer program products according to the multiple embodiments of the embodiments of the present application. In this regard, each box in the flowchart or block diagram can represent a module, a program segment or a part of the code, and a part of the module, program segment or code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than the order marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of boxes in the block diagram and / or flowchart can be implemented with a dedicated hardware-based system that performs a specified function or action, or can be implemented with a combination of dedicated hardware and computer instructions.

[0073] In addition, the functional modules in each embodiment of the present application can be integrated together to form an independent part, or each module can exist separately, or two or more modules can be integrated to form an independent part.

[0074] The above description is only an optional implementation manner of the embodiments of the present application, but the protection scope of the embodiments of the present application is not limited thereto. Any technician familiar with the technical field can easily think of changes or replacements within the technical scope disclosed in the embodiments of the present application, which should be covered within the protection scope of the embodiments of the present application.

Claims

1. A method for training an image recognition model, characterized in that: The method comprises: Using a model obtained by training based on the first type of training images as a teacher model; A target distillation algorithm is used to transfer the knowledge of the teacher model to the student model; wherein the target distillation algorithm includes: in the distillation process, based on the similarity between the teacher feature representation and the student feature representation, updating the distillation center; the teacher feature representation includes the feature representation of the training image by the teacher model, and the student feature representation includes the feature representation of the training image by the student model; Based on the second type of target image and the target recognition result of the target image, the student model is adjusted and trained to obtain a trained image recognition model.

2. The method according to claim 1, characterized in that The target distillation algorithm is used to transfer the knowledge of the teacher model to the student model, including: Dividing the training images into a plurality of image sets; wherein each of the image sets includes category images of multiple categories; Based on the similarity between the teacher category feature representation of the category image by the teacher model and the student category feature representation of the category image by the student model, updating the distilled class center of the category image; Calculating the distillation loss value of the student model based on the updated distillation class center; A model parameter of the student model is adjusted according to the distillation loss value to transfer the knowledge of the teacher model to the student model.

3. The method according to claim 2, characterized in that The calculating the distillation loss value of the student model based on the updated distillation class center includes: Sampling the updated distillation class center to obtain a sampled class center; The distillation loss value of the student model is calculated according to the sampled class center and the student class feature representation of the class image corresponding to the sampled class center.

4. The method according to claim 1, characterized in that: The step of adjusting and training the student model based on the second type of target image and the target recognition result of the target image to obtain a trained image recognition model includes: Based on the student model, obtaining a target student feature representation of the target image; Constructing a negative sample pair according to the target student feature representation of the target image and the target recognition result of the target image; wherein the negative sample pair includes two target student feature representations with different target recognition results; Calculating a first adjusted loss value of the student model according to the negative similarity corresponding to the negative sample pair; The model parameters of the student model are adjusted based on the first adjustment loss value to obtain the trained image recognition model.

5. The method according to claim 4, characterized in that in, The target recognition result includes the image category of the target image; The adjusting and training of the student model based on the second type of target image and the target recognition result of the target image to obtain a trained image recognition model also includes: Calculating a second adjusted loss value according to a distillation center corresponding to the image category of the target image and a target student feature representation of the target image by the student model; The step of adjusting the model parameters of the student model based on the first adjustment loss value to obtain the trained image recognition model includes: calculating an adjusted loss value based on the first adjusted loss value and the second adjusted loss value; The model parameters of the student model are adjusted based on the adjusted loss value to obtain the trained image recognition model.

6. The method according to claim 5, characterized in that The calculating the adjusted loss value based on the first adjusted loss value and the second adjusted loss value includes: According to the first adjusted loss value, the second adjusted loss value and L T =L cls +β*L far , calculate the adjusted loss value L T Among them, L far represents the first adjusted loss value, L cls represents the second adjusted loss value, and β represents the weight coefficient of the first adjusted loss value.

7. The method according to claim 4, characterized in that The step of calculating the first adjusted loss value of the student model according to the negative similarity corresponding to the negative sample pair includes: According to the negative similarity corresponding to the negative sample pair and Calculate the first adjusted loss value L far ; Wherein, N represents the number of negative sample pairs, si represents the negative similarity corresponding to the i-th negative sample pair, and T represents the preset similarity threshold.

8. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instructions are executed by a processor, the method according to any one of claims 1 to 7 is implemented.

9. An electronic device, characterized in that: The electronic device comprises: Memory; processor; The memory stores a computer program executable by the processor, and when the computer program is executed by the processor, the method according to any one of claims 1 to 7 is performed.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer program instructions, and when the computer program instructions are executed by a processor, the method according to any one of claims 1 to 7 is executed.