Model optimization method and device and storage medium

By determining the similarity thresholds of false positive and missed samples in the target recognition model, the model loss function is optimized, and the recall difference in training and testing is solved, and the performance of the model under specific false positive rates is improved.

CN120338037APending Publication Date: 2025-07-18ZHEJIANG DAHUA TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510293184.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-11
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

There are differences in false positive rate and recall rate during training and testing of existing target recognition models, resulting in low recall rates during testing.

Method used

Determine the current similarity threshold of the initial recognition model by entering the preset test set, determine the target loss value based on the false positive samples, missed samples and preset anchor samples, adjust the initial recognition model, and optimize the model to narrow the differences in the training and testing process.

Benefits of technology

This improves the recall rate of the model under specific false positive rates, reduces the risk of model overfitting, and achieves better consistency in the training and testing process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120338037A_ABST
    Figure CN120338037A_ABST
Patent Text Reader

Abstract

The invention discloses a model optimization method and device and a storage medium, and the method comprises the steps: inputting a preset test set into an initial recognition model, and obtaining a current similarity threshold value of the initial recognition model under a false alarm rate corresponding to the preset test set; according to a current similarity threshold value and similarity distribution of training samples in a training set of the initial recognition model, determining a false report sample and a missing report sample from the training set; determining a target loss value of the initial recognition model according to the misinformation sample, the misinformation sample, the preset anchor point sample and the model type of the initial recognition model; and adjusting the initial recognition model according to the target loss value to obtain a target recognition model. According to the scheme, the recognition model can be optimized, and the recognition effect of the recognition model is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of neural networks, and in particular, to a model optimization method, device, and storage medium. Background Art

[0002] In the current model training process, especially in the training process of object recognition models, most of them adopt classification training methods, and optimize the classification loss by adding different margins within and between classes, so that the model can finally have the ability to represent features with strong discrimination.

[0003] In the testing process, usually, feature extraction is first performed on two test samples, and then the feature similarity is calculated by calculating the cosine similarity or Euclidean distance, etc., and then compared with a preset similarity threshold to determine whether the two test samples belong to the same target object.

[0004] Generally, the similarity threshold is comprehensively evaluated and set by statistically calculating parameters such as false positive rate and recall rate under different similarity thresholds in a test scenario similar to the actual application scenario. It can be seen that the differences between the model training process and the model testing process are relatively large. Even if the model converges well in the training stage, there may be a situation where the recall rate is very low at a certain false positive rate in the testing stage. Summary of the Invention

[0005] This application provides at least one model optimization method, device, device, and computer-readable storage medium.

[0006] The first aspect of this application provides a model optimization method, including: inputting a preset test set into an initial recognition model to obtain a current similarity threshold of the initial recognition model at a false positive rate corresponding to the preset test set; determining false positive samples and missed detection samples from the training set according to the current similarity threshold and the similarity distribution of training samples in the training set of the initial recognition model; determining a target loss value of the initial recognition model according to the false positive samples, the missed detection samples, preset anchor samples, and the model type of the initial recognition model; and adjusting the initial recognition model according to the target loss value to obtain a target recognition model.

[0007] In one embodiment, the training set includes positive pairs of samples and negative pairs of samples. Determining false positive samples and false negative samples from the training set according to the current similarity threshold and the similarity distribution of training samples in the training set of the initial recognition model includes: obtaining the feature similarities of each positive pair of samples and the feature similarities of each negative pair of samples; determining the positive pair of samples with feature similarities less than the current similarity threshold as the false negative samples, and determining the negative pair of samples with feature similarities greater than the current similarity threshold as the false positive samples.

[0008] In one embodiment, determining the positive pair of samples with feature similarities less than the current similarity threshold as the false negative samples, and determining the negative pair of samples with feature similarities greater than the current similarity threshold as the false positive samples includes: determining the target similarity threshold range according to the current similarity threshold and a preset range parameter; determining the positive pair of samples with feature similarities within the target similarity range and less than the current similarity threshold as the false negative samples, and determining the negative pair of samples with feature similarities within the target similarity range and greater than the current similarity threshold as the false positive samples.

[0009] In one embodiment, the model type includes a non-student model. Determining the target loss value of the initial recognition model according to the false positive samples, the false negative samples, preset anchor samples, and the model type of the initial recognition model includes: in response to the initial recognition model being the non-student model, constructing a triplet according to the false positive samples, the false negative samples, and the preset anchor samples, and determining the triplet loss; determining the target loss according to the triplet loss and the classification loss of the initial recognition model.

[0010] In one embodiment, constructing a triplet according to the false positive samples, the false negative samples, and the preset anchor samples, and determining the triplet loss includes: obtaining the similarity improvement parameter corresponding to the false negative samples and the similarity reduction parameter of the false positive samples; determining a sample boundary parameter according to the similarity improvement parameter and the similarity reduction parameter; determining the triplet loss according to the feature similarity between the false negative samples and the anchor samples, the feature similarity between the false positive samples and the preset anchor samples, and the sample boundary parameter.

[0011] In one embodiment, the preset test set includes a preset negative pair set, and the model type includes a student model. Before determining false positive samples and false negative samples from the training set according to the current similarity threshold and the similarity distribution of training samples in the training set of the initial recognition model, the method further includes: in response to the initial recognition model being a student model, obtaining a teacher model corresponding to the student model; determining a teacher similarity threshold of the teacher model at the preset false positive rate according to the false positive rate corresponding to the preset negative pair set; and determining a target similarity range of the student model according to the teacher similarity threshold and the preset range parameter.

[0012] In one embodiment, after determining the target similarity range of the student model according to the teacher similarity threshold and the preset range parameter, the method further includes: obtaining false positive samples and false negative samples in the training set that are within the target similarity range; constructing a correct acceptance rate loss according to the false negative samples, and constructing a false acceptance rate loss according to the false positive samples; and determining the target loss according to the consistency loss between the student model and the teacher model, the correct acceptance rate loss, and the false acceptance rate loss.

[0013] In one embodiment, the method further includes: obtaining negative pair training samples with a feature similarity greater than a preset similarity threshold and positive pair training samples with a feature similarity less than the preset similarity threshold; and constructing a training set according to the negative pair training samples and the positive pair training samples, where the training set is used to perform model training on the initial recognition model.

[0014] A second aspect of the present application provides a model optimization device, including: a test module, configured to input a preset test set into an initial recognition model to obtain a current similarity threshold of the initial recognition model at the false positive rate corresponding to the preset test set; a sample determination module, configured to determine false positive samples and false negative samples from the training set according to the current similarity threshold and the similarity distribution of training samples in the training set of the initial recognition model; a loss determination module, configured to determine a target loss value of the initial recognition model according to the false positive samples, the false negative samples, preset anchor samples, and the model type of the initial recognition model; and a model adjustment module, configured to perform an adjustment process on the initial recognition model according to the target loss value to obtain a target recognition model.

[0015] A third aspect of the present application provides an electronic device, including a memory and a processor, where the processor is configured to execute program instructions stored in the memory to implement the above model optimization method.

[0016] A fourth aspect of the present application provides a computer-readable storage medium, on which program instructions are stored, and when the program instructions are executed by a processor, the above model optimization method is implemented.

[0017] In the above solution, by inputting a preset test set into the initial recognition model, the current similarity threshold of the initial recognition model at the false positive rate corresponding to the preset test set is obtained. If all samples are optimized during training, it is easy to cause model overfitting. Therefore, the false positive samples and missed detection samples that need to be optimized can be determined from the training set according to the current similarity threshold and the similarity distribution of the training samples in the training set of the initial recognition model, which reduces the calculation amount, is easy for model training, and has a better optimization effect. The target loss value of the initial recognition model is determined according to the false positive samples, missed detection samples, preset anchor samples, and the model type of the initial recognition model, so that different loss values can be determined for different types of models for adaptive optimization. Then, the initial recognition model is adjusted according to the target loss value to obtain the target recognition model, realizing model optimization.

[0018] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and do not limit this application. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] The drawings here are incorporated into the specification and constitute a part of this specification. These drawings show embodiments consistent with this application and are used together with the specification to explain the technical solutions of this application.

[0020] Figure 1 It is a schematic flowchart of an exemplary embodiment of the model optimization method of this application;

[0021] Figure 2 It is a schematic diagram for comparing exemplary model performance curves in the model optimization method of this application;

[0022] Figure 3 It is a schematic diagram of the positive-negative pair similarity distribution of an exemplary recognition model on a test set in the model optimization method of this application;

[0023] Figure 4 It is an exemplary optimization schematic diagram in the model optimization method of this application;

[0024] Figure 5 It is a schematic diagram of the positive-negative sample similarity distribution of an exemplary initial recognition model on a training set in the model optimization method of this application;

[0025] Figure 6 It is a block diagram of an exemplary model optimization device shown in an exemplary embodiment of this application;

[0026] Figure 7 It is a schematic structural diagram of an embodiment of an electronic device of this application;

[0027] Figure 8It is a schematic structural diagram of an embodiment of the computer-readable storage medium of the present application. Detailed implementation manners

[0028] The following will combine the accompanying drawings of the specification to elaborate in detail on the solutions of the embodiments of the present application.

[0029] In the following description, specific details such as specific system structures, interfaces, and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the present application.

[0030] The term "and / or" in this article is merely a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this article generally represents an "or" relationship between the front and rear associated objects. In addition, "multiple" in this article means two or more than two. In addition, the term "at least one" in this article represents any one of multiple types or any combination of at least two of multiple types. For example, including at least one of A, B, and C can represent including any one or more elements selected from the set composed of A, B, and C.

[0031] For ease of understanding, the application scenarios involved in the present application are now illustrated by examples. In the current model training process, especially in the training process of the target recognition model, most of them adopt the classification training method, and optimize the classification loss by adding different margins within and between classes, so that the model can finally have the ability to represent features with strong discrimination.

[0032] In the testing process, usually, feature extraction is first performed on two test samples, and then the feature similarity is calculated by calculating the cosine similarity or Euclidean distance, etc., and then compared with a pre-set similarity threshold to determine whether the two test samples belong to the same target object. This method has been widely used in scenarios such as 1:1 verification of target objects, 1:N retrieval of target objects, and clustering.

[0033] Generally, the similarity threshold is comprehensively evaluated and set by statistically parameters such as false positive rate and recall rate under different similarity thresholds in a test scenario similar to the actual application scenario. It can be seen that the differences between the model training process and the model testing process are relatively large. Even if the model converges well in the training stage, there may be a situation where the recall rate is very low at a certain false positive rate in the testing stage.

[0034] For example, the commonly used target recognition training method at present is to first perform classification training on N types of training data. When classifying, a classification loss function is used to obtain better results, and all types of data are basically treated equally. However, in actual testing, the performance of the neural network (recognition model) is judged by many indicators such as Rank1 (first hit), FAR (False Accept Rate), and TAR (True Accept Rate), and these indicators are often not specifically optimized during training. Therefore, in many cases, although the training results of the neural network seem to converge well, many problems will occur in actual testing. Therefore, it is very necessary to make the training process and testing process of the recognition model consistent to a certain extent.

[0035] In actual applications, the target recognition model usually selects a false alarm rate (such as 1e-4) as a benchmark, and on an existing test set (the data distribution is as close as possible to the actual application scenario), finds the similarity threshold at the selected false alarm rate to be used as the criterion for distinguishing positive pairs (sample pairs composed of positive samples and anchor samples) and negative pairs (sample pairs composed of negative samples and anchor samples) in the actual application scenario. Therefore, this application can optimize the model according to the recall index under the similarity threshold.

[0036] Please refer to Figure 1 , Figure 1 which is a schematic flowchart of an exemplary embodiment of the model optimization method of this application. Specifically, it may include the following steps:

[0037] Step S110, input the preset test set into the initial recognition model to obtain the current similarity threshold of the initial recognition model at the false alarm rate corresponding to the preset test set.

[0038] Among them, the preset test set refers to a pre-set test sample set, which may include positive samples, negative samples, and anchor samples. The corresponding positive sample and anchor sample can be called a positive pair (positive pair sample), and the corresponding negative sample and anchor sample can be called a negative pair (negative pair sample). The model task of the recognition model is usually to recognize the positive sample as similar or identical to the anchor sample (for example, the similarity is greater than or equal to the similarity threshold), and to recognize the negative sample as dissimilar or different from the anchor sample (for example, the similarity is less than the similarity threshold). This is the common principle of the recognition model and will not be elaborated here. Among them, there can be one or more anchor samples in this application, there can be one or more positive samples corresponding to one anchor sample, and there can be one or more negative samples corresponding to one anchor sample, which is not limited here.

[0039] Exemplarily, it can be referred to as Figure 2 shown in Figure 2It is a schematic diagram for comparing the model performance curves in the model optimization method of this application. Figure 2 Exemplarily, three curves are provided, corresponding to the ROC (Receiver Operating Characteristic Curve) test metrics of three recognition models, namely Model1 (recognition model 1), Model2 (recognition model 2), and Model3 (recognition model 3), under a certain same test set. It can be seen that for the TAR under different FARs, the performances of the three models have their own advantages and disadvantages. For example, when the FAR is 1e-8, the recall rate of model 1 is the best, followed by model 2, and the worst is model 3; but when the FAR is 1e-4, the recall rate of model 3 is the best, followed by model 1, and the worst is model 2. Therefore, it is necessary to pay attention to the recall effect of the recognition model at a specific false positive rate during the training of the recognition model.

[0040] Specifically, the false positive rate refers to the proportion of samples that are actually negative but are wrongly predicted as positive among the negative samples. Therefore, a negative pair set (negative pair test set) can be pre-constructed. The sample type and similarity distribution of this negative pair set need to simulate the actual application scenario as much as possible. Thus, a specific false positive rate is provided through this negative pair set to determine the similarity threshold of the initial recognition model at the specific false positive rate, and this similarity threshold is used as a reference for optimizing false positive samples and missed detection samples during training. And the positive samples do not affect the false positive rate parameter, so this application does not limit the positive samples (positive pair set) in the test set.

[0041] For example, a test set with a false positive rate of 1e-4 is pre-constructed and input into the initial recognition model to obtain the current similarity threshold of the initial recognition model at the preset false positive rate. Exemplarily, it can be referred to as Figure 3 as shown Figure 3 It is a schematic diagram of the positive and negative pair similarity distribution of an exemplary recognition model in the test set in the model optimization method of this application.

[0042] th is the similarity threshold of the initial recognition model at the false positive rate corresponding to the preset negative pair set. Generally, Figure 3 the samples on the left side of the similarity threshold in it are judged to be different from the corresponding anchor point samples, and the samples on the right side of the similarity threshold are judged to be the same as the corresponding anchor point samples.

[0043] Step S120, determine false positive samples and missed detection samples from the training set according to the current similarity threshold and the similarity distribution of the training samples in the training set of the initial recognition model.

[0044] Illustrated in combination with the foregoing embodiments, in Figure 3Under the shown similarity threshold, there will still be high-score false positives (judging negative samples (samples different from the anchor samples) as the same as the anchor samples) and low-score false negatives (judging positive samples (samples the same as the anchor samples) as different from the anchor samples) on both the left and right sides. Therefore, for the training process of the recognition model, this application optimizes the model for such false positive and false negative situations.

[0045] Exemplarily, during the model training process, after inputting the training set into the model, the similarity distribution of the training samples in the training set can be obtained in the same way. Therefore, according to the current similarity threshold determined from the test set, false positive samples, false negative samples, and the corresponding anchor samples can be determined from the training samples of the training set. Among them, false positive samples refer to negative samples in the training set whose feature similarity with the corresponding anchor sample is greater than or equal to the current similarity threshold, and false negative samples refer to positive samples in the training set whose feature similarity with the anchor sample is less than the current similarity threshold. Or, false positive samples refer to negative samples whose feature similarity with the anchor sample is greater than the current similarity threshold, and false negative samples refer to positive samples whose feature similarity with the anchor sample is less than or equal to the current similarity threshold. It specifically depends on the actual application scenario and is not limited here.

[0046] Optionally, for the training set, it will be very difficult to optimize all high-score false positives and low-score false negatives during training, and it is very easy to cause overfitting of the model. Therefore, in this application, a certain range of training samples can be optimized according to the obtained similarity threshold. As Figure 4 shown, Figure 4 is an exemplary optimization schematic diagram in the model optimization method of this application. If triplets are formed for the false positive samples and false negative samples near the similarity threshold (the target similarity range) for targeted optimization, so that the feature similarity between the false positive samples and the anchor samples is slightly lower than the current similarity threshold, and the feature similarity between the false negative samples and the anchor samples is slightly higher than the current similarity threshold, then the overall performance of the recognition model will be further improved.

[0047] Step S130, determine the target loss value of the initial recognition model according to the false positive samples, false negative samples, preset anchor samples, and the model type of the initial recognition model.

[0048] Among them, the preset anchor samples refer to the anchor samples that can be determined from the training set in the same way according to the method of the foregoing embodiments. There can be one or more preset anchor samples, which are not limited here.

[0049] It should also be noted that different model types may also represent different model training methods.

[0050] For example, the model types can be divided into student models and non-student models. Among them, student models usually involve the knowledge distillation technology of deep learning, and the learned spatial relationships need to be transmitted to the student models through their corresponding teacher models (such as probability distributions, attention weights, or compressed intermediate layer representations, etc.); while non-student models (other types of models except student models) do not need to use knowledge distillation technology for learning and training. For the specific explanation of student models, reference can be made to the prior art and will not be elaborated here.

[0051] Therefore, the present application provides different loss calculation methods for different types of models to determine the corresponding target loss values, so that appropriate loss values can be adaptively selected to optimize the models.

[0052] Step S140: Adjust the initial recognition model according to the target loss value to obtain the target recognition model.

[0053] Combined with the foregoing embodiments for explanation, after obtaining the target loss value, the model parameters of the initial recognition model can be adjusted according to the target loss value to obtain the target recognition model. Among them, the method for adjusting the model parameters of the initial model can refer to the model learning and training process in the prior art and will not be elaborated here.

[0054] It can be seen that the present application inputs the preset test set into the initial recognition model to obtain the current similarity threshold of the initial recognition model under the false alarm rate corresponding to the preset test set; if all samples are optimized during training, it is easy to cause model overfitting; therefore, the false alarm samples and missed alarm samples to be optimized can be determined from the training set according to the current similarity threshold and the similarity distribution of the training samples in the training set of the initial recognition model, reducing the calculation amount, facilitating model training and having a better optimization effect; determining the target loss value of the initial recognition model according to the false alarm samples, missed alarm samples, preset anchor samples, and the model type of the initial recognition model, so as to be able to determine different loss values for different types of models for adaptive optimization; and then adjusting the initial recognition model according to the target loss value to obtain the target recognition model, realizing model optimization.

[0055] Based on the above embodiments, the embodiments of the present application illustrate the steps of determining false alarm samples and missed alarm samples from the training set according to the current similarity threshold and the similarity distribution of the training samples in the training set of the initial recognition model. Among them, the training set includes positive pair samples and negative pair samples. Specifically, the method of this embodiment includes the following steps:

[0056] Obtain the feature similarity of each positive pair sample and the feature similarity of each negative pair sample; determine the missed detection samples as the positive pair samples with feature similarity less than the current similarity threshold, and determine the false alarm samples as the negative pair samples with feature similarity greater than the current similarity threshold.

[0057] It should be noted that the training samples of the training set in this application include positive pair samples and negative pair samples (equivalent to including positive samples, negative samples, and anchor samples). Among them, the feature similarity of the positive pair samples refers to the feature similarity between the positive sample and the corresponding anchor sample, and the feature similarity of the negative pair samples refers to the feature similarity between the negative sample and the corresponding anchor sample.

[0058] Combined with the foregoing embodiments, the missed detection samples refer to the samples that are missed when they are actually correct. Therefore, it is necessary to find the positive pair samples (positive samples) in the positive pair sample set whose feature similarity score with the corresponding anchor sample is less than the current similarity threshold. Similarly, the false alarm samples refer to the samples that are reported when they are actually wrong. Therefore, it is necessary to find the negative pair samples (negative samples) in the negative pair sample set whose feature similarity score with the corresponding anchor sample is greater than (or greater than or equal to) the current similarity threshold.

[0059] Based on the above embodiments, the embodiments of this application will describe the steps of determining the missed detection samples as the positive pair samples with feature similarity less than the current similarity threshold, and determining the false alarm samples as the negative pair samples with feature similarity greater than the current similarity threshold. Specifically, the method of this embodiment includes the following steps:

[0060] Determine the target similarity threshold range according to the current similarity threshold and the preset range parameter; determine the missed detection samples as the positive pair samples whose feature similarity is within the target similarity range and less than the current similarity threshold, and determine the false alarm samples as the negative pair samples whose feature similarity is within the target similarity range and greater than the current similarity threshold.

[0061] Combined with the foregoing embodiments, if all samples are optimized during training, not only the calculation amount is large, but also the model is prone to overfitting; therefore, the corresponding similarity range can be determined according to the current similarity threshold, and then the missed detection samples and false alarm samples within the similarity range in the training set can be selected for optimization, which reduces the calculation amount, is easy for model training, and has a better optimization effect.

[0062] Among them, the method for determining the target similarity range where the current similarity threshold is located may include, but is not limited to, determining the target similarity range according to the current similarity threshold and preset range parameters. For example, there are preset range parameters μ1 and μ2, and the size relationship and specific values between μ1 and μ2 are not limited here. Therefore, according to the current similarity threshold anchor_th obtained under a specific false alarm rate and the preset range parameters μ1 and μ2, the target similarity range corresponding to the current similarity threshold can be determined as [anchor_th - μ1, anchor_th + μ2].

[0063] Optionally, the present application can dynamically adjust the preset range parameters and / or the target similarity range. Exemplarily, when setting the preset range parameters, the device resource specifications of the corresponding model training device can be set (for example, one or more of the CPU resource amount, GPU resource amount, memory resource amount, etc., which are not limited here). When specifically implementing the present application, the preset range parameters and / or the target similarity range can be adjusted accordingly according to the current device resource specifications. Among them, the device resource specifications are positively correlated with the target similarity range, which means that the higher the device resource specifications, the wider the target similarity range (the absolute value of the difference between the left and right boundaries is larger). Specifically, it can be achieved by directly adjusting the target similarity range, or by reducing the preset range parameter corresponding to the left boundary and / or increasing the preset range parameter corresponding to the right boundary, which will not be elaborated here.

[0064] Based on the above embodiments, the steps for determining the target loss value of the initial recognition model according to the false alarm samples, missed alarm samples, preset anchor samples, and the model type of the initial recognition model are described in the embodiments of the present application. Among them, the model type includes non-student models. Specifically, the method of this embodiment includes the following steps:

[0065] In response to the initial recognition model being a non-student model, construct a triple according to the false alarm samples, missed alarm samples, and preset anchor samples, and determine the triple loss; determine the target loss according to the triple loss and the classification loss of the initial recognition model.

[0066] Combined with the foregoing embodiments for description, if the initial recognition model is a non-student model, it means that the initial recognition model does not need to be trained by the knowledge distillation technology. Then, a triple can be constructed according to the false alarm samples, missed alarm samples, and preset anchor samples, and the triple loss can be determined; then the target loss can be determined according to the triple loss and the classification loss of the initial recognition model.

[0067] Specifically, a triple refers to a set composed of three elements. An anchor sample and its corresponding positive and negative samples can form a triple. For an explanation of triples in the prior art, please refer to the relevant content, which will not be elaborated here. The mathematical expression for the relationship in a triple can be:

[0068]

[0069] where f is the feature extraction model, a is the anchor sample anchor, p is the positive sample, n is the negative sample, x refers to the target object to be recognized (such as an image), and i refers to the triple serial number.

[0070] Then, the triple loss Loss can be determined based on the features of the samples in the foregoing example and the preset triple loss function. triplet . Then, based on the triple loss Loss triplet and the classification loss Loss cls of the initial recognition model, the target loss Loss total is determined. Its mathematical expression can be:

[0071] Loss total = Loss triplet + Loss cls

[0072] where the classification loss can also be set with reference to the calculation method of the classification loss function in the prior art, which will not be elaborated here.

[0073] Furthermore, it should be noted that in the actual training process, the number of samples meeting the above triple requirements is relatively small. However, due to the slow feature drift of the recognition model (such as the FR model), the features extracted previously in this application can be considered as approximate values output by the current network model within a certain number of iterative training steps. Therefore, a feature queue Q ∈ R K*N*d can be constructed to store the features within a certain number of iterative training steps, where K is the number of iterative training times, d is the feature dimension, and N is the number of categories participating in the training. The resulting R refers to the feature matrix composed of all the features obtained in K iterations corresponding to the images of all categories during the training of this neural network model. Thus, more abundant triple samples can be formed by the features extracted through multiple iterative trainings to optimize the training effect. During the training process, by matching the features in the current iterative stage and the features in the existing feature queue, triples within the target similarity threshold range are obtained to form the triple loss Loss triplet .

[0074] Based on the above embodiments, the embodiments of the present application will describe the steps of constructing a triple according to false positive samples, false negative samples, and preset anchor samples, and determining the triple loss. Specifically, the method of this embodiment includes the following steps:

[0075] Obtain the similarity improvement parameter corresponding to the false negative sample and the similarity reduction parameter of the false positive sample; determine the sample boundary parameter according to the similarity improvement parameter and the similarity reduction parameter; determine the triple loss according to the feature similarity between the false negative sample and the anchor sample, the feature similarity between the false positive sample and the preset anchor sample, and the sample boundary parameter.

[0076] Combined with the foregoing embodiments for description, the triple loss is to ensure that in the feature embedding space, samples from the same category are closer to each other, while samples from different categories are farther away from each other. The triple loss requires three samples to calculate the loss, and these three samples are called the anchor sample, the positive sample, and the negative sample. Among them, the anchor sample is the sample used as the recognition benchmark and needs to be concerned about. The positive sample has the same class label as the anchor sample, and the negative sample has a different class label from the anchor sample. The mathematical expression of a common triple loss function can be:

[0077] Loss triplet =max(d(A,P)-d(A,N)+margin,0)

[0078] Where d(A,P) is the feature similarity (or feature distance) between the anchor sample and the positive sample, d(A,N) is the feature similarity (or feature distance) between the anchor sample and the negative sample, and "margin" is the sample boundary parameter, which is used to control the difference between the positive sample and the negative sample. Usually, it is expected that the distance between the anchor sample and the negative sample is at least larger than the distance between the anchor sample and the positive sample.

[0079] As can be seen from the above, when implementing this embodiment, the false negative sample is equivalent to the positive sample, and the false positive sample is equivalent to the negative sample. Refer to Figure 5 as shown Figure 5 is a distribution diagram of the similarity between positive and negative samples of an exemplary initial recognition model in the training set (sampled at a certain ratio) in the model optimization method of the present application. If it is necessary to increase the similarity between the anchor sample and the false negative sample to above anchor_th, then anchor_th + Δp, Δp > 0. If it is necessary to reduce the similarity of the false positive sample to below anchor_th, then anchor_th - Δn, Δn > 0. Among them, Δp is the similarity improvement parameter, and Δn is the similarity reduction parameter. Then the mathematical expression of the sample boundary parameter margin (Δm) of this triple can be:

[0080] Δm = Δp + Δn

[0081] Based on the above embodiments, the embodiments of the present application describe the steps before determining false positive samples and false negative samples from the training set according to the current similarity threshold and the similarity distribution of the training samples in the training set of the initial recognition model. Among them, the preset test set includes a preset negative pair set, and the model type includes a student model. Specifically, the method of this embodiment includes the following steps:

[0082] In response to the initial recognition model being a student model, obtain the teacher model corresponding to the student model; determine the teacher similarity threshold of the teacher model at the preset false positive rate according to the false positive rate corresponding to the preset negative pair set; determine the target similarity range of the student model according to the teacher similarity threshold and the preset range parameters.

[0083] It should be noted that in the actual application process, many product models are obtained through model distillation technology by distilling and learning on an existing larger teacher model. In actual applications, there will be a problem that the sample similarity distributions of the student model and the teacher model are inconsistent, that is, the deviation between the similarity threshold of the student model at a certain false positive rate and the similarity threshold of its corresponding teacher model at the same false positive rate is relatively large. In response to this situation, the present application proposes a method that can add certain constraints when distilling the student model to make the threshold of the student model at a certain false positive rate basically consistent with the teacher model, thereby avoiding repeatedly adjusting the similarity threshold of the student model.

[0084] The main difference between the method of this embodiment and the method of the foregoing embodiment is that during the execution of this embodiment, if the initial recognition model is a student model, it is necessary to obtain the teacher model corresponding to the student model; determine the teacher similarity threshold th1 of the teacher model at the preset false positive rate according to the false positive rate corresponding to the preset negative pair set. Then, by referring to the description of the foregoing embodiment, the target similarity range [th1 - μ3, th1 + μ4] of the student model can be determined according to the teacher similarity threshold and the preset range parameters μ3 and μ4. Among them, the relationship between μ3 and μ4 can be referred to the description of μ1 and μ2 in the foregoing, and will not be elaborated here.

[0085] Based on the above embodiments, the embodiments of the present application describe the steps after determining the target similarity range of the student model according to the teacher similarity threshold and the preset range parameters. Specifically, the method of this embodiment includes the following steps:

[0086] Obtain false positive samples and false negative samples in the training set that are within the target similarity range; construct the correct acceptance rate loss based on the false negative samples, and construct the false acceptance rate loss based on the false positive samples; determine the target loss according to the consistency loss, correct acceptance rate loss, and false acceptance rate loss between the student model and the teacher model.

[0087] Combined with the foregoing embodiments, after obtaining the target similarity range of the student model, false positive samples and false negative samples can be determined from the training set of the student model according to the comparison result between the sample similarity corresponding to each sample and the target similarity range where the current similarity threshold is located; determine the target loss value of the initial recognition model according to the false positive samples, false negative samples, preset anchor samples, and the model type of the initial recognition model; perform adjustment processing on the initial recognition model according to the target loss value to obtain the target recognition model. Among them, the method for determining false positive samples and false negative samples can still refer to the description of the foregoing embodiments. False positive samples are equivalent to negative samples, and false negative samples are equivalent to positive samples. The difference is that in this embodiment, false positive samples and false negative samples are determined from the training samples of the student model according to the teacher similarity threshold provided by the teacher model and the target similarity range.

[0088] Exemplarily, for positive and negative pairs within the threshold range [th1 - μ3, th1 + μ4], an FAR loss (false acceptance rate loss) and a TAR loss (correct acceptance rate loss) can be constructed to constrain the model training, so that the similarity threshold of the student model at a specific false positive rate is closer to the similarity threshold of the teacher model at this false positive rate. Then, based on the FAR loss, TAR loss, and the consistency loss (distillation loss) Loss between the student model and the corresponding teacher model fcd Determine the target loss Loss of the student model total , and its mathematical expression can be:

[0089] Loss total = Loss fcd + α * Loss f + β * Loss t

[0090] Among them, Loss fcd is the distillation loss, and α and β are the preset coefficients of the FAR loss and the TAR loss respectively. After adding the FAR loss and the TAR loss, the deviation between the similarity threshold of the student model at this false positive rate and the similarity threshold of the teacher model at this false positive rate will be significantly reduced.

[0091] Specifically, the mathematical expression of the FAR loss function Loss f is:

[0092]

[0093] Among them, s n represents the feature similarity between the selected negative sample and the anchor sample on the student model, and N n represents the number of selected negative samples, and τ is a pre-set hyperparameter.

[0094] The mathematical expression of the TAR loss function Loss t is as follows:

[0095]

[0096] Among them, s p represents the feature similarity between the selected positive sample and the anchor sample on the student model, and N p represents the number of selected positive samples.

[0097] Based on the above embodiments, it should also be noted that the process of constructing the training set in this embodiment may include but is not limited to: obtaining negative pair training samples with feature similarity greater than a preset similarity threshold and positive pair training samples with feature similarity less than the preset similarity threshold; constructing a training set according to the negative pair training samples and the positive pair training samples, and the training set is used to train the initial recognition model.

[0098] During the model training process, the present application can also be directed to complex recognition scenarios such as fuzzy, complex lighting, side direction recognition, target object occlusion, and low-quality images in the actual application scenario. A special training set can be formed by collecting some materials in the actual application or selecting some materials from the existing initial training set (the training set used to train the initial recognition model before) for the model optimization method of the present application. The positive and negative pair samples in this training set need to meet: the similarity of the negative pair samples is within a certain range higher than the preset similarity threshold, and the similarity of the positive pair samples is within a certain range lower than the preset similarity threshold. For example, if the preset similarity threshold is th0, then negative pair samples can be selected from the range [th0, th0 + u], and positive pair samples can be selected from the range [th0 - v, th0]. Specifically, reference can be made to the description of the foregoing embodiments by analogy, and details are not described here. And during the model training, focus on training this part of the special training set (for example, increasing the proportion of the sample quantity of the special training set in the total training set), thereby improving the performance of the model in this complex scenario.

[0099] In summary, the method provided in this application is as follows: 1. Regarding the differences between the current recognition model during the training process and the testing process, during the training process, this difference is narrowed by optimizing the positive and negative samples at a specific false alarm rate; the triplet loss is used to optimize the false alarms and missed detections near the threshold corresponding to this false alarm rate. The margin of the triplet loss is set for the triplets that meet the conditions near the threshold corresponding to this false alarm rate, solving the problems of difficult training convergence and easy overfitting caused by uniformly setting the margin for all positive and negative samples in the traditional triplet loss method. 2. Regarding the problem that the similarity distribution of positive and negative pairs between the student model and the teacher model after model distillation may be inconsistent, which may lead to a large difference in the effects of the student model and the teacher model at a specific false alarm rate, a method based on anchoring a specific false alarm rate is proposed to reduce this difference in effects, enabling the similarity threshold of the obtained student model at a specific false alarm rate to be basically aligned with that of the teacher model (optimizing the threshold point of the student model towards the threshold point of the teacher model during the training process), avoiding the problem of repeatedly adjusting and setting the threshold during the model iteration process. 3. Regarding some complex recognition scenarios, by collecting or constructing materials of the existing model in this scenario (missed detection samples within a certain range below the specific false alarm rate threshold (similarity threshold), and false alarms within a certain range above the specific false alarm rate threshold), forming triplets, and which can be used for model version iteration, to further improve the recognition performance of the model in scenarios such as blur, occlusion, complex lighting, large angles, etc.

[0100] Further, it should be noted that the execution subject of the model optimization method can be a model optimization device. For example, the model optimization method can be executed by a terminal device, a server, or other processing devices. Among them, the terminal device can be a user equipment (UE), a computer, a mobile device, a user terminal, a terminal, a cellular phone, a cordless phone, a personal digital assistant (PDA), a handheld device, a computing device, a vehicle-mounted device, a wearable device, etc. In some possible implementation manners, the model optimization method can be implemented by a processor calling computer-readable instructions stored in a memory.

[0101] Figure 6 It is a block diagram of a model optimization device shown in an exemplary embodiment of the present application. As Figure 6 shown, the exemplary model optimization device 600 includes: a test module 610, a sample determination module 620, a loss determination module 630, and a model adjustment module 640. Specifically:

[0102] The test module 610 is configured to input a preset test set into the initial recognition model to obtain the current similarity threshold of the initial recognition model at the false alarm rate corresponding to the preset test set.

[0103] A sample determination module 620, configured to determine false positive samples and false negative samples from a training set according to a current similarity threshold and a similarity distribution of training samples in a training set of an initial recognition model.

[0104] A loss determination module 630, configured to determine a target loss value of the initial recognition model according to the false positive samples, the false negative samples, preset anchor samples, and a model type of the initial recognition model.

[0105] A model adjustment module 640, configured to perform an adjustment process on the initial recognition model according to the target loss value to obtain a target recognition model.

[0106] In this exemplary model optimization device, by inputting a preset test set into the initial recognition model, a current similarity threshold corresponding to a false positive rate of the initial recognition model in the preset test set is obtained; if all samples are optimized during training, it is easy to cause model overfitting; therefore, false positive samples and false negative samples that need to be optimized can be determined from the training set according to the current similarity threshold and the similarity distribution of training samples in the training set of the initial recognition model, reducing the calculation amount, being easy for model training, and having a better optimization effect; the target loss value of the initial recognition model is determined according to the false positive samples, the false negative samples, the preset anchor samples, and the model type of the initial recognition model, so that different loss values can be determined for different types of models for adaptive optimization; then, the initial recognition model is adjusted according to the target loss value to obtain a target recognition model, realizing model optimization.

[0107] It should be noted that the device provided in the above embodiment and the method provided in the above embodiment belong to the same concept. The specific manners in which each module and unit perform operations have been described in detail in the method embodiment, and will not be repeated here. In practical applications, the device provided in the above embodiment can, according to needs, allocate the above functions to different functional modules, that is, divide the internal structure of the device into different functional modules to complete all or part of the functions described above. This is not limited here.

[0108] Among them, the functions of each module can be referred to in the model optimization method embodiment, and will not be repeated here.

[0109] Please refer to Figure 7 , Figure 7It is a schematic structural diagram of an embodiment of the electronic device of the present application. The electronic device 100 includes a memory 101 and a processor 102. The processor 102 is configured to execute program instructions stored in the memory 101 to implement the steps in any of the above-described model optimization method embodiments. In a specific implementation scenario, the electronic device 100 may include, but is not limited to, a microcomputer, a server. In addition, the electronic device 100 may also include mobile devices such as a laptop computer, a tablet computer, etc., which are not limited herein.

[0110] Specifically, the processor 102 is configured to control itself and the memory 101 to implement the steps in any of the above-described model optimization method embodiments. The processor 102 may also be referred to as a CPU (Central Processing Unit). The processor 102 may be an integrated circuit chip with signal processing capabilities. The processor 102 may also be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. Additionally, the processor 102 may be implemented jointly by integrated circuit chips.

[0111] In this exemplary electronic device, by inputting a preset test set into the initial recognition model, the current similarity threshold of the initial recognition model at the false alarm rate corresponding to the preset test set is obtained; if all samples are optimized during training, it is easy to cause model overfitting; therefore, the false alarm samples and missed alarm samples that need to be optimized can be determined from the training set according to the current similarity threshold and the similarity distribution of the training samples in the training set of the initial recognition model, reducing the computational amount, facilitating model training, and having a better optimization effect; the target loss value of the initial recognition model is determined according to the false alarm samples, missed alarm samples, preset anchor samples, and the model type of the initial recognition model, thereby being able to determine different loss values for different types of models for adaptive optimization; then, the initial recognition model is adjusted according to the target loss value to obtain the target recognition model, realizing model optimization.

[0112] Please refer to Figure 8 , Figure 8 It is a schematic structural diagram of an embodiment of the computer-readable storage medium of the present application. The computer-readable storage medium 110 stores program instructions 111 that can be run by a processor. The program instructions 111 are used to implement the steps in any of the above-described model optimization method embodiments.

[0113] In this exemplary storage medium, by inputting a preset test set into the initial recognition model, the current similarity threshold of the initial recognition model at the false alarm rate corresponding to the preset test set is obtained; if all samples are optimized during training, it is easy to cause model overfitting; therefore, the false alarm samples and missed alarm samples to be optimized can be determined from the training set according to the current similarity threshold and the similarity distribution of the training samples in the training set of the initial recognition model, which reduces the computational amount, is easy for model training, and has a better optimization effect; the target loss value of the initial recognition model is determined according to the false alarm samples, missed alarm samples, preset anchor samples, and the model type of the initial recognition model, so that different loss values can be determined for different types of models for adaptive optimization; then, the initial recognition model is adjusted according to the target loss value to obtain the target recognition model, realizing model optimization.

[0114] In some embodiments, the functions or modules included in the device provided by the embodiments of the present disclosure can be used to execute the methods described in the above method embodiments. The specific implementation can refer to the description of the above method embodiments. For the sake of brevity, it will not be repeated here.

[0115] The descriptions of the above embodiments tend to emphasize the differences between the embodiments. Their similarities or similarities can be referred to each other. For the sake of brevity, they will not be repeated here.

[0116] In several embodiments provided by the present application, it should be understood that the disclosed methods and devices can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in an electrical, mechanical or other form.

[0117] In addition, each functional unit in various embodiments of the present application may be integrated into one processing unit, or each unit may exist physically alone, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of a software functional unit. If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it may be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, may be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the methods in various embodiments of the present application. The foregoing storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc.

Claims

1. A model optimization method, characterized in that, The method includes: Inputting a preset test set into an initial recognition model to obtain a current similarity threshold of the initial recognition model at a false alarm rate corresponding to the preset test set; Determining false alarm samples and missed alarm samples from the training set according to the current similarity threshold and the similarity distribution of training samples in the training set of the initial recognition model; Determining a target loss value of the initial recognition model according to the false alarm samples, the missed alarm samples, preset anchor samples, and the model type of the initial recognition model; Performing an adjustment process on the initial recognition model according to the target loss value to obtain a target recognition model.

2. The method according to claim 1, characterized in that, The training set includes positive pair samples and negative pair samples. The determining false alarm samples and missed alarm samples from the training set according to the current similarity threshold and the similarity distribution of training samples in the training set of the initial recognition model includes: Obtaining the feature similarities of each positive pair sample and each negative pair sample; Determining the positive pair samples with feature similarities less than the current similarity threshold as the missed alarm samples, and determining the negative pair samples with feature similarities greater than the current similarity threshold as the false alarm samples.

3. The method according to claim 2, wherein The determining the positive pair samples with feature similarities less than the current similarity threshold as the missed alarm samples, and determining the negative pair samples with feature similarities greater than the current similarity threshold as the false alarm samples includes: Determining a target similarity threshold range according to the current similarity threshold and a preset range parameter; Determining the positive pair samples with feature similarities within the target similarity range and less than the current similarity threshold as the missed alarm samples, and determining the negative pair samples with feature similarities within the target similarity range and greater than the current similarity threshold as the false alarm samples.

4. The method according to claim 1, wherein The model type includes a non-student model. The determining the target loss value of the initial recognition model according to the false alarm samples, the missed alarm samples, preset anchor samples, and the model type of the initial recognition model includes: In response to the initial recognition model being the non-student model, constructing a triple according to the false alarm samples, the missed alarm samples, and the preset anchor samples, and determining a triple loss; Determining the target loss according to the triple loss and the classification loss of the initial recognition model.

5. The method according to claim 4, wherein The constructing a triple according to the false alarm samples, the missed alarm samples, and the preset anchor samples, and determining a triple loss includes: Obtaining a similarity improvement parameter corresponding to the missed alarm samples and a similarity reduction parameter of the false alarm samples; Determining a sample boundary parameter according to the similarity improvement parameter and the similarity reduction parameter; Determining the triple loss according to the feature similarity between the missed alarm samples and the anchor samples, the feature similarity between the false alarm samples and the preset anchor samples, and the sample boundary parameter.

6. The method according to claim 1, wherein The preset test set includes a preset negative pair set, and the model type includes a student model. Before determining false positive samples and false negative samples from the training set according to the current similarity threshold and the similarity distribution of the training samples in the training set of the initial recognition model, the method further includes: In response to the initial recognition model being a student model, obtaining a teacher model corresponding to the student model; Determining a teacher similarity threshold of the teacher model at the preset false positive rate according to the false positive rate corresponding to the preset negative pair set; Determining a target similarity range of the student model according to the teacher similarity threshold and the preset range parameter.

7. The method according to claim 6, wherein After determining the target similarity range of the student model according to the teacher similarity threshold and the preset range parameter, the method further includes: Obtaining false positive samples and false negative samples in the training set that are within the target similarity range; Constructing a correct acceptance rate loss according to the false negative samples and constructing a false acceptance rate loss according to the false positive samples; Determining the target loss according to the consistency loss between the student model and the teacher model, the correct acceptance rate loss, and the false acceptance rate loss.

8. The method according to claim 1, wherein The method further includes: Obtaining negative pair training samples with a feature similarity greater than a preset similarity threshold and positive pair training samples with a feature similarity less than the preset similarity threshold; Constructing a training set according to the negative pair training samples and the positive pair training samples, and the training set is used to perform model training on the initial recognition model.

9. An electronic device, characterized in that, It includes a memory and a processor, and the processor is configured to execute program instructions stored in the memory to implement the method according to any one of claims 1 to 8.

10. A computer-readable storage medium having program instructions stored thereon, characterized in that, When the program instructions are executed by the processor, the method according to any one of claims 1 to 8 is implemented.

Citation Information

Cited By

  • Model optimization method, electronic equipment and computer readable storage medium

    CN121561424A