A target classification and identification method and related device
By training the feature extraction network and determining the domain center features, the similarity loss is corrected, and the problem that different domain environments affect the target classification recognition accuracy is solved, and high-precision target classification recognition in different domains is achieved.
Patent Information
- Application Number
- CN202510132769.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-06
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2045-02-06
AI Technical Summary
When the prior art recognizes the target in an image and performs precise classification, it is susceptible to environmental and lighting factors in different domains, resulting in misidentified or misidentified, thereby reducing the accuracy of target classification recognition.
By obtaining the trained feature extraction network, the domain to which the sample belongs, and the domain center features are determined based on the classification, domain and features of the sample. Then, based on the reference similarity between the domain center feature and the sample feature, the original similarity is corrected to obtain a similarity correction loss. Through this loss and feature extraction network, the target network is obtained for target classification.
The accuracy of target classification recognition in different domains is improved, and the probability of misidentification or misidentification caused by the target being in different domains is reduced, so that a single threshold can be adapted to the target to be identified in different domains.
Smart Images

Figure CN119600372B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image recognition technology, and in particular to a target classification and recognition method and related devices. Background Art
[0002] As the demand for image recognition continues to refine, how to identify targets in images and accurately classify them has become an important branch. When determining whether targets belong to the same category, a threshold is usually set, and the threshold is used to divide targets into the same category. However, for images collected from different domains, conventional classification and recognition methods are easily affected by factors such as the environment and lighting in different domains when using a single threshold to divide targets into the same category, resulting in misidentification or missed identification, resulting in low accuracy of target classification and recognition. In view of this, how to improve the accuracy of target classification and recognition has become an urgent problem to be solved. Summary of the invention
[0003] The main technical problem solved by the present application is to provide a target classification and recognition method and related devices, which can improve the accuracy of target classification and recognition.
[0004] To solve the above technical problems, the first aspect of the present application provides a method for target classification and recognition, comprising: obtaining a trained feature extraction network; wherein the training of the feature extraction network is completed when classifying samples in a sample set, and the sample features are obtained after the samples are input into the feature extraction network; determining the domain to which the samples in the sample set belong, and determining the domain center features of the same type of samples in the domain to which they belong based on the classification to which the samples belong, the domain to which the samples belong, and the sample features of the samples; obtaining a sample pair consisting of every two samples in the sample set, and correcting the original similarity of the sample pair based on the reference similarity between the domain center features corresponding to the samples in the sample pair and the sample features, to obtain a similarity correction loss of the sample pair; wherein, when the sample pair includes samples of different classes, the correction value of the similarity correction loss minus the original similarity is a positive number, and the correction value is negatively correlated with the reference similarity; based on the similarity correction loss and the feature extraction network, training a target network corresponding to the feature extraction network; wherein the target network is used to classify the target to be identified.
[0005] To solve the above technical problems, the second aspect of the present application provides an electronic device, which includes: a memory and a processor coupled to each other, wherein the memory stores program data, and the processor calls the program data to execute the method described in the first aspect.
[0006] In order to solve the above technical problem, the third aspect of the present application provides a computer-readable storage medium on which program data is stored. When the program data is executed by a processor, the method described in the first aspect is implemented.
[0007] The above scheme obtains a trained feature extraction network, wherein the feature extraction network is trained when the samples in the sample set are classified, and the sample features are obtained when the samples in the sample set are input into the feature extraction network. The domain to which the samples in the sample set belong is determined, and based on the classification to which the samples belong, the domain to which the samples belong, and the sample features corresponding to the samples, the sample feature distribution of the same type of samples in the domain to which they belong is analyzed, and the domain center features of the same type of samples in the domain to which they belong are determined. Every two samples in the sample set are combined into a sample pair, and the reference similarity between the sample features corresponding to the samples in the sample pair and the domain center features corresponding to the samples is obtained. Based on the reference similarity, the original similarity between the two samples in the sample pair is corrected to obtain the similarity correction loss of the sample pair. When the sample pair includes samples of different classes, the similarity correction loss is greater than the original similarity, and the correction value obtained by the difference between the two is negatively correlated with the reference similarity. Therefore, for negative sample pairs with samples of different classes, the original similarity of the negative sample pair is reduced during the iteration process by increasing the loss, thereby reducing the probability that the original similarity of the negative sample pair is high due to the target being in a different domain. In addition, for the correction value, the larger the reference similarity, the smaller the correction value, so that the extracted features can approach the domain center features during the iteration process, and the accuracy of feature extraction in different domains is improved. Based on the similarity correction loss and the feature extraction network, the target network for classifying the target to be identified is further trained. Therefore, the target network can extract high-precision features of the target in the domain where it is located, and the original similarity between negative sample pairs is small. Therefore, setting a single threshold can adapt to the targets to be identified collected in different domains and improve the accuracy of target classification and recognition. BRIEF DESCRIPTION OF THE DRAWINGS
[0008] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work. Among them:
[0009] Figure 1 It is a flowchart of an implementation method of the target recognition and classification method of the present application;
[0010] Figure 2 This is a schematic diagram of an application scenario of an implementation method of determining a domain center feature of the present application;
[0011] Figure 3It is a flowchart of another implementation method of the target recognition and classification method of the present application;
[0012] Figure 4 This is a schematic diagram of an application scenario of an implementation method of determining similarity correction loss in the present application;
[0013] Figure 5 It is a structural schematic diagram of an embodiment of the electronic device of the present application;
[0014] Figure 6 It is a structural schematic diagram of an implementation method of a computer-readable storage medium of the present application. DETAILED DESCRIPTION
[0015] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments, and different implementation methods can be adaptively combined. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0016] The terms "system" and "network" are often used interchangeably in this article. The term "and / or" in this article is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / " in this article generally indicates that the associated objects before and after are in an "or" relationship. In addition, "many" in this article means two or more than two.
[0017] The target recognition and classification method provided in the present application is used to recognize and classify targets in an image, and its corresponding execution subject is a processing unit capable of performing data processing.
[0018] See also Figure 1 , Figure 1 This is a flow chart of an implementation method of a target recognition and classification method of the present application, the method comprising:
[0019] S101: Obtain a trained feature extraction network; wherein, the training of the feature extraction network is completed when the samples in the sample set are classified, and the sample features are obtained after the samples are input into the feature extraction network.
[0020] Specifically, a trained feature extraction network is obtained, wherein the feature extraction network is trained when the samples in the sample set are classified, and the sample features are obtained when the samples in the sample set are input into the feature extraction network.
[0021] It should be noted that the samples in the sample set are images corresponding to objects belonging to different categories, and at least some of the samples in the sample set are provided with classification labels.
[0022] In some implementation scenarios, the feature extraction network corresponds to a classification network, samples with classification labels are obtained from the sample set, and the feature extraction network and its corresponding classification network are supervised trained using the samples with classification labels to obtain a trained feature extraction network.
[0023] In some implementation scenarios, the feature extraction network corresponds to an identification network and a classification network, and some samples in the sample set are provided with classification labels. All samples in the sample set are used to perform semi-supervised training on the feature extraction network and its corresponding identification network and classification network to obtain a trained feature extraction network. Among them, the labeling network is used to set labels for unlabeled samples based on the similarity between sample features.
[0024] S102: Determine the domain to which the samples in the sample set belong, and determine the domain center features of the same type of samples in the domain to which they belong based on the classification to which the samples belong, the domain to which the samples belong, and the sample features of the samples.
[0025] Specifically, the domain to which the samples in the sample set belong is determined, and based on the classification to which the samples belong, the domain to which the samples belong, and the sample features corresponding to the samples, the sample feature distribution of similar samples in the domain to which they belong is analyzed to determine the domain center features of similar samples in the domain to which they belong.
[0026] It should be noted that the domain to which the sample belongs usually corresponds to the scene, and there are great differences in images of different scenes, such as day and night, occlusion and no occlusion, clear and blurred, and images of the same target at different times.
[0027] In some implementation scenarios, domain classification is performed based on sample features corresponding to samples in the sample set to determine the domains to which the samples in the sample set belong, obtain sample subsets consisting of similar samples in the sample set, determine the feature distribution of samples in the sample subset on each domain, and based on the feature distribution corresponding to the sample subsets, determine the domain center features corresponding to each sample subset, that is, the domain center features corresponding to similar samples on each domain.
[0028] In some implementation scenarios, the domains marked by the samples in the sample set are obtained, the feature distribution of all sample features corresponding to the samples belonging to the same category is obtained, the domain center features corresponding to all sample features are determined based on the feature distribution of all sample features, and the domain center features of the samples in the domain to which they belong are determined based on the sample features belonging to each domain.
[0029] For illustration purposes, see Figure 2 , Figure 2This is a schematic diagram of an application scenario of an implementation method of determining a domain center feature of the present application, wherein: Figure 2 The corresponding sample features of the same type of samples in the same domain are marked with the same filling color, that is, Figure 2 The feature set of similar samples in all domains. The sample features corresponding to each domain can obtain the domain center features of the corresponding domain, and the sample features on all domains can obtain the domain center features corresponding to all sample features, that is, Figure 2 The domain feature center set of the same type of samples in , where the domain center feature corresponding to all sample features corresponds to the darkest dot. Therefore, samples belonging to multiple categories can obtain their own domain center features, that is, Figure 2 The feature center set of each domain of multi-class samples.
[0030] S103: Obtain a sample pair consisting of every two samples in the sample set, and correct the original similarity of the sample pair based on the reference similarity between the domain center features corresponding to the samples in the sample pair and the sample features to obtain the similarity correction loss of the sample pair; wherein, when the sample pair includes samples of different classes, the correction value of the similarity correction loss minus the original similarity is a positive number, and the correction value is negatively correlated with the reference similarity.
[0031] Specifically, every two samples in the sample set are combined into a sample pair, and the reference similarity between the sample features corresponding to the samples in the sample pair and the domain center features corresponding to the samples is obtained. Based on the reference similarity, the original similarity between the two samples in the sample pair is corrected to obtain the similarity correction loss of the sample pair.
[0032] It should be noted that when the sample pair includes samples of different classes, the similarity correction loss is greater than the original similarity, and the correction value obtained by taking the difference between the two is negatively correlated with the reference similarity. Therefore, for negative sample pairs with samples of different classes, the original similarity of the negative sample pairs is reduced during the iteration process by increasing the loss, thereby reducing the probability that the original similarity of the negative sample pairs is higher due to the targets being in different domains.
[0033] In addition, for the correction value, the greater the reference similarity, the smaller the correction value, so that the extracted features can approach the domain center features during the iteration process, thereby improving the accuracy of feature extraction in different domains.
[0034] In some implementation scenarios, the samples in the sample set are combined in pairs to obtain multiple sample pairs, wherein a sample pair including two samples of the same type is a positive sample pair, and a sample pair including samples of different types is a negative sample pair. The reference similarity between the sample features corresponding to the samples in the sample pair and the domain center features corresponding to the samples is obtained, and a correction weight is determined based on the reference similarity, wherein the correction weight of the negative sample pair is greater than 1, and the greater the reference similarity, the smaller the correction weight, and the correction weight of the positive sample pair is between 0 and 1 or the correction weight of the positive sample pair is 1, and the similarity correction loss of the sample pair is obtained based on the product of the correction weight and the original similarity.
[0035] In some implementation scenarios, the samples in the sample set are combined in pairs to obtain multiple sample pairs, wherein a sample pair including two samples of the same type is a positive sample pair, and a sample pair including samples of different types is a negative sample pair. The reference similarity between the sample features corresponding to the samples in the sample pair and the domain center features corresponding to the samples is obtained, and a correction value is determined based on the reference similarity, wherein the correction value of the negative sample pair is a positive number, and the correction value of the positive sample pair is a negative number, and the greater the reference similarity, the smaller the absolute value of the correction value. The correction value is superimposed on the original similarity to obtain the similarity correction loss of the sample pair.
[0036] S104: Based on the similarity correction loss and the feature extraction network, a target network corresponding to the feature extraction network is trained; wherein the target network is used to classify the target to be identified.
[0037] Specifically, based on the similarity correction loss and the feature extraction network, a target network for classifying the target to be identified is further trained.
[0038] It is understandable that the target network can extract high-precision features of the target in the domain where it is located, and the original similarity between negative sample pairs is small, so setting a single threshold can adapt to the targets to be identified collected in different domains and improve the accuracy of target classification and recognition.
[0039] In some implementation scenarios, the feature extraction network is further optimized and trained based on the similarity correction loss until the feature extraction network meets the preset convergence conditions, thereby obtaining a target network corresponding to the feature extraction network.
[0040] In some implementation scenarios, the feature extraction network is used as a teacher model, a student model corresponding to the teacher model is obtained, the distillation loss of the teacher model and the student model when performing feature extraction on the same sample pair is determined, the student model is iteratively trained based on the distillation loss and the similarity correction loss until the student model meets the preset convergence conditions, and the trained student model is used as the target network corresponding to the feature extraction network.
[0041] Optionally, the preset convergence condition is related to at least one of the number of iterations and the training loss.
[0042] It is understandable that the target network is matched with a classification network, so as to classify the target to be identified and complete the clustering of multiple targets. In the process of obtaining the target network, it is possible to avoid using a large amount of spare data for testing, and the optimization can be completed during the training process. The features extracted by the target network can approach the domain center features, and the similarity of the negative sample pairs determined based on the features extracted by the target network is low, which can avoid repeated testing and repeated modification of thresholds during training, and use the domain center features to smooth the noise impact, avoid convergence anomalies, and accelerate convergence.
[0043] The above scheme obtains a trained feature extraction network, wherein the feature extraction network is trained when the samples in the sample set are classified, and the sample features are obtained when the samples in the sample set are input into the feature extraction network. The domain to which the samples in the sample set belong is determined, and based on the classification to which the samples belong, the domain to which the samples belong, and the sample features corresponding to the samples, the sample feature distribution of the same type of samples in the domain to which they belong is analyzed, and the domain center features of the same type of samples in the domain to which they belong are determined. Every two samples in the sample set are combined into a sample pair, and the reference similarity between the sample features corresponding to the samples in the sample pair and the domain center features corresponding to the samples is obtained. Based on the reference similarity, the original similarity between the two samples in the sample pair is corrected to obtain the similarity correction loss of the sample pair. When the sample pair includes samples of different classes, the similarity correction loss is greater than the original similarity, and the correction value obtained by the difference between the two is negatively correlated with the reference similarity. Therefore, for negative sample pairs with samples of different classes, the original similarity of the negative sample pair is reduced during the iteration process by increasing the loss, thereby reducing the probability that the original similarity of the negative sample pair is high due to the target being in a different domain. In addition, for the correction value, the larger the reference similarity, the smaller the correction value, so that the extracted features can approach the domain center features during the iteration process, and the accuracy of feature extraction in different domains is improved. Based on the similarity correction loss and the feature extraction network, the target network for classifying the target to be identified is further trained. Therefore, the target network can extract high-precision features of the target in the domain where it is located, and the original similarity between negative sample pairs is small. Therefore, setting a single threshold can adapt to the targets to be identified collected in different domains and improve the accuracy of target classification and recognition.
[0044] See also Figure 3 , Figure 3 : is a flow chart of another embodiment of the target recognition and classification method of the present application, the method comprising:
[0045] S201: Obtain a trained feature extraction network; wherein, the training of the feature extraction network is completed when classifying samples in the sample set, and the sample features are obtained after the sample is input into the feature extraction network.
[0046] Specifically, a trained feature extraction network is obtained, wherein the feature extraction network is trained when the samples in the sample set are classified, and the sample features are obtained when the samples in the sample set are input into the feature extraction network.
[0047] In some implementation scenarios, at least some samples in the sample set are provided with classification labels. Obtaining the trained feature extraction network includes: using the samples provided with classification labels, performing supervised training on the feature extraction network and its corresponding classification network to obtain the trained feature extraction network.
[0048] Specifically, the sample is input into the feature extraction network to obtain the sample feature, and the sample feature is input into the classification network to output the estimated classification, wherein the image is converted into an Embedding feature after passing through the feature extraction network, and the Embedding feature is converted into a sample category after passing through the classification network. Based on the estimated classification and classification label, the classification loss is determined, and the feature extraction network is iteratively optimized based on the classification loss to obtain the trained feature extraction network. Therefore, the feature extraction network is trained using samples with classification labels to improve the accuracy of the feature extraction network, so that the feature extraction network has a higher reference value when used as a teacher model, and the target network corresponding to the feature extraction network is finally combined with the classification network to classify the target to be identified.
[0049] S202: Determine the domain to which the samples in the sample set belong, and determine the domain center features of the same type of samples in the domain to which they belong based on the classification to which the samples belong, the domain to which the samples belong, and the sample features of the samples.
[0050] Specifically, the domain to which the samples in the sample set belong is determined, and based on the classification to which the samples belong, the domain to which the samples belong, and the sample features corresponding to the samples, the sample feature distribution of similar samples in the domain to which they belong is analyzed to determine the domain center features of similar samples in the domain to which they belong.
[0051] In some implementation scenarios, sample features corresponding to samples in a sample set are input into a domain classification network to obtain domains to which the samples in the sample set belong; wherein the domain classification network is trained using samples with domain labels; sample feature sets corresponding to samples of the same type in each domain to which they belong are obtained, and based on each sample feature set, domain center features of samples of the same type in each domain to which they belong are determined; and based on all sample feature sets, domain center features of samples of the same type in all domains are determined.
[0052] Specifically, the parameters of the trained feature extraction network are fixed, the feature extraction network is spliced with the domain classification network, and the samples with domain labels are used to supervise the domain classification network, thereby obtaining the trained domain classification network. The sample features corresponding to the samples in the sample set are input into the domain classification network to obtain the domains to which the samples in the sample set belong, so that all samples in the sample set complete the domain classification.
[0053] Further, please refer again to Figure 2 , obtain the sample feature set consisting of the sample features corresponding to the samples of the same type in each of the domains, that is, Figure 2 The dots with the same fill color in the figure form a feature set. For each sample feature set, determine the domain center feature of the same sample in each domain to which it belongs, that is, Figure 2 The cluster center corresponding to each filled color in the dot is determined by the cluster center corresponding to each filled color. For all sample feature sets, the domain center features of similar samples in all domains are determined, that is, Figure 2 The cluster center corresponding to all the filled color dots in the cluster is the darkest dot. Therefore, the domain center features corresponding to each domain to which the same samples belong, as well as the domain center features of all sample features corresponding to the same samples are obtained, so as to obtain the domain center features that can represent the unique features of all domains. When extracting features, the extracted features can be brought closer to the domain center features, thereby improving the accuracy of feature extraction.
[0054] S203: Obtain a sample pair consisting of every two samples in the sample set, and correct the original similarity of the sample pair based on the reference similarity between the domain center features corresponding to the samples in the sample pair and the sample features to obtain the similarity correction loss of the sample pair; wherein, when the sample pair includes samples of different categories, the correction value of the similarity correction loss minus the original similarity is a positive number, and the correction value is negatively correlated with the reference similarity.
[0055] Specifically, every two samples in the sample set are combined into a sample pair, and the reference similarity between the sample features corresponding to the samples in the sample pair and the domain center features corresponding to the samples is obtained. Based on the reference similarity, the original similarity between the two samples in the sample pair is corrected to obtain the similarity correction loss of the sample pair.
[0056] In some implementation scenarios, for each sample in a sample pair, the reference similarities between the sample features corresponding to the sample and each domain center feature are obtained and sorted, and the similarity sum values corresponding to a preset number of reference similarities are determined; based on the similarity sum value, the original similarity of the sample pair is corrected to obtain the similarity correction loss of the sample pair; wherein, the absolute value corresponding to the similarity correction loss minus the correction value of the original similarity is negatively correlated with the similarity sum value.
[0057] Specifically, for each sample in the sample pair, obtain the reference similarity between the sample feature corresponding to the sample and each domain center feature, sort the reference similarities, select the top preset number of reference similarities from the sorted reference similarities and sum them to obtain the similarity sum value. When the preset number is k, that is, select the top k reference similarities and sum them to obtain the similarity sum value.
[0058] For illustration purposes, see Figure 4 , Figure 4 : This is a schematic diagram of an application scenario of an implementation method of determining similarity correction loss in the present application. Taking a sample pair as a negative sample pair including samples of different classes as an example, the cosine distance between features is used as the similarity between features, wherein the dots filled with slashes correspond to the sample features corresponding to the two samples in the negative sample pair, and the dots covered by the dotted circles are the topk domain center features, wherein k is assumed to be 3, and the reference similarities are summed to obtain the similarity sum value, wherein the process of obtaining the similarity sum value is expressed by the following formula:
[0059] (1)
[0060] in, Indicates the similarity and value corresponding to one of the samples in the sample pair, represents the sample characteristics, represents the domain center feature,<a,b> Represents the similarity calculation function between feature a and feature b.
[0061] Furthermore, based on the similarity sum value, the original similarity between the sample features corresponding to the two samples in the sample pair is corrected to obtain the similarity correction loss of the sample pair. Among them, the absolute value corresponding to the correction value of the original similarity minus the similarity correction loss is negatively correlated with the similarity sum value, that is, when the sample feature is closer to the domain center feature and the similarity sum value is larger, the adjustment amplitude of the original similarity is smaller, and the obtained similarity correction loss is smaller than the change of the original similarity. Therefore, when the feature extraction network is further iterated and updated based on the similarity correction loss, the features extracted by the feature extraction network can be closer to the domain center feature, so that the extracted features are consistent with the characteristics of the corresponding domain, and the probability of misidentification or missed identification due to the target being in a different domain is reduced.
[0062] It should be noted that the original similarity of the sample pair is corrected based on the similarity and value to obtain the similarity correction loss of the sample pair, including: determining the correction weight corresponding to the original similarity based on the similarity and value and the classification of the samples in the sample pair; wherein, when the sample pair includes samples of different categories, the correction weight is used to increase the original similarity, and when the sample pair includes samples of the same category, the correction weight is used to reduce the original similarity; the original similarity of the sample pair is corrected using the correction weight to obtain the similarity correction loss of the sample pair.
[0063] Specifically, based on the similarity and value and the classification of samples in the sample pair, a corresponding correction weight is set for the original similarity, wherein when the sample pair includes samples of different classes, that is, when the sample pair is a negative sample pair, the correction weight is used to increase the original similarity, and when the sample pair includes samples of the same class, that is, when the sample pair is a positive sample pair, the correction weight is used to reduce the original similarity, and the original similarity of the corresponding sample pair is corrected using the correction weight, so that compared with the original similarity, the original similarity of the negative sample pair is increased, and the original similarity of the positive sample pair is reduced, thereby obtaining the similarity correction loss of the sample pair.
[0064] Optionally, the product of the correction weight and the original similarity corresponds to the similarity correction loss, wherein when the sample pair includes samples of different classes, the correction weight is greater than 1, and when the sample pair includes samples of the same class, the correction weight is between 0 and 1. The process of adjusting the original similarity using the correction weight is expressed by the following formula:
[0065] (2)
[0066] in, represents the similarity correction loss, The function is a weight function that sets the correction weight based on similarity and value, which is used to adjust the original similarity of the sample pair. Represents the original similarity of the sample pair.
[0067] S204: Input the sample pair into the teacher model and the student model respectively, and determine the comparison loss of the student model compared with the teacher model when extracting features from the samples in the sample pair; wherein the feature extraction network is the teacher model, and the teacher model corresponds to the student model.
[0068] Specifically, the feature extraction network is used as a teacher model, and the teacher model corresponds to a student model. The student model includes the same network structure as the teacher model or a lightweight network structure.
[0069] It can be understood that the samples in the sample pair are input into the teacher model and the student model respectively, and the comparison loss of the sample features obtained by the student model during feature extraction compared with the teacher model is determined.
[0070] In addition, after obtaining the sample features corresponding to the samples in the sample pairs output by the student model, the similarity correction loss of the student model can be determined based on the method adopted in the above steps.
[0071] S205: Based on the comparison loss and the similarity correction loss of the student model, the student model is iteratively optimized until the convergence condition is met, and the target network corresponding to the feature extraction network is obtained.
[0072] Specifically, a certain number of sample pairs are input into the teacher model and the student model in each iteration round. The loss is corrected based on the comparison loss and the similarity of the student model, and the student model is iteratively optimized until the student model meets the preset convergence conditions. The target network corresponding to the feature extraction network is obtained, thereby improving the training accuracy through knowledge distillation and accelerating the convergence process.
[0073] In some implementation scenarios, the sample pairs are respectively input into the teacher model and the student model, and the comparison loss of the student model compared with the teacher model when extracting features from the samples in the sample pairs is determined, including: inputting the sample pairs of the current iteration round into the teacher model and the student model, respectively, to obtain the first feature pair output by the teacher model and the second feature pair output by the student model; obtaining the distillation loss based on the feature deviation between the first feature pair and the second feature pair of the current iteration round, and obtaining the distribution loss based on the distribution deviation between the first feature pair and the second feature pair of the current iteration round; wherein the comparison loss includes the distillation loss and the distribution loss.
[0074] Specifically, the sample pairs of the current iteration round are input into the teacher model to obtain a first feature pair consisting of sample features corresponding to the two samples in the sample pairs output by the teacher model; the sample pairs of the current iteration round are input into the student model to obtain a second feature pair consisting of sample features corresponding to the two samples in the sample pairs output by the student model.
[0075] Furthermore, the feature deviation between the first feature pair and the second feature pair of the current iteration round is determined to obtain the distillation loss, the distribution deviation between the first feature pair and the second feature pair of the current iteration round is determined to obtain the distribution loss, and the distillation loss and the distribution loss are combined to form the comparison loss, thereby improving the accuracy of determining the loss between the teacher model and the student model, and using the distillation loss, similarity correction loss and distribution loss to iteratively optimize the student model to accelerate model convergence.
[0076] It should be noted that the distribution loss is obtained based on the distribution deviation between the first feature pair and the second feature pair in the current iteration round, including: obtaining a first similarity set composed of the similarities between two features in all first feature pairs in the current iteration round, obtaining a second similarity set composed of the similarities between two features in all second feature pairs in the current iteration round; and obtaining the distribution loss based on the first similarity set and the second similarity set.
[0077] Specifically, the similarity between two features in all first feature pairs of the current iteration round is obtained to obtain a first similarity set, and the similarity between two features in all second feature pairs of the current iteration round is obtained to obtain a second similarity set, so as to determine the probability distribution of the feature pairs extracted by the teacher model and the student model in terms of similarity, and determine the accurate distribution loss based on the first similarity set and the second similarity set, so that when the distribution loss is used to optimize the student model, the student model can better approach the teacher model in terms of similarity distribution.
[0078] Optionally, the difference between the student and teacher models is evaluated based on the KL divergence of their probability distributions to constrain the student model training and enhance the stability and similarity consistency of the training.
[0079] In this embodiment, the feature extraction network is used as a teacher model, and the teacher model corresponds to a student model. The samples in the sample pair are respectively input into the teacher model and the student model, and the comparison loss of the sample features obtained by the student model when performing feature extraction compared with the teacher model is determined. For each sample in the sample pair, the reference similarity between the sample feature corresponding to the sample and each domain center feature is obtained, the reference similarity is sorted, and the first preset number of reference similarities are selected from the sorted reference similarities to sum up to obtain the similarity sum value. Based on the similarity sum value, the original similarity between the sample features corresponding to the two samples in the sample pair is corrected to obtain the similarity correction loss of the sample pair. A certain number of sample pairs are input into the teacher model and the student model in each iteration round, and the loss is corrected based on the comparison loss and the similarity correction loss of the student model. The student model is iteratively optimized until the student model meets the preset convergence conditions, and the target network corresponding to the feature extraction network is obtained, thereby improving the training accuracy through knowledge distillation and accelerating the convergence process.
[0080] See also Figure 5 , Figure 5 It is a structural diagram of an embodiment of an electronic device of the present application, wherein the electronic device 30 includes a memory 301 and a processor 302 coupled to each other, wherein the memory 301 stores program data (not shown), and the processor 302 calls the program data to implement the method in any of the above embodiments. For descriptions of related contents, please refer to the detailed description of the above method embodiments, which will not be repeated here.
[0081] See also Figure 6 , Figure 6 It is a structural diagram of an embodiment of a computer-readable storage medium of the present application. The computer-readable storage medium 40 stores program data 400. When the program data 400 is executed by a processor, the method in any of the above embodiments is implemented. For descriptions of related contents, please refer to the detailed description of the above method embodiments, which will not be repeated here.
[0082] It should be noted that the units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the present implementation scheme.
[0083] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0084] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) or a processor to execute all or part of the steps of each implementation method of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk and other media that can store program code.
[0085] The above description is only an implementation method of the present application, and does not limit the protection scope of the present application. Any equivalent structure or equivalent process transformation made using the contents of the present application specification and drawings, or directly or indirectly used in other related technical fields, are also included in the protection scope of the present application.
Claims
1. A target recognition and classification method, characterized in that: The method comprises: Obtaining a trained feature extraction network; wherein the training of the feature extraction network is completed when the samples in the sample set are classified, and the sample features are obtained after the samples are input into the feature extraction network; Determine the domain to which the samples in the sample set belong, and determine the domain center feature of the same type of samples in the domain to which they belong based on the classification to which the samples belong, the domain to which the samples belong, and the sample features of the samples; wherein the domain to which the samples belong corresponds to a scene, and the sample corresponds to an image corresponding to a target; Obtain a sample pair consisting of every two samples in the sample set, and based on the reference similarity between the domain center feature and the sample feature corresponding to the samples in the sample pair, correct the original similarity of the sample pair to obtain the similarity correction loss of the sample pair; wherein, when the sample pair includes samples of different classes, the correction value of the similarity correction loss minus the original similarity is a positive number, and the correction value is negatively correlated with the reference similarity; Based on the similarity correction loss and the feature extraction network, a target network corresponding to the feature extraction network is trained; wherein the target network is used to classify the target to be identified, and the target network is matched with a classification network; Using the target network to identify and classify the targets to be identified in the image, and completing clustering of multiple targets; The determining of the domain to which the samples in the sample set belong, based on the classification to which the samples belong, the domain to which the samples belong and the sample features of the samples, determines the domain center features of the same type of samples in the domain to which they belong, including: Inputting sample features corresponding to samples in the sample set into a domain classification network to obtain domains to which the samples in the sample set belong; wherein the domain classification network is trained using samples with domain labels; Obtain a sample feature set corresponding to the same type of sample in each domain to which it belongs, determine a domain center feature of the same type of sample in each domain to which it belongs based on each of the sample feature sets, and determine a domain center feature of the same type of sample in all domains to which it belongs based on all of the sample feature sets.
2. The target recognition and classification method according to claim 1, characterized in that: The correcting the original similarity of the sample pair based on the reference similarity between the domain center feature and the sample feature corresponding to the sample in the sample pair to obtain the similarity correction loss of the sample pair includes: For each sample in the sample pair, obtain the reference similarities between the sample feature corresponding to the sample and each domain center feature and sort them, and determine the similarities and values corresponding to the first preset number of the reference similarities; The original similarity of the sample pair is corrected based on the similarity sum value to obtain a similarity correction loss of the sample pair; wherein an absolute value corresponding to the similarity correction loss minus the correction value of the original similarity is negatively correlated with the similarity sum value.
3. The target recognition and classification method according to claim 2, characterized in that: The correcting the original similarity of the sample pair based on the similarity sum value to obtain the similarity correction loss of the sample pair includes: Based on the similarity and value and the classification of the samples in the sample pair, determining a correction weight corresponding to the original similarity; wherein when the sample pair includes samples of different classes, the correction weight is used to increase the original similarity, and when the sample pair includes samples of the same class, the correction weight is used to decrease the original similarity; The original similarity of the sample pair is corrected using the correction weight to obtain a similarity correction loss of the sample pair.
4. The target recognition and classification method according to claim 1, characterized in that: The feature extraction network is a teacher model, and the teacher model corresponds to a student model; The training of obtaining a target network corresponding to the feature extraction network based on the similarity correction loss and the feature extraction network includes: Inputting the sample pair into the teacher model and the student model respectively, and determining the comparison loss of the student model compared with the teacher model when extracting features from the samples in the sample pair; Based on the comparison loss and the similarity correction loss of the student model, the student model is iteratively optimized until a convergence condition is met, thereby obtaining a target network corresponding to the feature extraction network.
5. The target recognition and classification method according to claim 4, characterized in that: The step of inputting the sample pair into the teacher model and the student model respectively, and determining a comparison loss of the student model compared with the teacher model when extracting features from the samples in the sample pair, comprises: Inputting the sample pairs of the current iteration round into the teacher model and the student model respectively, obtaining a first feature pair output by the teacher model and a second feature pair output by the student model; Based on the feature deviation between the first feature pair and the second feature pair in the current iteration round, a distillation loss is obtained, and based on the distribution deviation between the first feature pair and the second feature pair in the current iteration round, a distribution loss is obtained; wherein the comparison loss includes the distillation loss and the distribution loss.
6. The target recognition and classification method according to claim 5, characterized in that: The obtaining of the distribution loss based on the distribution deviation between the first feature pair and the second feature pair in the current iteration round includes: Obtain a first similarity set consisting of similarities between two features in all first feature pairs of the current iteration round, and obtain a second similarity set consisting of similarities between two features in all second feature pairs of the current iteration round; The distribution loss is obtained based on the first similarity set and the second similarity set.
7. The target recognition and classification method according to claim 1, characterized in that: At least some of the samples in the sample set are provided with classification labels; The step of obtaining a trained feature extraction network includes: The feature extraction network and its corresponding classification network are supervisedly trained using samples with the classification labels to obtain the trained feature extraction network; wherein the target network is combined with the classification network to classify the target to be identified.
8. An electronic device, characterized in that: include: A memory and a processor coupled to each other, wherein the memory stores program data, and the processor calls the program data to execute the method according to any one of claims 1 to 7.
9. A computer-readable storage medium having program data stored thereon, characterized in that: When the program data is executed by a processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Unsupervised pedestrian re-recognition system and method based on sample separation
CN113065516A
Target recognition model training method, target recognition method and related device
CN115439707A