Method, apparatus, device and storage medium for determining labels of training images

The clustering feature matrix assists in determining the target label of the training image, which solves the problem of poor training sample labels in the prior art, improves the accuracy of the neural network and reduces the workload of manual calibration.

CN119380145BActive Publication Date: 2025-05-27JINAN BOGUAN INTELLIGENT TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411965112.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-30
Publication Date
2025-05-27
Estimated Expiration
2044-12-30

AI Technical Summary

Technical Problem

The labels of training samples determined by label smoothing techniques or data cleaning techniques in the prior art are not effective when training network models.

Method used

By acquiring pre-noted multiple training images and their original category labels, multiple target clustering parameters of the preset clustering algorithm are determined, and the training images are clustered based on these parameters to obtain the feature matrix. Then, the target category label of the training image is determined using the similarity between the feature matrix output by the neural network and the cluster feature matrix.

Benefits of technology

Reduces the dependence of neural networks on labels, improves the accuracy of the final trained neural network, and reduces the workload of manual label calibration.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119380145B_ABST
    Figure CN119380145B_ABST
Patent Text Reader

Abstract

The present invention provides a method, apparatus, device and storage medium for determining labels of training images, which are applied to the field of computer technology. The method includes: obtaining a plurality of training images and a plurality of target clustering parameters of a preset clustering algorithm; clustering the plurality of training images according to the total number of original class labels and each target clustering parameter to determine a plurality of first feature matrices of the plurality of training images; the number of the plurality of first feature matrices is the same as the total number of class labels; inputting each training image into a first neural network model for classification processing to determine a second feature matrix corresponding to each training image; for each training image, obtaining a target first feature matrix corresponding to the original class label of the training image from the plurality of first feature matrices, and determining a target class label corresponding to the training image according to the similarity between the second feature matrix and the target first feature matrix. By adopting the technical solution of the present invention, the accuracy of the finally trained neural network can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technologies, and in particular, to a method, an apparatus, a device, and a storage medium for determining labels of training images. Background Art

[0002] With the rapid development of artificial intelligence technologies, the current expectations for artificial intelligence-related algorithm metrics are also getting higher and higher, which requires training artificial intelligence-related network models with higher precision and accuracy.

[0003] In related technologies, during the training process of a network model, for example, different training samples (such as different training images) can be collected in advance, and then these training samples can be labeled separately, and then the network model can be trained with the labeled training samples. However, during the labeling process, it may be prone to mislabeling due to difficulties in calibrating data labels or other reasons, such as blurred images making it difficult to determine the category, mislabeling caused by the labeler himself, etc. These may all lead to inaccurate labels finally obtained. Currently, there are also some methods to improve the accuracy of the labeled labels, such as label smoothing technology, data cleaning technology, etc.

[0004] However, the labels of the training samples determined by the above technologies still have poor effects when training the network model. Summary of the Invention

[0005] The present invention provides a method, an apparatus, a device, and a storage medium for determining labels of training images, to solve the defect that the labels of training samples determined by label smoothing technology or data cleaning technology in the prior art still have poor effects, and to realize determining multiple target clustering parameters of a preset clustering algorithm based on a training image and its original category label, and determining the target category label of each training image according to the similarity between the feature matrix obtained after clustering multiple training images based on the target clustering parameters and the feature matrix output by a neural network. In this way, determining the target label of a training image with the assistance of the clustered feature matrix can reduce the dependence of the neural network on the label, improve the accuracy of the finally trained neural network, and reduce the workload of manual label calibration.

[0006] The present invention provides a method for determining labels of training images, including:

[0007] Obtaining a plurality of pre-annotated training images and obtaining a plurality of target clustering parameters corresponding to a preset clustering algorithm; each training image includes its own corresponding original category label, and the above-mentioned plurality of target clustering parameters are determined in advance according to the plurality of training images and the original category label of each training image;

[0008] Cluster multiple training images according to the total number of categories corresponding to the original category labels and each target clustering parameter to determine multiple first feature matrices corresponding to the multiple training images; the number of the multiple first feature matrices is the same as the total number of categories, and each first feature matrix corresponds to the features of the training images under one category label.

[0009] Input each training image into the first neural network model for classification processing to determine the second feature matrix corresponding to each training image.

[0010] For each training image, obtain the target first feature matrix corresponding to the original category label of the training image from the multiple first feature matrices, and determine the target category label corresponding to the training image according to the similarity between the second feature matrix and the target first feature matrix; the target category label is used to train the first neural network model.

[0011] According to a method for determining the label of a training image provided by the present invention, the target category label of the training image is represented by a label vector, and the label vector includes the probabilities corresponding to various category labels in the total categories. The step of determining the target category label corresponding to the training image according to the similarity between the second feature matrix and the target first feature matrix includes:

[0012] Calculate the first similarity between the second feature matrix and the target first feature matrix.

[0013] If the first similarity is within the first threshold range, set the probability corresponding to the original category label of the training image in the label vector to 1 and set the probabilities corresponding to other category labels to 0 to obtain the target category label corresponding to the training image; the other category labels are the other category labels in the total categories except the original category label of the training image.

[0014] If the first similarity exceeds the first threshold range, calculate the second similarity between the second feature matrix and other first feature matrices in the multiple first feature matrices, and determine the target category label corresponding to the training image according to the second similarity and the second threshold range; the other first feature matrices are the remaining first feature matrices in the multiple first feature matrices except the target first feature matrix.

[0015] Wherein, the similarity corresponding to the first threshold range is greater than the similarity corresponding to the second threshold range.

[0016] According to a method for determining the label of a training image provided by the present invention, before determining the target category label corresponding to the training image according to the similarity between the second feature matrix and the target first feature matrix, the method further includes:

[0017] Obtain the maximum and minimum values in the second feature matrix, and obtain the maximum and minimum values in the target first feature matrix;

[0018] According to the maximum and minimum values in the second feature matrix and the maximum and minimum values in the target first feature matrix, perform feature stretching processing on the second feature matrix to determine the second feature matrix after feature stretching; the eigenvalue ranges corresponding to the stretched second feature matrix and the target first feature matrix are the same;

[0019] Correspondingly, the above-mentioned determining the target class label corresponding to the training image according to the similarity between the second feature matrix and the target first feature matrix includes:

[0020] Determine the target class label corresponding to the training image according to the similarity between the second feature matrix after feature stretching and the target first feature matrix.

[0021] According to a method for determining the label of a training image provided by the present invention, the above-mentioned determining the target class label corresponding to the training image according to the similarity between the second feature matrix after feature stretching and the target first feature matrix includes:

[0022] According to the second feature matrix after feature stretching, calculate the mean and standard deviation corresponding to the second feature matrix after feature stretching;

[0023] Obtain the preset maximum number of digits and minimum number of digits, and perform outlier removal on the second feature matrix after feature stretching according to the maximum number of digits, minimum number of digits, mean, and standard deviation to determine the second feature matrix after outlier removal;

[0024] Determine the target class label corresponding to the training image according to the similarity between the second feature matrix after outlier removal and the target first feature matrix.

[0025] According to a method for determining the label of a training image provided by the present invention, the above-mentioned obtaining multiple target clustering parameters corresponding to a preset clustering algorithm includes:

[0026] Divide multiple training images into multiple sets of training images according to their original class labels; each set of training images includes training images under one class label;

[0027] Perform distance calculation processing on the training images in each set of training images respectively to determine multiple target clustering parameters corresponding to the preset clustering algorithm.

[0028] According to a method for determining the label of a training image provided by the present invention, the above-mentioned multiple target clustering parameters include a target neighborhood radius and a target minimum number of samples. The above-mentioned performing distance calculation processing on the training images in each set of training images respectively to determine multiple target clustering parameters corresponding to the preset clustering algorithm includes:

[0029] For each category label, input the training image set corresponding to the category label into the second neural network model, calculate the distance between every two training images in the second neural network model, obtain multiple first distances corresponding to the training image set, and determine the target neighborhood radius under the category label corresponding to the training image set according to the multiple first distances;

[0030] For each category label, input the training images in the training image set corresponding to the category label into the third neural network model in different batches for distance calculation, determine the model evaluation indexes corresponding to different batches respectively, and determine the target batch among different batches and use the number of training images corresponding to the target batch as the target minimum sample number under the category label corresponding to the training image set according to the model evaluation indexes corresponding to different batches respectively.

[0031] According to a method for determining the label of a training image provided by the present invention, the above multiple target clustering parameters include a target distance metric threshold, and the training images in each type of training image set are respectively subjected to distance calculation processing to determine multiple target clustering parameters corresponding to a preset clustering algorithm, including:

[0032] For each category label, calculate the distance between each training image in the training image set corresponding to the category label and its own original category label, and obtain multiple second distances;

[0033] Eliminate the second distances with large differences among the multiple second distances, and select multiple candidate second distances with the top-ranked median in the eliminated second distances;

[0034] Calculate the average value corresponding to the multiple candidate second distances, determine the target second distance among the multiple candidate second distances according to the average value, and determine the target second distance as the target distance metric threshold under the category label corresponding to the training image set.

[0035] The present invention also provides a device for determining the label of a training image, including the following modules:

[0036] An acquisition module, configured to acquire multiple pre-annotated training images and acquire multiple target clustering parameters corresponding to a preset clustering algorithm; each training image includes its own corresponding original category label, and the above multiple target clustering parameters are determined in advance according to the multiple training images and the original category label of each training image;

[0037] A clustering module, configured to cluster the multiple training images according to the total number of category labels corresponding to the original category labels and each target clustering parameter, and determine multiple first feature matrices corresponding to the multiple training images; the number of the above multiple first feature matrices is the same as the total number of category labels, and each first feature matrix corresponds to the features of the training images under one category label;

[0038] A model processing module, configured to input each training image into a first neural network model for classification processing to determine a second feature matrix corresponding to each training image;

[0039] A target label determination module, configured to, for each training image, obtain a target first feature matrix corresponding to the original class label of the training image from multiple first feature matrices, and determine a target class label corresponding to the training image according to the similarity between the second feature matrix and the target first feature matrix; the target class label is used to train the first neural network model.

[0040] The present invention further provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the method for determining the label of a training image as described in any one of the above is implemented.

[0041] The present invention further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the method for determining the label of a training image as described in any one of the above is implemented.

[0042] The present invention further provides a computer program product, including a computer program. When the computer program is executed by a processor, the method for determining the label of a training image as described in any one of the above is implemented.

[0043] The method, device, equipment, and storage medium for determining the label of a training image provided by the present invention obtain multiple training images with pre-labeled original class labels and obtain multiple target clustering parameters corresponding to them determined according to the multiple training images and their original class labels in a preset clustering algorithm. Then, according to the total number of classes corresponding to the original class labels and each target clustering parameter, the multiple training images are clustered to determine multiple first feature matrices corresponding to the multiple training images and having the same number as the total number of classes. At the same time, each training image is input into a first neural network model for classification processing to determine a second feature matrix and a predicted class corresponding to each training image. Then, for each training image, a target first feature matrix corresponding to the original class label of the training image is obtained from multiple first feature matrices, and a target class label corresponding to the training image is determined according to the similarity between the second feature matrix and the target first feature matrix; wherein, each first feature matrix corresponds to the features of the training images under one class label. In this method, since the feature matrix obtained by clustering can be used to assist the feature matrix output by the neural network model to determine the target label of the training image, the dependence of the neural network model on its own original label during the training process can be reduced, the accuracy of the finally trained neural network model can be improved, and at the same time, the workload of manual label calibration can be reduced. Description of the Drawings

[0044] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the accompanying drawings required in the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on these drawings.

[0045] Figure 1 It is one of the schematic flowcharts of the method for determining the label of the training image provided by the embodiment of the present invention.

[0046] Figure 2 It is another schematic flowchart of the method for determining the label of the training image provided by the embodiment of the present invention.

[0047] Figure 3 It is the third schematic flowchart of the method for determining the label of the training image provided by the embodiment of the present invention.

[0048] Figure 4 It is the schematic structural diagram of the device for determining the label of the training image provided by the embodiment of the present invention.

[0049] Figure 5 It is the schematic structural diagram of the electronic device provided by the embodiment of the present invention. Specific Embodiments

[0050] To make the objectives, technical solutions, and advantages of the present invention clearer, the following will clearly and completely describe the technical solutions in the present invention with reference to the accompanying drawings in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention.

[0051] Some existing methods for improving the label accuracy of calibration, such as label smoothing technology, data cleaning technology, etc., still have poor effects when determining the labels of training samples for training the network model. Based on this, the embodiments of the present invention provide a method, device, equipment, and storage medium for determining the label of a training image, which can solve this problem.

[0052] The following will describe the method for determining the label of the training image in the embodiments of the present invention in combination with Figures 1 - 3 Describe the method for determining the label of the training image in the embodiments of the present invention.

[0053] It should be noted that the execution subject of the embodiments of the present invention can be the method for determining the label of the training image, or an electronic device, or other devices or equipment, which is not specifically limited here. The following embodiments will be described by taking an electronic device as the execution subject as an example. The electronic device can be a terminal or a server.

[0054] Figure 1 is one of the schematic flowcharts of the method for determining the label of the training image provided by the embodiment of the present invention. As Figure 1 shown, the method includes the following steps:

[0055] Step 102, obtain a plurality of pre-annotated training images and obtain a plurality of target clustering parameters corresponding to a preset clustering algorithm; each training image includes its own corresponding original class label, and the above-mentioned plurality of target clustering parameters are determined in advance according to the plurality of training images and the original class label of each training image.

[0056] Among them, before training the first neural network model, a plurality of different types of training images can be collected in advance, and then a class label is set for each training image. The label set here can be recorded as the original class label, and the original class label can be a manually calibrated class label. Here, different class labels can include, for example, people, vehicles, animals, plants, etc.

[0057] At the same time, a preset clustering algorithm can be obtained, and then the clustering parameters corresponding to the preset clustering algorithm are determined. For example, for the density clustering algorithm, its corresponding clustering parameters can include parameters such as the neighborhood radius, the minimum number of samples, and the distance metric threshold.

[0058] After obtaining the plurality of training images and setting the original class label for each training image, the target value corresponding to the clustering parameters of the preset clustering algorithm can be determined through each training image and its own original class label, and then each clustering parameter and its target value can be recorded as the target clustering parameter, so that a plurality of target clustering parameters can be obtained.

[0059] It can be understood that the plurality of target clustering parameters finally determined here are for the plurality of training images, so the plurality of target clustering parameters determined are more targeted at clustering the plurality of training images themselves during clustering.

[0060] Step 104, cluster the plurality of training images according to the total number of classes corresponding to the original class label and each target clustering parameter, and determine a plurality of first feature matrices corresponding to the plurality of training images; the number of the above-mentioned plurality of first feature matrices is the same as the total number of classes, and each first feature matrix corresponds to the features of the training images under one class label.

[0061] Among them, after obtaining each training image with the original category label set, the total number of category labels, that is, the total number of categories, can also be statistically obtained through the set original category labels. For example, there are 100 training images, among which the original category labels of 30 training images are people, the original category labels of 30 training images are vehicles, and the original category labels of 40 training images are animals. Then the total number of categories is three, that is, the three category labels of people, vehicles, and animals are included in the multiple training images in total.

[0062] Then, after obtaining the multiple target clustering parameters of the preset clustering algorithm, assuming that the multiple target clustering parameters are the multiple target clustering parameters determined for each category label, that is, each category label has multiple target clustering parameters. Then, all the training images can be re-clustered according to the multiple target clustering parameters of each category label to obtain multiple clustering results equal to the total number of categories. For example, three clustering results are obtained, corresponding to the clustering of the training images of people, vehicles, and animals respectively. Among them, each clustering result includes the training images under the corresponding category label. Then, the feature extraction can be performed on the training images under each category label to obtain the feature matrix corresponding to all the training images under this category label, which is all denoted as the first feature matrix. For example, the first feature matrix corresponding to the category of people, the first feature matrix corresponding to the category of vehicles, and the first feature matrix corresponding to the category of animals can be obtained here.

[0063] It can be seen that the number of the first feature matrices here is the same as the number of clustering results and also the same as the total number of categories.

[0064] Step 106: Input each training image into the first neural network model for classification processing to determine the second feature matrix corresponding to each training image.

[0065] In this step, after obtaining the multiple training images and setting the original category labels, each training image can also be input into the first neural network model for classification processing to output the predicted category corresponding to each training image. At the same time, the feature matrix corresponding to each training image can be output through the convolutional layer before the softmax classification layer of the first neural network model. The feature matrix output by each training image can be denoted as the second feature matrix. That is to say, the second feature matrix of each training image here can represent the predicted category of this training image.

[0066] The first neural network model here can be an initially untrained neural network model, or it can also be a neural network model that has started training but is not yet fully trained. In short, the first neural network model can be an incompletely trained neural network model. There is no specific limitation on the specific architecture of the first neural network model here. For example, it can be a CNN (Convolutional Neural Networks), an RNN (Recurrent Neural Network), etc.

[0067] In addition, the main function of the first neural network model here is to classify the input image, that is, the first neural network model can also be a classification model. Of course, the first neural network model can also perform other image processing on the input image, such as image detection processing, etc., which is not specifically limited here.

[0068] It should be noted that for the execution order of the above steps 104 and 106, the two can be executed synchronously, or they can be executed sequentially. For example, step 104 can be executed first and then step 106, or step 106 can be executed first and then step 104. The above step number order is only an example.

[0069] Step 108, for each training image, obtain the target first feature matrix corresponding to the original class label of the training image from multiple first feature matrices, and determine the target class label corresponding to the training image according to the similarity between the second feature matrix and the target first feature matrix; the target class label is used to train the first neural network model.

[0070] Among them, the predicted class of each of the above training images may be the same as its original class label, or it may be different from its original class label. When the predicted class is different from the original class label, it may be that the parameters of the network model are not adjusted properly, or it may be that the originally labeled class label by humans is incorrect, etc. If it is the case where the originally labeled class label by humans is incorrect, then when training the first neural network model subsequently, using this incorrect original class label for training will affect the accuracy of the finally trained first neural network model. Therefore, in this step, similarity calculation is used to determine the final class label for model training, so as to improve the accuracy of the trained model.

[0071] Specifically, after obtaining the first feature matrix of various category labels by re-clustering multiple training images using multiple target clustering parameters of a preset clustering algorithm, and after obtaining the second feature matrix of a certain training image output by the first neural network model, the first feature matrix corresponding to the original category label of this training image can be selected from multiple first feature matrices and denoted as the target first feature matrix. Then, the similarity between the target first feature matrix and the second feature matrix of this training image output by the first neural network model can be calculated, and then whether the original category label of this training image needs to be corrected can be determined based on the calculated similarity. For example, if the target first feature matrix is relatively similar to the second feature matrix of a certain training image, it indicates that the original category label of this training image is correctly calibrated and does not need to be corrected, and the original category label of this training image can be used as its final target category label; for another example, if the target first feature matrix is not very similar to the second feature matrix of a certain training image, it indicates that there may be an error in the calibration of the original category label of this training image, and the original category label needs to be corrected to obtain a more accurate target category label.

[0072] After determining the target category label of each training image, the target category label of each training image can be used as the calibrated label and directly input into the loss function for loss calculation, or the target category label and the predicted category can also be input into the loss function for loss calculation; then the first neural network model is trained based on the calculated loss, and finally the trained first neural network model is obtained.

[0073] As can be seen from the above description, when training the first neural network model, the accuracy of the calibration of the original category label can be determined by the similarity between the features of the clustered images and the features output by the network model, and a more accurate target category label can be determined and input into the loss function to train the network model. This can reduce the dependence of the network model training on the original category label, effectively avoid the problem that the original category label is incorrect due to reasons such as blurred images and unclear classification criteria in the dataset (i.e., training images), which affects the training accuracy of the model, thereby effectively improving the accuracy of the finally trained model, and there is no need to manually calibrate the original category label repeatedly, so the workload of manual label calibration can also be reduced.

[0074] In this embodiment, by obtaining a plurality of training images with pre-labeled original class labels and obtaining a plurality of corresponding target clustering parameters determined according to the plurality of training images and their original class labels in a preset clustering algorithm, then clustering the plurality of training images according to the total number of classes corresponding to the original class labels and each target clustering parameter to determine a plurality of first feature matrices corresponding to the plurality of training images and the same as the total number of classes. At the same time, inputting each training image into a first neural network model for classification processing to determine a second feature matrix and a predicted class corresponding to each training image, and then for each training image, obtaining a target first feature matrix corresponding to the original class label of the training image among the plurality of first feature matrices, and determining the target class label corresponding to the training image according to the similarity between the second feature matrix and the target first feature matrix; wherein, each first feature matrix corresponds to the features of the training images under one class label. In this method, since the feature matrix obtained by clustering can be used to assist the feature matrix output by the neural network model to determine the target label of the training image, this can reduce the dependence of the neural network model on its own original label during the training process, improve the accuracy of the finally trained neural network model, and at the same time reduce the workload of manual label calibration.

[0075] The following embodiments illustrate the process of determining the target class label corresponding to the training image through similarity.

[0076] In some embodiments, the step of "determining the target class label corresponding to the training image according to the similarity between the second feature matrix and the target first feature matrix" in step 108 above may include the following steps:

[0077] Calculate the first similarity between the second feature matrix and the target first feature matrix;

[0078] If the first similarity is within the first threshold range, set the probability corresponding to the original class label of the training image in the label vector to 1 and set the probabilities corresponding to other class labels to 0 to obtain the target class label corresponding to the training image; the above other class labels are other class labels in the total classes except the original class label of the training image;

[0079] If the first similarity exceeds the first threshold range, calculate the second similarity between the second feature matrix and other first feature matrices among the plurality of first feature matrices, and determine the target class label corresponding to the training image according to the second similarity and the second threshold range; the above other first feature matrices are the remaining first feature matrices among the plurality of first feature matrices except the target first feature matrix; wherein, the similarity corresponding to the first threshold range is greater than the similarity corresponding to the second threshold range.

[0080] Among them, in this embodiment, it is proposed to represent the target category label of the training image by a label vector, and the label vector includes the probabilities corresponding to various category labels in the total categories. Assuming the total categories are: 1, 2, 3, ..., N, then the label vector of the target category label of a certain training image can be expressed as [prob1, prob2, prob3, ..., probN], where it can include the respective probabilities prob corresponding to each category label. For example, prob1 refers to the probability corresponding to category label 1, and probN refers to the probability corresponding to category label N. The range of prob is 0 to 1. By default, in the label vector corresponding to a certain category label, the probability of this category label is 1, and the other probabilities are all 0, and the sum of the probabilities corresponding to all category labels of each original category label is 1, that is, prob1 + prob2 + prob3 + ... + probN = 1. Exemplarily, assuming that the original category label of a certain training image is category label 1 and its calibration is correct, then the label vector corresponding to the target category label of this training image is [1, 0, 0, ..., 0], that is, the probability of category label 1 corresponding to this training image is 1, and the probabilities of other category labels are all 0.

[0081] On the basis of the above description, for each training image, after obtaining the similarity between the second feature matrix of the training image and the target first feature matrix corresponding to its original category label, this similarity can represent the feature similarity between this training image and a group of training images corresponding to its original category label. The larger the similarity value, the more similar this training image and a group of training images corresponding to its original category label are.

[0082] Based on this, multiple threshold ranges for the similarity can be preset here. For example, a first threshold range and a second threshold range can be set. The similarity corresponding to the first threshold range is greater than the similarity corresponding to the second threshold range. For example, the first threshold range is [0.8, 1], and the second threshold range is [0.5, 0.8).

[0083] Then for each training image, the similarity between the second feature matrix of the training image and the target first feature matrix can be calculated first, denoted as the first similarity. Then it can be judged whether this first similarity falls within the first threshold range. If the first similarity of this training image falls within the first threshold range, it means that this training image and a group of training images corresponding to its original category label are relatively similar, and at the same time, it can also mean that the calibration of the original category label corresponding to this training image is correct. At this time, in the label vector corresponding to the target category label of this training image, the probability corresponding to the original category label can be set to 1 and the probabilities corresponding to other category labels can be set to 0. The finally obtained label vector can be used as the target category label corresponding to this training image.

[0084] If the first similarity of the training image does not fall within the first threshold range, that is, it exceeds the first threshold range, it indicates that the training image may not be very similar to a set of training images corresponding to its original class label. At this time, the remaining first feature matrices except the target feature matrix can be obtained from multiple first feature matrices, denoted as other first feature matrices, and then the similarity between the training image and other first feature matrices is calculated, all denoted as the second similarity. Then, it can be determined whether each second similarity of the training image falls within the second threshold range. If a certain second similarity falls within the second threshold range, it indicates that the true class label of the training image may be its original class label, or it may also be the class label corresponding to the second similarity that falls within the second threshold range. At this time, for the convenience of training the model, the probability corresponding to the original class label of the training image can be set to 1 - r in the label vector, the probability corresponding to the class label corresponding to the second similarity that falls within the second threshold range can be set to r, and the probabilities of other remaining class labels are all set to 0. The finally obtained label vector can be used as the target class label corresponding to the training image. Among them, the value of r can be set according to the actual situation, such as 0.3, 0.4, etc.

[0085] Furthermore, if the first similarity of the training image does not fall within the first threshold range and all second similarities of the training image do not fall within the second threshold range, that is, the first similarity of the training image exceeds the first threshold range and all second similarities exceed the second threshold range, it is determined that the original class label of the training image is incorrect. At this time, in order to prevent the training image from affecting the training accuracy of the model during training, the training image and its original class label can be excluded and not participate in the subsequent training of the first neural network model; or the original class label of the training image can be re-verified and calibrated manually, and the calibrated label is then used to repeat the above steps to calculate the similarity and perform similarity matching to determine the target class label. In this way, only the labels of some training images need to be re-verified manually, thus greatly reducing the workload of manual calibration.

[0086] In this embodiment, when the similarity between the features predicted by the training image and the features of a set of training images corresponding to its original class label falls within different similarity threshold ranges, the probability / weight of its original class label is re-adjusted, so that the dependence of model training on the original class label can be further weakened, and the accuracy of the finally trained model can be further improved.

[0087] Furthermore, in another embodiment, the process of determining the target class label of the training image and training the first neural network model is described when the target class label is not represented by a label vector.

[0088] After obtaining the similarity between the second feature matrix of a certain training image and the target first feature matrix, the similarity can also be compared with a preset similarity threshold. If the similarity is greater than the preset similarity threshold, it indicates that the original class label of the training image is correctly calibrated, that is, the target class label of the training image is its original class label. Subsequently, the first neural network model can be trained jointly by the predicted class of the training image and the target class label.

[0089] If the above similarity is not greater than the preset similarity threshold, it indicates that the original class label of the training image is incorrectly calibrated. At this time, the training image and its original class label can be directly excluded and not participate in the subsequent training of the first neural network model; or the original class label of the training image can be manually re-reviewed and calibrated, and the calibrated label can be used to repeat the above steps to calculate the similarity and determine the target class label, and then participate in the training of the first neural network model.

[0090] In addition, the size of the above preset similarity threshold can be set according to the actual situation, such as 0.8, 0.85, etc.

[0091] In this embodiment, by comparing the calculated similarity with the similarity threshold and determining whether the training image participates in the model training based on the comparison result, the process is relatively simple and intuitive, so the model training efficiency can be improved; at the same time, the training images with calculated similarities less than the similarity threshold and their original class labels can be excluded from the model training, which can improve the accuracy of the trained model.

[0092] When actually calculating the similarity between the features predicted by the training image and the features of a set of training images corresponding to its original class label, there may be a situation where the size of the features output by the model is inconsistent with the features obtained by clustering, which ultimately affects the accuracy of the model training. Based on this, in this embodiment, a technical solution for feature stretching of the features output by the model is proposed before calculating the similarity. The following embodiments will illustrate this process.

[0093] Figure 2 is the second flow chart of the method for determining the label of the training image provided by the embodiment of the present invention. As Figure 2 shown, before the above step 108 "determine the target class label corresponding to the training image according to the similarity between the second feature matrix and the target first feature matrix", the above method may further include the following steps:

[0094] Step 202, obtain the maximum value and the minimum value in the second feature matrix, and obtain the maximum value and the minimum value in the target first feature matrix.

[0095] In this step, for each training image, after obtaining the second feature matrix output by the first neural network model, assuming that the second feature matrix output by the first neural network model is of size B, C, H, W, where B is the batch size, C is the number of channels, H is the height of the feature output by the model, and W is the width of the feature output by the model, then the second feature matrix can include multiple eigenvalues. Then, the maximum eigenvalue (denoted as the maximum value) and the minimum eigenvalue (denoted as the minimum value) can be found from the multiple eigenvalues of the second feature matrix.

[0096] Step 204: According to the maximum and minimum values in the second feature matrix and the maximum and minimum values in the target first feature matrix, perform feature stretching processing on the second feature matrix to determine the second feature matrix after feature stretching; the eigenvalue range of the stretched second feature matrix is the same as that of the target first feature matrix.

[0097] In this step, a new network layer (denoted as the feature conversion layer) can be added to the first neural network model to perform feature stretching processing on the feature matrix.

[0098] After obtaining the target first feature matrix corresponding to the original class label, the target first feature matrix also includes multiple eigenvalues. Then, the maximum and minimum values in the target first feature matrix can also be obtained. Then, through the maximum and minimum values in the second feature matrix and the maximum and minimum values in the target first feature matrix, the feature conversion layer in the first neural network model can be used to perform feature stretching processing on the second feature matrix to obtain the second feature matrix after feature stretching. For example, the following formula can be used for feature stretching processing:

[0099] 。

[0100] Among them, X’ represents the second feature matrix after feature stretching; X represents the second feature matrix output by the first neural network model; X max and X min respectively represent the maximum and minimum values in the second feature matrix output by the first neural network model; new max and new min respectively represent the maximum and minimum values in the target first feature matrix; argmin represents taking the minimum value; argmax represents taking the maximum value; C represents the target first feature matrix (i.e., the core feature array obtained during clustering), and C i represents the i-th eigenvalue in the target first feature matrix, with a range of 1 to n, where n is greater than 1.

[0101] After stretching the features of the second feature matrix output by the first neural network model, the eigenvalue range of the second feature matrix can be adjusted to be consistent with that of the target first feature matrix. For example, the eigenvalue range of the second feature matrix is 0-100, and the eigenvalue range of the target first feature matrix is 10-20. After feature stretching here, the eigenvalue range of the second feature matrix can also be adjusted to 10-20, which is convenient for calculating the similarity between the two in subsequent calculations.

[0102] After feature stretching, correspondingly, determining the target class label corresponding to the training image according to the similarity between the second feature matrix and the target first feature matrix in step 108 above includes:

[0103] Step 206, determining the target class label corresponding to the training image according to the similarity between the second feature matrix after feature stretching and the target first feature matrix.

[0104] That is to say, in this step, after obtaining the second feature matrix after feature stretching, the similarity between the second feature matrix after feature stretching and the target first feature matrix can be directly calculated, and then the target class label corresponding to the training image can be determined according to the similarity; or, after obtaining the second feature matrix after feature stretching, the second feature matrix after feature stretching can also be converted into a low-dimensional feature. For example, the second feature matrix after feature stretching is converted from B, C, H, W to a low-dimensional feature B, C*H*W, and then the similarity between the low-dimensional feature and the target first feature matrix is calculated, and then the target class label corresponding to the training image is determined according to the similarity.

[0105] Furthermore, when calculating the similarity, in order to avoid the influence of outliers and other abnormal values on the similarity, the second feature matrix after feature stretching and converted into a low-dimensional feature can be processed to remove the centroid points. As an optional embodiment, the mean and standard deviation corresponding to the second feature matrix after feature stretching can be calculated according to the second feature matrix after feature stretching; the preset maximum number of digits and minimum number of digits are obtained, and the second feature matrix after feature stretching is removed from outliers according to the maximum number of digits, minimum number of digits, mean, and standard deviation to determine the second feature matrix after outlier removal; the target class label corresponding to the training image is determined according to the similarity between the second feature matrix after outlier removal and the target first feature matrix.

[0106] Among them, after obtaining the second feature matrix after feature stretching and converted into a low-dimensional feature, the mean values of multiple eigenvalues therein can be calculated, and the standard deviation can be calculated based on the calculated mean value. At the same time, the preset minimum number of digits Q1 and maximum number of digits Q2 (the sizes of Q1 and Q2 can be set according to the actual situation) can be obtained, and then the upper and lower boundaries of the outlier points are calculated through the following formula:

[0107] 。

[0108] Among them, Lower Bound represents the lower boundary, Upper Bound represents the upper boundary, α is an empirical value, such as 1.5, x represents the eigenvalue in the second feature matrix after feature stretching and conversion to low-dimensional features, μ represents the mean, represents the standard deviation.

[0109] The upper and lower boundaries of the eigenvalue calculated above can be used to obtain the eigenvalue range [Lower Bound, UpperBound]. For the eigenvalue in the second feature matrix after feature stretching and conversion to low-dimensional features, if it exceeds the eigenvalue range [Lower Bound, Upper Bound], it is removed, otherwise it is retained. Through such operations, the second feature matrix after finally removing outliers and other abnormal values can be obtained. Then, the similarity between the second feature matrix after outlier removal and the target first feature matrix can be calculated, and the target class label of the training image can be determined based on the calculated similarity.

[0110] The calculation of the above similarity can be performed, for example, using the following formula:

[0111] 。

[0112] Among them, r is the calculated similarity; x i ’ represents the i-th eigenvalue in the second feature matrix after outlier removal, with a range of 1 to n, where n is greater than 1; represents the mean of the eigenvalues in the second feature matrix after outlier removal; Ci represents the i-th eigenvalue in the target first feature matrix; represents the mean of the eigenvalues in the target first feature matrix.

[0113] In this embodiment, the second feature matrix is feature-stretched by the maximum and minimum values in the second feature matrix and the target first feature matrix respectively, so as to stretch the eigenvalues to the same eigenvalue range as the target first feature matrix, that is, within a fixed space. This can facilitate the calculation of the similarity between the second feature matrix and the target first feature matrix, improve the accuracy of the calculated similarity, and further improve the accuracy of determining the target class label of the training image subsequently. In addition, by performing outlier removal processing on the second feature matrix after feature stretching, such as removing outliers, the accuracy of the calculated similarity can be further improved, and the accuracy of determining the target class label of the training image subsequently can be further improved.

[0114] In the above embodiments, the process of calculating the similarity and determining the target class label of the training image through multiple target clustering parameters of the preset clustering algorithm is illustrated. The following embodiments will illustrate the determination process of multiple target clustering parameters of the preset clustering algorithm.

[0115] Figure 3 FIG. 3 is a third schematic flowchart of the method for determining the label of the training image provided by the embodiment of the present invention. As Figure 3 shown, the above step 102, "obtaining multiple target clustering parameters corresponding to the preset clustering algorithm", may include the following steps:

[0116] Step 302, dividing multiple training images into multiple sets of training images according to their original class labels; each set of training images includes training images under one class label.

[0117] Among them, after initially obtaining multiple training images and obtaining the original class label corresponding to each training image, the training images with the same class label can be aggregated together according to the original class label of each training image, so that multiple sets of training images corresponding to their respective class labels can be obtained. The multiple class labels here are the same as the total number of classes corresponding to the original class labels. For example, if the total number of classes is 3, then 3 sets of training images can be obtained after classification, and each image set corresponds to the training images under one class label.

[0118] Step 304, performing distance calculation processing on the training images in each set of training images respectively to determine multiple target clustering parameters corresponding to the preset clustering algorithm.

[0119] In this step, after determining the preset clustering algorithm, the clustering parameters corresponding to the preset clustering algorithm can be obtained, and then the distance calculation can be performed through the above-mentioned classified training image sets with multiple class labels to determine the target values of the multiple clustering parameters of the preset clustering algorithm, that is, to determine multiple target clustering parameters. For example, if the preset clustering algorithm is the density clustering algorithm, the corresponding clustering parameters may include parameters such as the neighborhood radius eps, the minimum number of samples min_sample, and the distance metric threshold metric. Among them, the neighborhood radius can determine whether the points within the radius are regarded as core points, the minimum number of samples represents the minimum number of samples required to become a core point, and the distance metric threshold represents the metric threshold for calculating the distance between sample points.

[0120] To better illustrate how to obtain multiple target clustering parameters of the preset clustering algorithm through distance calculation, the following will take the density clustering algorithm as an example of the preset clustering algorithm. Optionally, the multiple target clustering parameters corresponding to the density clustering algorithm may include the target neighborhood radius, the target minimum number of samples, and the target distance metric threshold. The calculation processes of these three target clustering parameters will be described below respectively.

[0121] Optionally, the above-mentioned multiple target clustering parameters include a target neighborhood radius. The process of calculating the target neighborhood radius may include:

[0122] For each category label, input the training image set corresponding to the category label into the second neural network model. Calculate the distance between every two training images in the second neural network model to obtain a plurality of first distances corresponding to the training image set, and determine the target neighborhood radius under the category label corresponding to the training image set according to the plurality of first distances.

[0123] Among them, after the above classification, a training image set with multiple category labels can be obtained. Hereinafter, a training image set with one category label is taken as an example for illustration. For this category label, a neural network model based on an arbitrary distance-related loss function can be pre-created, denoted as the second neural network model, and the second neural network model can be pre-trained.

[0124] Then, each training image in the training image set corresponding to the category label can be input into the second neural network model. Calculate the distance between every two training images in the second neural network model to obtain a plurality of first distances and output them; or when there are too many training images in the training image set corresponding to the category label, the training images can also be input into the second neural network model in batches for pairwise distance calculation, and then the plurality of distances calculated for each batch can be averaged to obtain the average distance corresponding to each batch. In this way, a plurality of average distances can be obtained for multiple batches, all denoted as first distances. Then, find the inflection point among the plurality of first distance values, and use the first distance corresponding to the inflection point as the target neighborhood radius corresponding to the category label. Here, the inflection point can be the first distance at which the data mutates among the plurality of first distances, or it can also be the largest first distance among the plurality of first distances. For example, if the plurality of first distances are 0.1, 0.13, 0.10, 0.2, 0.15, 0.11, then the inflection point is 0.2, that is, the target neighborhood radius corresponding to the category label can be 0.2.

[0125] Through the above distance calculation method, the similarity between the training data under the same category label can be reflected, so as to quickly determine the corresponding neighborhood radius.

[0126] It should be noted that each category label can obtain a corresponding target neighborhood radius. For other category labels, their corresponding target neighborhood radii can also be calculated in the above manner.

[0127] Optionally, the above-mentioned multiple target clustering parameters include a target minimum number of samples. The process of calculating the target minimum number of samples may include:

[0128] For each type of category label, the training images in the training image set corresponding to the category label are input into the third neural network model in different batches for distance calculation, the model evaluation metrics corresponding to different batches are determined, and according to the model evaluation metrics of different batches, the target batch is determined among different batches and the number of training images corresponding to the target batch is used as the target minimum number of samples under the category label corresponding to the training image set.

[0129] Among them, for a certain category label, another neural network model based on an arbitrary distance-related loss function can be pre-created, denoted as the third neural network model, and the third neural network model can be pre-trained.

[0130] Then, the training images in the training image set corresponding to the category label can be input into the third neural network model in different batches respectively for distance calculation to obtain the distances corresponding to the training images in different batches. Then, after each different batch of training images is iterated for a certain number of rounds, the model evaluation metrics corresponding to the training images in the corresponding batch are calculated. Here, the model evaluation metrics are not limited to loss functions, distance metric functions, and accuracy metrics, etc. It can be understood that the model evaluation metrics of the training images in different batches are represented by the same metric, which is convenient for comparison with each other. The above-mentioned training images in different batches refer to different numbers of training images input into the third neural network model in each batch. For example, it can be from 1 to N, where N is the total number of training images in the training image set corresponding to the category label, and N is greater than 1.

[0131] After that, the model evaluation metrics of the training images in different batches can be compared, and a batch with the optimal model evaluation metric (i.e., the optimal convergence) is selected and denoted as the target batch, and the number of training images corresponding to the input of the target batch is used as the target minimum number of samples corresponding to the category label.

[0132] It should be noted that each category label can obtain a corresponding target minimum number of samples. For other category labels, their corresponding target minimum number of samples can also be calculated in the above manner.

[0133] Optionally, the above-mentioned multiple target clustering parameters include a target distance metric threshold, and the process of calculating the target distance metric threshold can include the following steps:

[0134] For each type of category label, calculate the distance between each training image in the training image set corresponding to the category label and its own original category label to obtain a plurality of second distances;

[0135] Eliminate the second distances with relatively large differences among the plurality of second distances, and select a plurality of candidate second distances with the top-ranked median in the eliminated second distances;

[0136] An average value corresponding to multiple candidate second distances is calculated, a target second distance is determined from the multiple candidate second distances according to the average value, and the target second distance is determined as a target distance measurement threshold under a category label corresponding to the training image set.

[0137] Among them, for a certain category label, the distance between the data of each training image in the training image set corresponding to the category label and the data corresponding to its original category label can be calculated to obtain the distance calculated for each training image, which is recorded as the second distance. The calculation formula of the second distance is as follows:

[0138] .

[0139] Among them, l represents the training image data itself, p represents the data corresponding to the original category label, S represents the covariance matrix, T represents the transpose, and d(l,p) represents the second distance of each training image.

[0140] After obtaining multiple second distances, the multiple second distances can be sorted in ascending order, and one or more second distances ranked at the top in the sorting result (such as the first two second distances in the sorting result) and one or more distances ranked at the bottom (such as the last two second distances in the sorting result) are eliminated, that is, the second distances with large differences among the multiple second distances are eliminated. Then, multiple (such as three or other numbers) second distances ranked at the top in the majority digit can be selected from the remaining multiple second distances, and all are recorded as candidate second distances. Here, the majority digit selection refers to selecting several second distances with a large number of occurrences from multiple second distances.

[0141] Then, the mean of the multiple candidate second distances can be calculated to obtain the average value, and a candidate second distance closest to the average value is selected from the multiple candidate second distances as the target distance measurement threshold corresponding to the final category label.

[0142] It should be noted that each category label can obtain a corresponding target distance measurement threshold, and for other category labels, their corresponding target distance measurement thresholds can also be calculated in the above manner.

[0143] In addition, after obtaining the above three target clustering parameters, a preset clustering algorithm may be used to cluster the multiple training images to obtain multiple first feature matrices.

[0144] Furthermore, although the above is an explanation of the process of calculating the target clustering parameters for the density clustering algorithm, for clustering parameters of other clustering algorithms, a suitable method may be selected in this manner to calculate and determine the target clustering parameters.

[0145] In this embodiment, after classifying multiple training images according to their original class labels, distance calculations are performed on the training image sets under each class label to determine the target clustering parameters corresponding to each class label, and then clustering is performed. Compared with the prior art in which clustering parameters are determined based on empirical values or random values, the present application determines clustering parameters based on training images and original class label information, which can improve the reliability of subsequent clustering results, and further improve the reliability of the target class labels of the finally determined training images. In addition, each clustering parameter in the preset clustering algorithm can be calculated according to different methods, which can improve the accuracy of the determined target clustering parameters, further improve the reliability of subsequent clustering results, and the reliability of the target class labels of the finally determined training images.

[0146] The label determination device for training images provided by the present invention will be described below. The label determination device for training images described below can be correspondingly referred to the label determination method for training images described above.

[0147] Figure 4 is a schematic structural diagram of the label determination device for training images provided by an embodiment of the present invention. Refer to Figure 4 As shown, the device may include:

[0148] An acquisition module 410, configured to acquire a plurality of pre-annotated training images and a plurality of target clustering parameters corresponding to a preset clustering algorithm; each training image includes its own corresponding original class label, and the above-mentioned plurality of target clustering parameters are pre-determined according to a plurality of training images and the original class label of each training image;

[0149] A clustering module 420, configured to cluster a plurality of training images according to the total number of class labels corresponding to the original class labels and each target clustering parameter, and determine a plurality of first feature matrices corresponding to the plurality of training images; the number of the above-mentioned plurality of first feature matrices is the same as the total number of class labels, and each first feature matrix corresponds to the features of the training images under one class label;

[0150] A model processing module 430, configured to input each training image into a first neural network model for classification processing to determine a second feature matrix corresponding to each training image;

[0151] A target label determination module 440, configured to, for each training image, obtain a target first feature matrix corresponding to the original class label of the training image from a plurality of first feature matrices, and determine a target class label corresponding to the training image according to the similarity between the second feature matrix and the target first feature matrix; the target class label is used to train the first neural network model.

[0152] In some embodiments, the above-mentioned target category label is represented by a label vector, and the label vector includes the probabilities corresponding to various category labels in the total categories. The above-mentioned target label determination module 440 is specifically configured to, if the similarity between the second feature matrix and the target first feature matrix is within the first threshold range, set the probability corresponding to the original category label of the training image in the label vector to 1 and set the probabilities corresponding to other category labels to 0, so as to obtain the target category label corresponding to the training image; the above-mentioned other category labels are other category labels in the total categories except the original category label of the training image; if the similarity between the second feature matrix and the target first feature matrix is within the second threshold range, set the probability corresponding to the original category label of the training image in the label vector to the first probability and set the probability corresponding to the predicted category of the training image to the second probability, so as to obtain the target category label corresponding to the training image; the sum of the above-mentioned first probability and the second probability is 1; wherein, the similarity corresponding to the first threshold range is greater than the similarity corresponding to the second threshold range.

[0153] Optionally, before the above-mentioned target label determination module 440 determines the target category label corresponding to the training image according to the similarity between the second feature matrix and the target first feature matrix, the above-mentioned device further includes:

[0154] A feature stretching module, configured to obtain the maximum value and the minimum value in the second feature matrix, and obtain the maximum value and the minimum value in the target first feature matrix; perform feature stretching processing on the second feature matrix according to the maximum value and the minimum value in the second feature matrix and the maximum value and the minimum value in the target first feature matrix, so as to determine the second feature matrix after feature stretching; the eigenvalue range corresponding to the stretched second feature matrix is the same as that of the target first feature matrix;

[0155] Correspondingly, the above-mentioned target label determination module 440 is specifically configured to determine the target category label corresponding to the training image according to the similarity between the second feature matrix after feature stretching and the target first feature matrix.

[0156] Optionally, the above-mentioned target label determination module 440 is specifically configured to calculate the mean value and the standard deviation corresponding to the second feature matrix after feature stretching according to the second feature matrix after feature stretching; obtain the preset maximum number of digits and minimum number of digits, and perform outlier removal on the second feature matrix after feature stretching according to the maximum number of digits, the minimum number of digits, the mean value, and the standard deviation, so as to determine the second feature matrix after outlier removal; determine the target category label corresponding to the training image according to the similarity between the second feature matrix after outlier removal and the target first feature matrix.

[0157] In some embodiments, the above-mentioned acquisition module 410 includes:

[0158] A classification unit for classifying multiple training images into multiple sets of training images according to their respective original class labels; each set of training images includes training images under one class label.

[0159] A clustering parameter determination unit for respectively performing distance calculation processing on the training images in each set of training images to determine multiple target clustering parameters corresponding to a preset clustering algorithm.

[0160] Optionally, the above multiple target clustering parameters include a target neighborhood radius and a target minimum sample number. The clustering parameter determination unit is specifically configured to, for each class label, input the set of training images corresponding to the class label into a second neural network model, calculate the distance between every two training images in the second neural network model, obtain multiple first distances corresponding to the set of training images, and determine the target neighborhood radius under the class label corresponding to the set of training images according to the multiple first distances.

[0161] For each class label, input the training images in the set of training images corresponding to the class label into a third neural network model in different batches for distance calculation, determine the model evaluation indexes corresponding to different batches respectively, and determine the target batch according to the model evaluation indexes of different batches respectively, and use the number of training images corresponding to the target batch as the target minimum sample number under the class label corresponding to the set of training images.

[0162] Optionally, the above multiple target clustering parameters include a target distance metric threshold. The clustering parameter determination unit is specifically configured to, for each class label, calculate the distance between each training image in the set of training images corresponding to the class label and its own original class label, obtain multiple second distances; remove the second distances with relatively large differences among the multiple second distances, and select multiple candidate second distances with a higher median ranking from the removed second distances; calculate the average value corresponding to the multiple candidate second distances, determine the target second distance from the multiple candidate second distances according to the average value, and determine the target second distance as the target distance metric threshold under the class label corresponding to the set of training images.

[0163] It should be noted here that the above device provided by the embodiment of the present invention can implement all the method steps implemented by the above method embodiment, and can achieve the same technical effect. The same parts and beneficial effects as those in the method embodiment will not be specifically described in this embodiment.

[0164] Figure 5 The structural schematic diagram of the electronic device provided by the embodiment of the present invention is as Figure 5As shown in the figure, the electronic device may include: a processor 510, a communications interface 520, a memory 530, and a communication bus 540. Among them, the processor 510, the communications interface 520, and the memory 530 complete communication with each other through the communication bus 540. The processor 510 may call logic instructions in the memory 530 to execute a method for determining labels of training images. The method includes: obtaining a plurality of pre-annotated training images and obtaining a plurality of target clustering parameters corresponding to a preset clustering algorithm; each training image includes its own corresponding original class label, and the above-mentioned plurality of target clustering parameters are determined in advance according to the plurality of training images and the original class label of each training image; clustering the plurality of training images according to the total number of classes corresponding to the original class label and each target clustering parameter to determine a plurality of first feature matrices corresponding to the plurality of training images; the number of the above-mentioned plurality of first feature matrices is the same as the total number of classes, and each first feature matrix corresponds to the features of the training images under one class label; inputting each training image into a first neural network model for classification processing to determine a second feature matrix corresponding to each training image; for each training image, obtaining a target first feature matrix corresponding to the original class label of the training image among the plurality of first feature matrices, and determining a target class label corresponding to the training image according to the similarity between the second feature matrix and the target first feature matrix; the target class label is used to train the first neural network model.

[0165] In addition, when the logic instructions in the above-mentioned memory 530 are implemented in the form of software functional units and sold or used as an independent product, they may be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, may be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk, or an optical disc that can store program codes.

[0166] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the method for determining the label of a training image provided by each of the above methods. The method includes: obtaining a plurality of pre-annotated training images and obtaining a plurality of target clustering parameters corresponding to a preset clustering algorithm; each training image includes its own corresponding original class label, and the above-mentioned plurality of target clustering parameters are determined in advance according to the plurality of training images and the original class label of each training image; clustering the plurality of training images according to the total number of class labels corresponding to the original class labels and each target clustering parameter to determine a plurality of first feature matrices corresponding to the plurality of training images; the number of the above-mentioned plurality of first feature matrices is the same as the total number of class labels, and each first feature matrix corresponds to the features of the training images under one class label; inputting each training image into a first neural network model for classification processing to determine a second feature matrix corresponding to each training image; for each training image, obtaining a target first feature matrix corresponding to the original class label of the training image from the plurality of first feature matrices, and determining a target class label corresponding to the training image according to the similarity between the second feature matrix and the target first feature matrix; the target class label is used to train the first neural network model.

[0167] On another aspect, the present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it realizes the method for determining the label of a training image provided by each of the above methods. The method includes: obtaining a plurality of pre-annotated training images and obtaining a plurality of target clustering parameters corresponding to a preset clustering algorithm; each training image includes its own corresponding original class label, and the above-mentioned plurality of target clustering parameters are determined in advance according to the plurality of training images and the original class label of each training image; clustering the plurality of training images according to the total number of class labels corresponding to the original class labels and each target clustering parameter to determine a plurality of first feature matrices corresponding to the plurality of training images; the number of the above-mentioned plurality of first feature matrices is the same as the total number of class labels, and each first feature matrix corresponds to the features of the training images under one class label; inputting each training image into a first neural network model for classification processing to determine a second feature matrix corresponding to each training image; for each training image, obtaining a target first feature matrix corresponding to the original class label of the training image from the plurality of first feature matrices, and determining a target class label corresponding to the training image according to the similarity between the second feature matrix and the target first feature matrix; the target class label is used to train the first neural network model.

[0168] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative work.

[0169] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0170] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features. And these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for determining labels of training images, characterized in that: include: Acquire a plurality of pre-labeled training images and obtain a plurality of target clustering parameters corresponding to a preset clustering algorithm; Each of the training images includes its own corresponding original category label, and the plurality of target clustering parameters are determined in advance according to the plurality of training images and the original category label of each of the training images; Clustering the plurality of training images according to the total number of categories corresponding to the original category labels and the target clustering parameters, and determining a plurality of first feature matrices corresponding to the plurality of training images; the number of the plurality of first feature matrices is the same as the total number of categories, and each first feature matrix corresponds to features of all training images under one category label; Input each of the training images into the first neural network model for classification processing, and determine a second feature matrix corresponding to each of the training images; For each of the training images, obtaining a target first feature matrix corresponding to the original category label of the training image from the multiple first feature matrices, and determining whether the original category label of the training image needs to be corrected and determining the target category label corresponding to the training image based on the similarity between the second feature matrix and the target first feature matrix; The target category label is used to train the first neural network model, and the similarity represents the feature similarity between the training image and a group of training images corresponding to its original category label.

2. The method for determining labels of training images according to claim 1, characterized in that: The target category label of the training image is represented by a label vector, and the label vector includes the probabilities corresponding to various category labels in the total category. The determining the target category label corresponding to the training image according to the similarity between the second feature matrix and the target first feature matrix includes: Calculating a first similarity between the second feature matrix and the target first feature matrix; If the first similarity is within a first threshold range, the probability corresponding to the original category label of the training image in the label vector is set to 1 and the probability corresponding to other category labels is set to 0, so as to obtain the target category label corresponding to the training image; the other category label is the other category label in the total category except the original category label of the training image; If the first similarity exceeds the first threshold range, calculating the second similarity between the second feature matrix and other first feature matrices in the plurality of first feature matrices, and determining the target category label corresponding to the training image according to the second similarity and the second threshold range; the other first feature matrices are the remaining first feature matrices in the plurality of first feature matrices except the target first feature matrix; The similarity corresponding to the first threshold range is greater than the similarity corresponding to the second threshold range.

3. The method for determining labels of training images according to claim 1, characterized in that: Before determining the target category label corresponding to the training image according to the similarity between the second feature matrix and the target first feature matrix, the method further includes: Obtaining the maximum value and the minimum value in the second feature matrix, and obtaining the maximum value and the minimum value in the target first feature matrix; According to the maximum value and the minimum value in the second feature matrix and the maximum value and the minimum value in the target first feature matrix, the second feature matrix is ​​subjected to feature stretching processing to determine a second feature matrix after feature stretching; the second feature matrix after stretching has the same feature value range as that corresponding to the target first feature matrix; Accordingly, determining the target category label corresponding to the training image according to the similarity between the second feature matrix and the target first feature matrix includes: The target category label corresponding to the training image is determined according to the similarity between the second feature matrix after feature stretching and the target first feature matrix.

4. The method for determining labels of training images according to claim 3, characterized in that: The step of determining the target category label corresponding to the training image according to the similarity between the second feature matrix after feature stretching and the target first feature matrix includes: According to the second feature matrix after feature stretching, calculating the mean and standard deviation corresponding to the second feature matrix after feature stretching; Obtaining a preset maximum number of digits and a minimum number of digits, and removing outliers from the second feature matrix after feature stretching according to the maximum number of digits, the minimum number of digits, the mean, and the standard deviation, to determine the second feature matrix after the outliers are removed; The target category label corresponding to the training image is determined according to the similarity between the second feature matrix after the outliers are removed and the target first feature matrix.

5. The method for determining labels of training images according to any one of claims 1 to 4, characterized in that: The obtaining of a plurality of target clustering parameters corresponding to a preset clustering algorithm includes: Dividing the plurality of training images into multiple types of training image sets according to their respective original category labels; each type of training image set includes training images under one category label; Distance calculation processing is performed on each type of training images in the training image set to determine multiple target clustering parameters corresponding to the preset clustering algorithm.

6. The method for determining labels of training images according to claim 5, characterized in that: The multiple target clustering parameters include a target neighborhood radius and a target minimum sample number. The distance calculation process is performed on each type of training images in the training image set to determine the multiple target clustering parameters corresponding to the preset clustering algorithm, including: For each category label, a training image set corresponding to the category label is input into a second neural network model, a distance between every two training images is calculated in the second neural network model, a plurality of first distances corresponding to the training image set are obtained, and a target neighborhood radius under the category label corresponding to the training image set is determined according to the plurality of first distances; For each category label, the training images in the training image set corresponding to the category label are input into the third neural network model in different batches for distance calculation, and the model evaluation indicators corresponding to the different batches are determined. According to the model evaluation indicators of the different batches, the target batch is determined among the different batches, and the number of training images corresponding to the target batch is used as the target minimum number of samples under the category label corresponding to the training image set.

7. The method for determining labels of training images according to claim 5, characterized in that: The multiple target clustering parameters include a target distance measurement threshold, and the distance calculation process is performed on each type of training images in the training image set to determine the multiple target clustering parameters corresponding to the preset clustering algorithm, including: For each category label, calculate the distance between each training image in the training image set corresponding to the category label and its original category label to obtain multiple second distances; Eliminate second distances with relatively large differences among the plurality of second distances, and select a plurality of candidate second distances with the most-complex digits ranked top among the eliminated second distances; An average value corresponding to the multiple candidate second distances is calculated, a target second distance is determined from the multiple candidate second distances according to the average value, and the target second distance is determined as a target distance measurement threshold under the category label corresponding to the training image set.

8. A device for determining labels of training images, characterized in that: include: An acquisition module, used to acquire a plurality of pre-labeled training images and to acquire a plurality of target clustering parameters corresponding to a preset clustering algorithm; Each of the training images includes its own corresponding original category label, and the plurality of target clustering parameters are determined in advance according to the plurality of training images and the original category label of each of the training images; A clustering module, used for clustering the plurality of training images according to the total number of categories corresponding to the original category labels and each of the target clustering parameters, and determining a plurality of first feature matrices corresponding to the plurality of training images; the number of the plurality of first feature matrices is the same as the total number of categories, and each first feature matrix corresponds to features of all training images under one category label; A model processing module, used for inputting each of the training images into the first neural network model for classification processing, and determining a second feature matrix corresponding to each of the training images; a target label determination module, configured to obtain, for each of the training images, a target first feature matrix corresponding to the original category label of the training image from the plurality of first feature matrices, and determine, based on the similarity between the second feature matrix and the target first feature matrix, whether the original category label of the training image needs to be corrected and determine the target category label corresponding to the training image; The target category label is used to train the first neural network model, and the similarity represents the feature similarity between the training image and a group of training images corresponding to its original category label.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the method for determining a label of a training image according to any one of claims 1 to 7 is implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for determining a label of a training image according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Image classification and neural network training method and device, equipment and storage medium

    CN111259967A