Image recognition method and device, electronic equipment and storage medium
By using a feature extraction model trained by self-supervised contrast learning in image recognition, the problem of misrecognition of similar images within the class in image recognition is solved, and the accuracy and accuracy of recognition are improved.
Patent Information
- Application Number
- CN202311549790.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-17
- Publication Date
- 2025-05-20
AI Technical Summary
The prior art is difficult to effectively distinguish similar images within the classification in image recognition tasks, resulting in high error recognition rate and low accuracy.
An image recognition method is adopted to obtain the feature extraction model by inputting the strong enhancement sample data and weak enhancement sample data into the initial feature extraction model and the momentum update model, and performing self-supervised comparison learning training. This model can output features with high distinction and reduce the misidentification of similar images within the class.
The accuracy and accuracy of image recognition are improved, the error recognition rate of similar images within the class is reduced, and the feature extraction model can more effectively identify target objects in the image.
Smart Images

Figure CN120020899A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and in particular, to an image recognition method, apparatus, electronic device, and storage medium. Background Art
[0002] Image recognition, also known as image classification, is to use a computer to process, analyze, and understand an image to identify target objects of different patterns or categories.
[0003] In the image recognition task, the related technology uses a direct classification method for image classification, that is, a classification model is trained using sample images marked with target categories, and the classification model is directly used for image classification. This method cannot enable the model to learn features with large discrimination degrees. Especially in the fine-grained image recognition task, similar images within a class are likely to cause misrecognition, resulting in low accuracy of image recognition. Summary of the Invention
[0004] The present invention provides an image recognition method, apparatus, electronic device, and storage medium to solve the problem of misrecognition of similar images of the same type in the prior art and improve the accuracy of image recognition.
[0005] The present invention provides an image recognition method, including:
[0006] Obtain an image to be recognized, and input the image to be recognized into a feature extraction model to obtain a target feature vector output by the feature extraction model;
[0007] Based on the target feature vector, recognize the classification result of the target object in the image to be recognized;
[0008] Wherein, the feature extraction model is obtained by respectively inputting strongly augmented sample data and weakly augmented sample data into an initial feature extraction model and a momentum update model corresponding to the initial feature extraction model, and performing self-supervised contrastive learning training on the initial feature extraction model and the momentum update model; the strongly augmented sample data is obtained by performing strong augmentation processing on the first sample image corresponding to each target sample category; the weakly augmented sample data is obtained by performing weak augmentation processing on the second sample image corresponding to each target sample category.
[0009] According to an image recognition method provided by the present invention, the feature extraction model is trained through the following steps:
[0010] Randomly obtain the first training sample images of each training batch from the training dataset, and the first training sample images include the first sample images and the second sample images corresponding to each target sample category;
[0011] Perform strong enhancement processing on the first sample image to obtain the strongly enhanced sample data;
[0012] Perform weak enhancement processing on the second sample image to obtain the weakly enhanced sample data;
[0013] Input both the strongly enhanced sample data and the weakly enhanced sample data into the initial feature extraction model and the momentum update model, and pair up the first sample feature vectors output by the initial feature extraction model and the second sample feature vectors output by the momentum update model one by one;
[0014] Determine the first feature similarity of each pair of paired sample feature vectors, and add a similarity label for identifying whether it is a feature of the same category to the first feature similarity;
[0015] Determine the loss function value based on the first feature similarity and the similarity label, and adjust the model parameters of the initial feature extraction model and the momentum update model based on the loss function value until the training ends, and determine the trained initial feature extraction model as the feature extraction model.
[0016] According to an image recognition method provided by the present invention, the randomly obtaining the first training sample image of each training batch from the training data set includes:
[0017] For each training batch, randomly sort all sample categories of the training data set, and determine the first preset number of sample categories after sorting as the target sample categories of the training batch;
[0018] Obtain two sample images from the sample images of each target sample category to obtain the first sample image and the second sample image corresponding to each target sample category.
[0019] According to an image recognition method provided by the present invention, the adjusting the model parameters of the initial feature extraction model and the momentum update model based on the loss function value includes:
[0020] Based on the loss function value, adjust the model parameters of the initial feature extraction model by backpropagation of gradients;
[0021] Based on the preset momentum update parameter and the adjusted model parameters of the initial feature extraction model, adjust the model parameters of the momentum update model.
[0022] According to an image recognition method provided by the present invention, the recognizing the classification result of the target object in the image to be recognized based on the target feature vector includes:
[0023] Input the target feature vector into the classification layer to obtain the classification probability output by the classification layer;
[0024] Determine the first classification category corresponding to the maximum classification probability in the classification probability as the classification result of the target object in the image to be recognized;
[0025] Among them, the classification layer is obtained by separately training the initial classification layer based on the second training sample image, the category label corresponding to the second training sample image, and the feature extraction model.
[0026] According to an image recognition method provided by the present invention, the recognizing the classification result of the target object in the image to be recognized based on the target feature vector includes:
[0027] Determine the second feature similarity between the target feature vector and each feature vector in the obtained first template feature vector;
[0028] Sort the second feature similarities in descending order, and obtain the second classification categories corresponding to the first second preset number of feature vectors in the sorting result;
[0029] Determine the classification category with the largest number in the second classification categories as the classification result of the target object in the image to be recognized.
[0030] According to an image recognition method provided by the present invention, it further includes:
[0031] Input the test sample image into the feature extraction model to obtain the first feature vector corresponding to each test classification category output by the feature extraction model;
[0032] Determine the third feature similarity between the feature vectors in the first feature vector, and determine the first third preset number of feature vectors corresponding in the first feature vector in ascending order of the third feature similarity to obtain the second template feature vector;
[0033] For each template feature vector in the second template feature vector, determine the feature center vector based on the fourth feature similarity between the template feature vector and the second feature vector, and update the template feature vector with the feature center vector to obtain the third template feature vector after updating the second template feature vector; the second feature vector is the other feature vectors in the first feature vector except the second template feature vector;
[0034] Determine the third template feature vectors corresponding to each test classification category as the first template feature vector.
[0035] The present invention also provides an image recognition device, including:
[0036] A feature extraction module, configured to obtain an image to be recognized and input the image to be recognized into a feature extraction model, so as to obtain a target feature vector output by the feature extraction model;
[0037] An identification module, configured to identify a classification result of a target object in the image to be recognized based on the target feature vector;
[0038] Wherein, the feature extraction model is obtained by respectively inputting strongly augmented sample data and weakly augmented sample data into an initial feature extraction model and a momentum update model corresponding to the initial feature extraction model, and performing self-supervised contrastive learning training on the initial feature extraction model and the momentum update model; the strongly augmented sample data is obtained by performing strong augmentation processing on a first sample image corresponding to each target sample category; the weakly augmented sample data is obtained by performing weak augmentation processing on a second sample image corresponding to each target sample category.
[0039] The present invention further provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, where when the processor executes the computer program, the image recognition method as described in any one of the above is implemented.
[0040] The present invention further provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the image recognition method as described in any one of the above is implemented.
[0041] The present invention further provides a computer program product, including a computer program, and when the computer program is executed by a processor, the image recognition method as described in any one of the above is implemented.
[0042] The image recognition method, device, electronic device, and storage medium provided by the present invention input the acquired image to be recognized into a feature extraction model to obtain a target feature vector output by the feature extraction model, and then identify the classification result of the target object in the image to be recognized based on the extracted target feature vector, thereby realizing the recognition of the image to be recognized. Among them, the feature extraction model is obtained by respectively inputting strongly augmented sample data and weakly augmented sample data into an initial feature extraction model and a momentum update model corresponding to the initial feature extraction model, and performing self-supervised contrast learning training on the initial feature extraction model and the momentum update model. The strongly augmented sample data is obtained by strongly augmenting the first sample image corresponding to each target sample category, and the weakly augmented sample data is obtained by weakly augmenting the second sample image corresponding to each target sample category. In this way, during the model training process, contrast learning between sample images of the same class and contrast learning between models can be carried out, thereby improving the feature extraction ability of the feature extraction model, enabling the trained feature extraction model to output features with higher discrimination, reducing the misrecognition of similar images within the class, and further improving the accuracy and precision of image recognition. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0044] Figure 1 is a flowchart of the image recognition method provided by an embodiment of the present invention;
[0045] Figure 2 is a schematic diagram of the training principle of the feature extraction model in an embodiment of the present invention;
[0046] Figure 3 is a flowchart of the training method of the feature extraction model in an embodiment of the present invention;
[0047] Figure 4 is a schematic diagram of the structure of the image recognition device provided by an embodiment of the present invention;
[0048] Figure 5 is a schematic diagram of the structure of the electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0049] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below with reference to the accompanying drawings in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without making creative efforts shall fall within the protection scope of the present invention.
[0050] It should be noted that the serial numbers assigned to the objects described in the present invention itself, such as "first", "second", etc., are only used to distinguish the described objects and do not have any sequential or technical meanings.
[0051] The following Figures 1-3 describes the image recognition method of the present invention. This image recognition method can be applied to electronic devices such as terminal devices or servers. Among them, terminal devices can include mobile phones, computers, in-vehicle devices, tablet computers, wearable devices, smart home devices, etc.; servers can include independent servers, cluster servers or cloud servers, etc. This image recognition method can also be applied to an image recognition device provided in an electronic device such as a terminal device or a server, and this image recognition device can be implemented by software, hardware or a combination of both.
[0052] Figure 1 Exemplarily shown is a schematic flowchart of the image recognition method provided by an embodiment of the present invention. Referring to Figure 1 as shown, this image recognition method may include the following steps 110 to 120.
[0053] Step 110: Obtain an image to be recognized, and input the image to be recognized into a feature extraction model to obtain a target feature vector output by the feature extraction model.
[0054] Among them, the feature extraction model is obtained by respectively inputting strongly augmented sample data and weakly augmented sample data into an initial feature extraction model and a momentum update model corresponding to the initial feature extraction model, and performing self-supervised contrast learning training on the initial feature extraction model and the momentum update model. Among them, performing self-supervised contrast learning training on the initial feature extraction model and the momentum update model means performing self-supervised contrast learning training on the initial feature extraction model in comparison with the momentum update model, and the feature extraction model is the trained initial feature extraction model. The network structure of the momentum update model is the same as that of the initial feature extraction model, and the momentum update model is the momentum weighting of the initial feature extraction model, that is, the model parameters of the momentum update model follow the preset momentum update parameters and the model parameters of the initial feature extraction model for updating.
[0055] Exemplarily, the strongly augmented sample data and the weakly augmented sample data can be input into the initial feature extraction model to obtain a first sample feature vector output by the initial feature extraction model. The first sample feature vector can include a first sub-sample feature vector corresponding to the strongly augmented sample data and a second sub-sample feature vector corresponding to the weakly augmented sample data. The strongly augmented sample data and the weakly augmented sample data can be input into the momentum update model to obtain a second sample feature vector output by the momentum update model. The second sample feature vector can include a third sub-sample feature vector corresponding to the strongly augmented sample data and a fourth sub-sample feature vector corresponding to the weakly augmented sample data. Then, the first sub-sample feature vector, the second sub-sample feature vector, the third sub-sample feature vector, and the fourth sub-sample feature vector can be pairwise compared between the models to determine the similarity between the paired sample feature vectors. Furthermore, the loss function value can be determined based on the similarity, and the model parameters of the initial feature extraction model can be adjusted using the loss function value. And the model parameters of the momentum update model can be adjusted based on the adjusted model parameters until the model converges, the training is ended, and the trained initial feature extraction model is determined as the feature extraction model.
[0056] For example, a similarity label for identifying whether it is a same-class feature can be added to the similarity calculated by pairwise pairing, and then the loss function value can be determined based on the similarity and the similarity label.
[0057] Exemplarily, the cross-entropy loss function can be selected as the loss function.
[0058] Among them, the strongly augmented sample data is obtained by strongly augmenting the first sample image corresponding to each target sample category, and the weakly augmented sample data is obtained by weakly augmenting the second sample image corresponding to each target sample category. It can be understood that the strongly augmented sample data and the weakly augmented sample data contain the same classification types. The classification types of each sample image in each type of augmented sample data are different, and the classification types of the sample images between the two types of augmented sample data correspond one by one, that is, two sample images of the same classification type are respectively divided into the strongly augmented sample data and the weakly augmented sample data.
[0059] Exemplarily, the electronic device can perform random batch sampling on the training data set to obtain the sample images of each training batch, and the sample categories of each training batch are all different. Specifically, in the sampling process of each training batch, the sample images of the same target sample category can be collected twice, so that each target sample category includes two sample images. Then, one of the sample images is strongly augmented and the other sample image is weakly augmented to obtain the strongly augmented sample data and the weakly augmented sample data in this way. In this way, sample data of two different fields of view, different distortions, and different deformation types can be obtained.
[0060] Among them, the training data set may include the CUB-200-2011 data set, the Cars-196 data set, and the Stanford Online Products (SOP) data set, but is not limited thereto.
[0061] The weak augmentation processing may include random scaling and cropping, horizontal mirroring, and normalization processing of the image. The strong augmentation processing may include random scaling and cropping, horizontal mirroring, color jittering, random grayscaling, adding Gaussian noise, overexposure, and normalization processing of the image. It should be noted that the methods of weak augmentation processing and strong augmentation processing here are only for illustrative purposes and do not constitute the sole limitation of the present invention. It can be understood that the strong augmentation processing performs more data augmentation processing on the image than the weak augmentation processing.
[0062] Step 120: Identify the classification result of the target object in the image to be recognized based on the target feature vector.
[0063] The extracted target feature vector can characterize the features of the target object in the image to be recognized, and the target object can be recognized using this target feature vector. For example, the extracted target feature vector can be input into a pre-trained classification layer, and the classification layer is used to determine the classification result of the target object in the image to be recognized. Alternatively, the extracted target feature vector can be compared with a pre-constructed template feature vector to match the classification result of the target object in the image to be recognized, where the template feature vector is used to characterize the correspondence between the classification category and the feature vector, that is, the feature vector of the known classification category.
[0064] The image recognition method provided by the embodiments of the present invention inputs the obtained image to be recognized into a feature extraction model to obtain the target feature vector output by the feature extraction model, and then identifies the classification result of the target object in the image to be recognized based on the extracted target feature vector, realizing the recognition of the image to be recognized. Among them, the feature extraction model is obtained by inputting both the strongly augmented sample data and the weakly augmented sample data into the initial feature extraction model and the momentum update model corresponding to the initial feature extraction model respectively, and performing self-supervised contrastive learning training on the initial feature extraction model and the momentum update model. The strongly augmented sample data is obtained by performing strong augmentation processing on the first sample image corresponding to each target sample category, and the weakly augmented sample data is obtained by performing weak augmentation processing on the second sample image corresponding to each target sample category. In this way, during the model training process, contrastive learning between similar sample images of the same class and contrastive learning between models can be performed, thereby improving the feature extraction ability of the feature extraction model, enabling the trained feature extraction model to output features with higher discrimination, reducing the misrecognition of similar images within the class, and further improving the accuracy and precision of image recognition.
[0065] Based on Figure 1 the corresponding embodiment of the image recognition method, Figure 2An exemplary schematic diagram of the training principle of the feature extraction model in an embodiment of the present invention is shown. Referring to Figure 2 as shown, the first training sample images of each training batch can be collected from the training dataset to obtain two images of different sample categories, namely the first sample image and the second sample image of each sample category. Then, strong augmentation processing is performed on the first sample image, and weak augmentation processing is performed on the second sample image to obtain strongly augmented sample data and weakly augmented sample data. Then, both types of augmented data are respectively input into the encoding network Q and the encoding network K for feature extraction, and the extracted feature vectors are paired and compared between the encoding networks to determine the similarity between the paired feature vectors. Furthermore, based on this similarity, the loss function value is determined, and based on the loss function value, the model parameters of the encoding network Q and the encoding network K are adjusted until the model converges and the training ends. The trained encoding network Q is the required feature extraction model. Among them, the encoding network Q is the initial feature extraction model, the encoding network K is the momentum update model, and the network structures of the encoding network Q and the encoding network K are the same. For example, both can adopt a residual neural network (ResNet), such as ResNet50.
[0066] Specifically, in combination with Figure 2 , Figure 3 An exemplary schematic diagram of the training method of the feature extraction model in an embodiment of the present invention is shown. It is a self-supervised contrastive learning training method. Referring to Figure 3 as shown, the feature extraction model can be trained through the following steps 310 to 360.
[0067] Step 310: Randomly obtain the first training sample images of each training batch from the training dataset. The first training sample images include the first sample image and the second sample image corresponding to each target sample category.
[0068] Among them, the training dataset can include the CUB-200-2011 dataset, the Cars-196 dataset, and the Stanford Online Products (SOP) dataset, but is not limited thereto.
[0069] For each training batch, a batch of samples can be randomly collected from the training dataset to ensure that the categories of each Batch training sample image are different, and each category in each Batch includes 2 sample images.
[0070] Specifically, randomly obtaining the first training sample images of each training batch from the training dataset may include: for each training batch, randomly sorting all the sample categories of the training dataset, and determining the first preset number of sample categories at the front of the sorted sample categories as the target sample categories of this training batch; obtaining two sample images from the sample images of each target sample category, to obtain the first sample image and the second sample image corresponding to each target sample category. Wherein, the first preset number is less than the number of all the sample categories of the training dataset.
[0071] Exemplarily, the target sample categories of the training batch may be determined from the sorted sample categories based on the number of processors used during model training. It can be understood that the more the number of processors, the larger the amount of data that can be processed, and the corresponding first preset number is also larger. For example, in the case where the number of processors is 1, that is, in the single-card network training mode, the first sample categories of the first preset number at the front of the sorted sample categories may be determined as the target sample categories of the batch. In the case where the number of processors is greater than 1, that is, in the multi-card network training mode, according to the order of the sorted sample categories, every fourth preset number of second sample categories, the second sample categories are allocated to the corresponding processors, and the second sample categories allocated by each processor are determined as the target sample categories of this training batch, where the processors and the second sample categories are in one-to-one correspondence, then the first preset number is the product of the fourth preset number and the number of processors. Wherein, the processors used during model training may be, for example, Graphics Processing Units (GPUs).
[0072] For example, before model training, the image paths, image indexes, and sample category labels of the entire training dataset may be read first, the sample category labels are counted to obtain the number of sample categories of the entire training dataset, and the image indexes with the same category labels are classified into the same list, to implement the classification of each sample image in the training dataset. Wherein, the image path is used to indicate the storage location of the sample image; the image index may be used to identify the sample image, that is, to distinguish different sample images; the sample category label is used to identify the classification category.
[0073] In the single - card network training mode, assuming that the number of categories in the training dataset is m, the sample category labels can be set as [0, m), that is, labeled as 0 to m - 1. In each training batch, these m sample category labels can be shuffled, and the first n sample categories are selected as the target sample categories of the current batch. For each of the first n sample categories selected, 2 image indices are randomly selected from its corresponding image index list as the sample indices selected for this sample category, and then the corresponding sample images are read according to the selected sample indices as the first training sample images of the current batch. Among them, n < m, the first preset quantity N in the single - card network training mode is N = n, and the number of classification categories of the first training sample images of the current batch is also n. By repeating this process, the first training sample images can be obtained, and these m sample category labels need to be shuffled before sampling in each training batch to ensure that the training sample images with different sample categories can be randomly collected in each training batch.
[0074] In the multi - card network training mode, assuming that a total of d processors are used to train the model, the number of categories in the training dataset is m, and the sample category labels are marked as [0, m). In each training batch, these m sample category labels can be randomly sorted first, and the first N sample categories are distributed to each processor. n sample categories can be allocated to each processor, so N = d * n < m, and the first preset quantity N in the multi - card network training mode is N = d * n. Specifically, for the k - th processor, the sample categories in the interval from k * n to (k + 1) * n can be selected from the sorted sample categories as the sample categories of the k - th processor in the current batch. For each sample category in the interval from k * n to (k + 1) * n, 2 image indices are randomly selected from the image index list corresponding to this sample category as the sample indices selected for this sample category, and then the corresponding sample images are read according to the selected sample indices as the training sample images on the k - th processor in the current batch, where k ∈ [0, d). In this way, the training sample images on each processor can be obtained. By repeating this process, the first training sample images of the current batch can be obtained, and these m sample category labels need to be shuffled before sampling in each batch to ensure that the training sample images with different sample categories can be randomly collected in each batch.
[0075] Step 320: Perform strong enhancement processing on the first sample image to obtain strongly enhanced sample data.
[0076] Step 330: Perform weak enhancement processing on the second sample image to obtain weakly enhanced sample data.
[0077] After sampling in step 310, each sample category of the first training sample image contains a first sample image and a second sample image, a total of 2 sample images. The first sample images of each sample category can be grouped into one data set for strong enhancement processing to obtain strongly enhanced sample data. The second sample images of each sample category can be grouped into one data set for weak enhancement processing to obtain weakly enhanced sample data.
[0078] Among them, the weak enhancement processing can include random scaling and cropping, horizontal mirroring, and normalization processing of the image. The strong enhancement processing can include random scaling and cropping, horizontal mirroring, color jittering, random grayscaling, adding Gaussian noise, overexposure, and normalization processing of the image.
[0079] Step 340: Input both the strongly enhanced sample data and the weakly enhanced sample data into the initial feature extraction model and the momentum update model, and pair up the first sample feature vectors output by the initial feature extraction model and the second sample feature vectors output by the momentum update model pairwise.
[0080] Among them, the first sample feature vectors output by the initial feature extraction model can include the first sub-sample feature vectors corresponding to the strongly enhanced sample data and the second sub-sample feature vectors corresponding to the weakly enhanced sample data. The second sample feature vectors output by the momentum update model can include the third sub-sample feature vectors corresponding to the strongly enhanced sample data and the fourth sub-sample feature vectors corresponding to the weakly enhanced sample data. The first sub-sample feature vectors and the second sub-sample feature vectors can be paired with the third sub-sample feature vectors and the fourth sub-sample feature vectors pairwise.
[0081] Step 350: Determine the first feature similarity of each pair of paired sample feature vectors, and add a similarity label for identifying whether it is a same-category feature to the first feature similarity.
[0082] For example, combined with Figure 2 , assuming that the strongly enhanced sample data is D s , and the weakly enhanced sample data is D w , input both D s and D w into two deep learning networks, the encoding network Q and the encoding network K, for feature extraction respectively. The first sub-sample feature vector Q s corresponding to D s and the second sub-sample feature vector Q w corresponding to D w can be obtained, and the third sub-sample feature vector K s corresponding to D s and the fourth sub-sample feature vector K w corresponding to D w output by the encoding network K can be obtained. Then Q sRespectively with K s and K w are paired, and Q w is respectively paired with K s and K w is paired, and the first feature similarity between the paired sample feature vectors is calculated, and the feature similarity matrices dist A and dist B can be obtained as follows:
[0083]
[0084]
[0085] Among them, represents the first feature similarity between the second sub-sample feature vector Q w and the fourth sub-sample feature vector K w , represents the first feature similarity between the second sub-sample feature vector Q w and the third sub-sample feature vector K s , represents the first feature similarity between the first sub-sample feature vector Q s and the fourth sub-sample feature vector K w , represents the first feature similarity between the first sub-sample feature vector Q s and the third sub-sample feature vector K s .
[0086] After obtaining the first feature similarity between the paired sample feature vectors, the feature similarity between the two sample features paired with the same sample category in the first feature similarity can be added with the first similarity label, and the feature similarity between the two sample features paired with different sample categories can be added with the second similarity label. Among them, the first similarity label is, for example, 1, and the second similarity label is, for example, 0. In this way, the first similarity label and the second similarity label can be used to distinguish whether the two paired features corresponding to the feature similarity are sample features of the same category.
[0087] Step 360: Determine the loss function value based on the first feature similarity and the similarity label, and adjust the model parameters of the initial feature extraction model and the momentum update model based on the loss function value until the training ends, and determine the trained initial feature extraction model as the feature extraction model.
[0088] Among them, the loss function can be selected as the binary cross-entropy loss function. After adding similarity labels to the first feature similarity, the loss function value of the binary cross-entropy loss function can be determined according to the first feature similarity and the corresponding similarity labels. Specifically, the loss function value loss can be determined using the following formula (3):
[0089]
[0090] where dist is the feature similarity matrix, including the above-mentioned feature similarity matrices dist A and dist B ; T is a variable that controls the curvature of the classification boundary and can be preset. For example, in the embodiments of the present invention, it can be set to T = 0.2; max is to find the maximum value; i is the index for summing all sample images in the current Batch, that is, the row number of the feature similarity matrix; j is the index for summing all paired sample categories in the current Batch, that is, the column number of the feature similarity matrix; y j is the similarity label of the feature similarity matrix, with 1 at the diagonal position and 0 at other positions.
[0091] Exemplarily, adjusting the model parameters of the initial feature extraction model and the momentum update model based on the loss function value may include: adjusting the model parameters of the initial feature extraction model through gradient backpropagation based on the loss function value; adjusting the model parameters of the momentum update model based on the preset momentum update parameter and the adjusted model parameters of the initial feature extraction model.
[0092] For example, combined with Figure 2 , after obtaining the loss function value, the model parameters of the encoding network Q can be updated through gradient backpropagation. The encoding network K is the momentum weighting of the encoding network Q, and the model parameters of the encoding network K can be updated iteratively through momentum weighting. Specifically, the model parameters of the encoding network K can be updated using the following formula (4):
[0093] Param k = Param k * f + Param q * (1 - f)(4)
[0094] where Param k is the model parameter of the encoding network K, Param q is the model parameter of the encoding network Q, and f is the preset momentum update parameter. In the embodiments of the present invention, f can be set to 0.999, for example.
[0095] Exemplarily, after the end of one epoch of training, the trained encoding network Q can be tested using the test data set. The test sample images are input into the trained encoding network Q, and the trained encoding network Q extracts features from the test sample images to obtain the feature vectors corresponding to each test sample image. Based on the feature vectors, evaluation metrics of the model are calculated. For example, the recall rate metric Recall@K of the model can be calculated, and the effect of the trained encoding network Q in this training is evaluated through the evaluation metrics of the model. After the test is completed, the trained encoding network Q with the Recall@K metric greater than the preset metric threshold can be used as the final feature extraction model.
[0096] Specifically, the trained encoding network Q can be used to extract the sample features of the test sample images in the entire test data set, calculate the similarity between the sample features of each test sample image and all the sample features in the entire test data set respectively, and sort the similarities from high to low. For each test sample image, the class labels corresponding to the top K similarities of the sample features are selected as the prediction labels, and the sample class label of the test sample image is used as the actual label. Then, the number of labels that are the same as the actual label in the prediction labels is determined, and the ratio of this number of labels to the total number of prediction labels is determined as the Recall@K metric. Among them, the sample classes of the test sample images in the test data set are completely different from the sample classes of the training sample images in the training data set.
[0097] For the training method of the feature extraction model provided in this exemplary embodiment, on the one hand, the training sample images of each training batch can be randomly obtained from the training data set, so that the sample classes of the training sample images of each training batch are different, and in the training sample images of each training batch, each sample class includes 2 sample images, one of which is subjected to strong augmentation processing and the other is subjected to weak augmentation processing. Then, the obtained strongly augmented sample data and weakly augmented sample data are used for contrast learning between models, which can enhance the contrast learning of the same-class sample images, enable the initial feature extraction model to learn more discriminative features, and improve the feature extraction ability of the trained feature extraction model. On the other hand, during the self-supervised contrast learning process, similarity labels are added to the feature similarities between the two sample feature vectors compared between models. Through the similarity labels, it can be distinguished whether the two paired features are the same sample features, avoiding the problem that data with the same class label is forcibly separated and data with the same example is used to calculate the feature similarity, and further improving the feature extraction ability of the trained feature extraction model.
[0098] Considering the situation where the number of categories is constantly increasing in image recognition tasks. For example, in vehicle model recognition tasks, new vehicle models are constantly added. In this case, if the traditional direct classification method is used in vehicle model recognition tasks, only the classification categories involved in the training of the image recognition model can be recognized. The classification categories are fixed and cannot recognize other categories outside the classification categories involved in the training. When new classification categories are added, the image recognition model needs to be retrained.
[0099] Based on this, in an exemplary embodiment of the present invention, a trained classification layer can be added after the feature extraction model. The classification layer is used to classify the target feature vector extracted by the feature extraction model to obtain the classification result of the target object in the image to be recognized. Specifically, recognizing the classification result of the target object in the image to be recognized based on the target feature vector may include: inputting the target feature vector into the classification layer to obtain the classification probability output by the classification layer; determining the first classification category corresponding to the maximum classification probability in the classification probability as the classification result of the target object in the image to be recognized; wherein, the classification layer is obtained by separately training the initial classification layer based on the feature extraction model, the second training sample image, and the category label corresponding to the second training sample image.
[0100] Specifically, an initial classification layer can be added after the feature extraction model. The feature extraction model is used to extract features from the second training sample image, and the extracted feature vector is input into the initial classification layer to obtain the predicted classification category output by the initial classification layer. Then, based on the predicted classification category and the category label, the loss value is determined, and based on this loss value, the classification layer parameters of the initial classification layer are updated using the backpropagation algorithm. This iterative training is continued until the initial classification layer converges and the training ends. The trained initial classification layer is determined as the classification layer. During this training process, only the classification layer parameters are updated, and the model parameters of the feature extraction model are not updated.
[0101] In this way, after the target feature vector of the image to be recognized is extracted by the feature extraction model provided in the embodiment of the present invention, the classification layer is further used to classify the target feature vector, which can improve the accuracy and precision of image recognition. When the number of image categories increases, only the classification layer needs to be retrained, and there is no need to retrain the feature extraction model, which greatly saves the training time.
[0102] In another exemplary embodiment of the present invention, a template feature vector with known classification categories can be constructed. After the target feature vector of the image to be recognized is extracted by the feature extraction model, the target feature vector can be compared with the template feature vector to match the classification category corresponding to the target feature vector to obtain the classification result of the target object in the image to be recognized.
[0103] Specifically, the classification result of the target object in the image to be recognized based on the target feature vector may include: determining the second feature similarity between the target feature vector and each feature vector in the obtained first template feature vector; sorting the second feature similarities in descending order, and obtaining the second classification categories corresponding to the first second preset number of feature vectors in the sorting result; and determining the classification category with the largest number in the second classification categories as the classification result of the target object in the image to be recognized.
[0104] Exemplarily, the first template feature vector can be determined by the following method: inputting the test sample image into the feature extraction model to obtain the first feature vector corresponding to each test classification category of the test sample image output by the feature extraction model; determining the third feature similarity between the feature vectors in the first feature vector, and determining the first third preset number of feature vectors corresponding in the first feature vector in ascending order of the third feature similarity to obtain the second template feature vector; for each template feature vector in the second template feature vector, determining the feature center vector based on the fourth feature similarity between the template feature vector and the second feature vector, and updating the template feature vector with the feature center vector to obtain the third template feature vector after updating the second template feature vector; and determining the third template feature vectors corresponding to each test classification category as the first template feature vector. Wherein, the second feature vector is the other feature vectors in the first feature vector except the second template feature vector; and the test sample image is a sample image with a known classification category.
[0105] Specifically, the test sample images of each test classification category in the test sample set can be input into the feature extraction model, and the feature extraction model can be used to extract features to obtain the first feature vectors of each test sample image corresponding to each test classification category. The first feature vector can distinguish the features of each test sample image with a high degree of discrimination. Then, for the first feature vectors corresponding to each test classification category, the feature similarity of each feature vector in the first feature vector is calculated to obtain the third feature similarity, and the third feature similarity is sorted in ascending order. Each third feature similarity corresponds to two feature vectors. Starting from the smallest third feature similarity, the first N feature vectors can be selected as the second template feature vectors corresponding to the test classification category, that is, the initial template feature vectors, according to the sorting of the third feature similarity. For each feature vector in the N feature vectors, such as the feature vector N1, the fourth feature similarity between the feature vector N1 and the other feature vectors in the first feature vector corresponding to the test classification category except N1 can be determined. For example, if the first feature vector includes Y feature vectors, the feature vector N1 can determine Y - 1 fourth feature similarities. Then, the Y - 1 fourth feature similarities can be sorted in descending order, and the feature vectors corresponding to the first M1 (such as 32) fourth feature similarities in the sorting result are taken together with the feature vector N1 to calculate the feature center vector, and the feature center vector is used to replace the feature vector N1, where M1 ≤ Y - 1. The same processing as that of the feature vector N1 is performed on each of the N feature vectors to complete the update of the initial template feature vectors of the test classification category, and the third template feature vectors corresponding to the test classification category are obtained. The same processing is performed on the first feature vectors corresponding to each test classification category, and the third template feature vectors corresponding to each test classification category can be obtained. These third template feature vectors can be determined as the first template feature vectors.
[0106] Exemplarily, the obtained first template feature vectors can be stored in a storage device, and the first template feature vectors can be obtained from the storage device during image recognition.
[0107] In the image recognition task, after the target feature vector of the image to be recognized is extracted by using the feature extraction model provided in the embodiment of the present invention, the second feature similarity between the target feature vector and each feature vector in the first template feature vector can be calculated, the second feature similarity is sorted in descending order, and the second classification categories corresponding to the first M2 (such as 8) feature vectors in the sorting result are obtained. Then, the second classification categories are counted, the number of each classification category in the second classification categories is obtained, and further, the classification category with the largest number is determined, and the classification category with the largest number is used as the classification result of the image to be recognized.
[0108] When a new classification category is added, the test sample image corresponding to this classification category can be obtained. By performing the same process of determining the template feature vector on this test sample image, the third template feature vector corresponding to the newly added classification category can be obtained, and this third template feature vector can be added to the first template feature vector to update the first template feature vector.
[0109] In this way, after the target feature vector of the image to be recognized is extracted by using the feature extraction model provided in the embodiment of the present invention, the target feature vector is further compared with the first template feature vector, and the most similar classification category is matched as the classification result of the image to be recognized, which can improve the accuracy and precision of image recognition. When the number of image categories increases, only the template feature vector of the new classification category needs to be added to the first template feature vector, and there is no need to retrain the feature extraction model, greatly saving the training time.
[0110] The embodiment of the present invention provides two optional image recognition methods, that is, the method of adding a classification layer after the feature extraction model for image recognition and the method of using the first template feature vector for image recognition. In practical applications, it can be flexibly selected according to the requirements of the usage scenario.
[0111] Next, the image recognition device provided by the present invention will be described. The image recognition device described below can be correspondingly referred to the image recognition method described above.
[0112] Figure 4 An exemplary structural schematic diagram of the image recognition device provided by the embodiment of the present invention is shown. Referring to Figure 4 as shown, the image recognition device may include: a feature extraction module 410, configured to obtain an image to be recognized and input the image to be recognized into the feature extraction model to obtain the target feature vector output by the feature extraction model; a recognition module 420, configured to recognize the classification result of the target object in the image to be recognized based on the target feature vector; wherein, the feature extraction model is obtained by respectively inputting the strongly augmented sample data and the weakly augmented sample data into the initial feature extraction model and the momentum update model corresponding to the initial feature extraction model, and performing self-supervised contrastive learning training on the initial feature extraction model and the momentum update model; the strongly augmented sample data is obtained by strongly augmenting the first sample image corresponding to each target sample category; the weakly augmented sample data is obtained by weakly augmenting the second sample image corresponding to each target sample category.
[0113] In an exemplary embodiment, the image recognition device may further include: a model training module for training a feature extraction model. Exemplarily, the model training module may include: a sample acquisition unit for randomly acquiring first training sample images of each training batch from a training dataset, where the first training sample images include first sample images and second sample images corresponding to each target sample category; a strong enhancement unit for performing strong enhancement processing on the first sample images to obtain strongly enhanced sample data; a weak enhancement unit for performing weak enhancement processing on the second sample images to obtain weakly enhanced sample data; a feature pairing unit for inputting both the strongly enhanced sample data and the weakly enhanced sample data into an initial feature extraction model and a momentum update model, and pairwise pairing the first sample feature vectors output by the initial feature extraction model and the second sample feature vectors output by the momentum update model; a similarity determination unit for determining a first feature similarity of each pair of paired sample feature vectors and adding a similarity label for identifying whether they are features of the same category to the first feature similarity; a training unit for determining a loss function value based on the first feature similarity and the similarity label, and adjusting the model parameters of the initial feature extraction model and the momentum update model based on the loss function value until the training is completed, and determining the trained initial feature extraction model as the feature extraction model.
[0114] In an exemplary embodiment, the sample acquisition unit is specifically configured to: for each training batch, randomly sort all sample categories of the training dataset, and determine the first preset number of sample categories after sorting as the target sample categories of the training batch; and acquire two sample images from the sample images of each target sample category to obtain the first sample image and the second sample image corresponding to each target sample category.
[0115] In an exemplary embodiment, the training unit is specifically configured to: adjust the model parameters of the initial feature extraction model by gradient backpropagation based on the loss function value; and adjust the model parameters of the momentum update model based on a preset momentum update parameter and the adjusted model parameters of the initial feature extraction model.
[0116] In an exemplary embodiment, the recognition module 420 is specifically configured to: input the target feature vector into a classification layer to obtain a classification probability output by the classification layer; and determine the first classification category corresponding to the maximum classification probability in the classification probabilities as the classification result of the target object in the image to be recognized; where the classification layer is obtained by separately training an initial classification layer based on the second training sample images, the category labels corresponding to the second training sample images, and the feature extraction model.
[0117] In an exemplary embodiment, the recognition module 420 is specifically configured to: determine the second feature similarity between the target feature vector and each feature vector in the obtained first template feature vector; sort the second feature similarities in descending order, and obtain the second classification categories corresponding to the first second preset number of feature vectors in the sorting result; determine the classification category with the largest number in the second classification categories as the classification result of the target object in the image to be recognized.
[0118] In an exemplary embodiment, the image recognition device may further include a template feature vector determination module. The template feature vector determination module may be configured to: input the test sample image into the feature extraction model to obtain the first feature vector corresponding to each test classification category output by the feature extraction model; determine the third feature similarity between the feature vectors in the first feature vector, and determine the first third preset number of feature vectors corresponding in the first feature vector in ascending order of the third feature similarity to obtain the second template feature vector; for each template feature vector in the second template feature vector, determine the feature center vector based on the fourth feature similarity between the template feature vector and the second feature vector, and update the template feature vector using the feature center vector to obtain the third template feature vector after updating the second template feature vector; the second feature vector is the other feature vectors in the first feature vector except the second template feature vector; determine the third template feature vectors corresponding to each test classification category as the first template feature vector.
[0119] Figure 5 An exemplary structural diagram of an electronic device is shown as Figure 5 shown, the electronic device may include: a processor 510, a communication interface 520, a memory 530, and a communication bus 540. Among them, the processor 510, the communication interface 520, and the memory 530 complete mutual communication through the communication bus 540. The processor 510 may call the logical instructions in the memory 530 to execute the image recognition method provided in any of the above method embodiments.
[0120] In addition, when the logical instructions in the above-mentioned memory 530 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.
[0121] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the image recognition method provided in any of the above method embodiments.
[0122] In yet another aspect, the present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is configured to execute the image recognition method provided in any of the above method embodiments.
[0123] Exemplarily, the computer-readable storage medium includes a non-transitory computer-readable storage medium.
[0124] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative efforts.
[0125] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0126] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. An image recognition method, characterized in that: include: Acquire an image to be identified, and input the image to be identified into a feature extraction model to obtain a target feature vector output by the feature extraction model; Identify a classification result of a target object in the image to be identified based on the target feature vector; Among them, the feature extraction model is obtained by respectively inputting the strongly enhanced sample data and the weakly enhanced sample data into the initial feature extraction model and the momentum update model corresponding to the initial feature extraction model, and performing self-supervised contrastive learning training on the initial feature extraction model and the momentum update model; the strongly enhanced sample data is obtained by performing strong enhancement processing on the first sample image corresponding to each target sample category; the weakly enhanced sample data is obtained by performing weak enhancement processing on the second sample image corresponding to each target sample category.
2. The image recognition method according to claim 1, characterized in that: The feature extraction model is trained through the following steps: Randomly acquiring a first training sample image of each training batch from a training data set, wherein the first training sample image includes the first sample image and the second sample image corresponding to each of the target sample categories; Performing strong enhancement processing on the first sample image to obtain the strongly enhanced sample data; Performing weak enhancement processing on the second sample image to obtain the weakly enhanced sample data; Inputting the strongly enhanced sample data and the weakly enhanced sample data into the initial feature extraction model and the momentum update model respectively, and pairing the first sample feature vector output by the initial feature extraction model with the second sample feature vector output by the momentum update model; Determine the first feature similarity of each pair of paired sample feature vectors, and add a similarity label for identifying whether the first feature similarity is a feature of the same category; A loss function value is determined based on the first feature similarity and the similarity label, and model parameters of the initial feature extraction model and the momentum update model are adjusted based on the loss function value until the training is completed, and the trained initial feature extraction model is determined as the feature extraction model.
3. The image recognition method according to claim 2, characterized in that: The step of randomly acquiring the first training sample image of each training batch from the training data set comprises: For each training batch, all sample categories of the training data set are randomly sorted, and a first preset number of sample categories of the sorted sample categories are determined as target sample categories of the training batch; Two sample images are obtained from the sample images of each target sample category to obtain the first sample image and the second sample image corresponding to each target sample category.
4. The image recognition method according to claim 2, characterized in that: The adjusting the model parameters of the initial feature extraction model and the momentum update model based on the loss function value includes: Based on the loss function value, adjusting the model parameters of the initial feature extraction model through gradient back propagation; Based on the preset momentum update parameters and the adjusted model parameters of the initial feature extraction model, the model parameters of the momentum update model are adjusted.
5. The image recognition method according to any one of claims 1 to 4, characterized in that: The classification result of identifying the target object in the image to be identified based on the target feature vector includes: Inputting the target feature vector into a classification layer to obtain a classification probability output by the classification layer; Determine a first classification category corresponding to a maximum classification probability among the classification probabilities as a classification result of the target object in the image to be identified; The classification layer is obtained by separately training the initial classification layer based on the second training sample image, the category label corresponding to the second training sample image and the feature extraction model.
6. The image recognition method according to any one of claims 1 to 4, characterized in that: The classification result of identifying the target object in the image to be identified based on the target feature vector includes: Determine the second feature similarity between the target feature vector and each feature vector in the acquired first template feature vector; Sorting the second feature similarities in descending order, and obtaining the second classification categories corresponding to the first second preset number of feature vectors in the sorting result; The classification category with the largest number in the second classification categories is determined as the classification result of the target object in the image to be identified.
7. The image recognition method according to claim 6, characterized in that: Also includes: Inputting the test sample image into the feature extraction model to obtain a first feature vector corresponding to each test classification category output by the feature extraction model; Determine a third feature similarity between each feature vector in the first feature vector, and determine a first third preset number of feature vectors corresponding to the first feature vector in order from low to high of the third feature similarity, to obtain a second template feature vector; For each template feature vector in the second template feature vector, a feature center vector is determined based on a fourth feature similarity between the template feature vector and the second feature vector, and the template feature vector is updated using the feature center vector to obtain a third template feature vector after the second template feature vector is updated; the second feature vector is other feature vectors in the first feature vector except the second template feature vector; The third template feature vectors corresponding to each test classification category are all determined as the first template feature vectors.
8. An image recognition device, characterized in that: include: A feature extraction module is used to obtain an image to be identified, and input the image to be identified into a feature extraction model to obtain a target feature vector output by the feature extraction model; A recognition module, which recognizes a classification result of a target object in the image to be recognized based on the target feature vector; Among them, the feature extraction model is obtained by respectively inputting the strongly enhanced sample data and the weakly enhanced sample data into the initial feature extraction model and the momentum update model corresponding to the initial feature extraction model, and performing self-supervised contrastive learning training on the initial feature extraction model and the momentum update model; the strongly enhanced sample data is obtained by performing strong enhancement processing on the first sample image corresponding to each target sample category; the weakly enhanced sample data is obtained by performing weak enhancement processing on the second sample image corresponding to each target sample category.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the image recognition method according to any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the image recognition method according to any one of claims 1 to 7 is implemented.