Rotation network-based classification model training and classification method and related device
By generating the difference calculation of pseudo-label and rotating network, the model robustness problem of target domain samples is solved, and efficient image classification of passive domain data is achieved.
Patent Information
- Application Number
- CN202510332814.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-20
- Publication Date
- 2025-07-04
AI Technical Summary
The prior art cannot fully utilize the potential information of the target domain when the target domain samples are not marked, resulting in poor model training robustness.
By acquiring the target domain image set, enhancing image set and rotating image set, a classification network and clustering algorithm are used to generate pseudo-labels, combining the rotating classification network and the mapping network, calculating the degree of difference and updating the model parameters, forming a classification model based on the rotating network.
The robustness and adaptability of target domain image classification are improved, the generalization performance of the image classification model for target domain samples is improved, and the image classification model training of passive domain data is realized.
Smart Images

Figure CN120259751A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image recognition, and particularly to a training and classification method for a classification model based on a rotation network and related devices. Background Art
[0002] As an important basic task in computer vision, image recognition aims to classify an input image into a predefined class label. Currently, deep learning-based image classification models have been widely studied. Around the classification accuracy and computational efficiency of the models, researchers have proposed various architectures, such as AlexNet, VGG, GoogLeNet, and ResNet, etc. Among them, typically, ResNet proposed by Kaiming He et al. solves the problem of difficult training of deep networks by introducing residual learning, enabling the model to effectively perform deeper feature extraction and significantly improving the classification performance. However, in actual applications, when a traditional image recognition task is directly applied to the target domain after being trained on the source domain, the performance often drops significantly due to the difference in data distribution.
[0003] To improve the above problems, the domain adaptation image recognition method has emerged. It draws on the concept of transfer learning and learns the common domain-invariant features between the source domain and the target domain by using the labeled data in the source domain and a small amount of labeled or a large amount of unlabeled data in the target domain. However, in the case where the target domain samples are unlabeled, this method cannot fully utilize the potential information of the target domain, resulting in poor robustness of the finally trained model. Therefore, it is necessary to provide a training and classification method for a classification model based on a rotation network and related devices. Summary of the Invention
[0004] In view of the above-mentioned disadvantages of the prior art, the purpose of the present invention is to provide a training and classification method for a classification model based on a rotation network and related devices, which improves the problem that in the case where the target domain samples are unlabeled, the potential information of the target domain cannot be fully utilized, resulting in poor robustness of model training.
[0005] To achieve the above and other related objectives, the present invention provides a training method for a classification model based on a rotation network, including: obtaining a target domain image set, an augmented image set, a rotated image set, and a true label set; wherein, the augmented images are obtained by augmenting the target domain image data, the rotated images are generated by rotating the augmented images, and the true labels are the rotation angles at which the augmented images are rotated to the corresponding rotated images; inputting the augmented image set, the rotated image set, and the target domain image set into the feature extraction network of the classification model to extract corresponding augmented features, rotated features, and target features; inputting the target features into the classification network of the classification model to generate a prediction result, and calculating a classification difference degree based on the prediction result and the corresponding pseudo-label; wherein, the pseudo-label is obtained by clustering all the target features input into the classification network; inputting the augmented features and the rotated features into the rotation classification network of the classification model, and calculating a rotation prediction difference degree based on the predicted rotation angle generated by the rotation classification network and the corresponding true label; inputting the augmented features and the rotated features into the mapping network of the classification model, and calculating a contrastive learning difference degree based on the projection vectors of the augmented features and the rotated features obtained by the mapping network; weighting the rotation prediction difference degree, the classification difference degree, and the contrastive learning difference degree to obtain a total difference degree, and updating the parameters of the rotation network and the feature extraction network based on the total difference degree to obtain a trained classification model, and extracting the feature extraction network and the classification network therefrom to form an image classification model.
[0006] In an embodiment of the present invention, the generation process of the pseudo-label includes: clustering all the target features according to the corresponding prediction results to obtain a plurality of category clusters; for each target feature input into the classification network: calculating the distances between the target feature and each category cluster; selecting the category cluster with the closest distance to the target feature, and taking the corresponding category as the pseudo-label of the target feature.
[0007] In an embodiment of the present invention, the augmented images include image augmentation in a first manner and image augmentation not in the first manner for the target domain images. For each target domain image, the process of generating augmented features, rotated features, and target features through the augmented images and the rotated images includes: forming an augmented sample pair by combining the augmented image obtained by image augmentation in the first manner and the rotated image corresponding to the image augmentation not in the first manner; inputting the augmented sample pair and the corresponding target domain image into the feature extraction network to extract the augmented features, selected features, and target features correspondingly.
[0008] In an embodiment of the present invention, inputting the target feature into the classification network of the classification model to generate a prediction result, and calculating the classification difference degree based on the prediction result and the corresponding pseudo-label includes: inputting the target feature into the classification network to generate a prediction result; calculating the cross-entropy loss between the prediction result and the corresponding pseudo-label; obtaining a classification weight according to the prediction result and the total number of categories predicted by the classification network; and performing weighted processing on the cross-entropy loss based on the classification weight to obtain the classification difference degree.
[0009] In an embodiment of the present invention, the classification weight is where K is the total number of categories, ε is a preset stable coefficient constant, and w i is the classification weight corresponding to the target domain image, and is the prediction result corresponding to the target domain image.
[0010] In an embodiment of the present invention, the classification model is ResNet-101.
[0011] In an embodiment of the present invention, a classification method is further provided, including: obtaining a target image to be classified; inputting the target image into the feature extraction network of a trained image classification model to extract the target feature of the target image; and inputting the target feature into the classification network of the image classification model to obtain the classification result of the target image.
[0012] In an embodiment of the present invention, a training device for a classification model based on a rotation network is further provided. The system includes: a data acquisition module, configured to acquire a target domain image set, an enhanced image set, a rotated image set, and a true label set; wherein, the enhanced images are obtained by augmenting the target domain image data, the rotated images are generated by rotating the enhanced images, and the true label is the rotation angle at which the enhanced image is rotated to the corresponding rotated image; a feature extraction module, configured to input the enhanced image set, the rotated image set, and the target domain image set into the feature extraction network of the classification model to extract corresponding enhanced features, rotated features, and target features; a classification module, configured to input the target features into the classification network of the classification model to generate a prediction result, and calculate a classification difference degree based on the prediction result and the corresponding pseudo label; wherein, the pseudo label is obtained by clustering all the target features input into the classification network; a rotation angle prediction module, configured to input the enhanced features and the rotated features into the rotation classification network of the classification model, and calculate a rotation prediction difference degree based on the predicted rotation angle generated by the rotation classification network and the corresponding true label; a contrastive learning module, configured to input the enhanced features and the rotated features into the mapping network of the classification model, and calculate a contrastive learning difference degree based on the projection vector of the enhanced features and the projection vector of the rotated features obtained by the mapping network; a parameter update module, configured to perform weighted processing on the rotation prediction difference degree, the classification difference degree, and the contrastive learning difference degree to obtain a total difference degree, and update the parameters of the rotation network and the feature extraction network based on the total difference degree to obtain a trained classification model, and extract the feature extraction network and the classification network therefrom to form an image classification model.
[0013] In an embodiment of the present invention, an electronic device is further provided, including: one or more processors; a storage device, configured to store one or more programs, which, when executed by the one or more processors, cause the electronic device to implement the training method for the classification model based on the rotation network according to any one of the above.
[0014] In an embodiment of the present invention, a computer-readable storage medium is further provided, on which a computer program is stored, which, when executed by a processor of a computer, causes the computer to execute the training method for the classification model based on the rotation network according to any one of the above.
[0015] As described above, a training and classification method for a classification model based on a rotation network and related devices of the present invention have the following beneficial effects: performing data augmentation and rotation on target domain images to correspondingly obtain augmented images and rotated images, and generating pseudo-labels of target features through a classification network and a clustering algorithm. The learning ability of the model for geometric transformations is improved through a rotation classification network and a mapping network, and the parameters of the feature extraction network and the rotation classification network are updated by calculating the final total loss. The finally obtained image classification model consists of a feature extraction network and a classification network. In this way, the robustness of target domain image classification is improved, and the data features are fully mined through multi-task collaborative optimization, enhancing the adaptability and generalization performance of the image classification model to target domain samples, and enabling the training of an image classification model for source-free domain data. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 FIG. is a schematic flowchart of a method for training an image classification model provided by an embodiment of the present invention;
[0017] Figure 2 FIG. is a schematic diagram for obtaining enhanced sample pairs of the present invention;
[0018] Figure 3 FIG. is an overall schematic diagram of the training process of the image classification model of the present invention;
[0019] Figure 4 FIG. shows a structural block diagram of a training system for an image classification model provided by an embodiment of the present invention;
[0020] Figure 5 FIG. shows a schematic structural diagram of an electronic device according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0021] The following specifically illustrates the embodiments of the present invention through specific examples. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that, without conflict, the following embodiments and the features in the embodiments can be combined with each other.
[0022] It should be noted that the diagrams provided in the following embodiments only illustrate the basic concept of the present invention in a schematic manner. Therefore, only the components related to the present invention are shown in the diagrams, rather than being drawn according to the number, shape, and size of the components in actual implementation. The types, quantities, and ratios of the components in actual implementation can be arbitrarily changed, and the component layout type may also be more complex.
[0023] In the following description, numerous details are explored to provide a more thorough explanation of embodiments of the present invention. However, it is obvious to those skilled in the art that the embodiments of the present invention can be implemented without these specific details. In other embodiments, well-known structures and devices are shown in the form of block diagrams rather than in detail to avoid making the embodiments of the present invention difficult to understand.
[0024] The inventors found that the effectiveness of current image recognition deep networks is attributed to two basic conditions: one is that the training samples and test samples come from the same data set or different data sets with similar distributions; the other is that a large number of labeled samples are required in the training stage. Generally, the data set collected from the actual scenario of the task is called the target domain (unlabeled), while the publicly available data set (labeled) of the relevant scenario is called the source domain. First, due to differences in factors such as background, lighting conditions, and object appearance, there are significant differences in the styles of target domain and source domain images, resulting in large differences in data distributions. Second, the supervised samples for image recognition tasks require the annotation of classification labels for each image. Especially when the number of categories is large, the cost of manual annotation is very high. Therefore, the accuracy and robustness of cross-domain image classification models face huge challenges.
[0025] Domain Adaptive Image Classification was first officially proposed at the CVPR (IEEE Conference on Computer Vision and Pattern Recognition) conference in 2018. Drawing on the concept of transfer learning, it learns the common domain-invariant features between the source domain and the target domain by using the labeled data in the source domain and a small amount of labeled or a large amount of unlabeled data in the target domain. Researchers have proposed various solutions from multiple perspectives, such as based on distribution alignment, parallel training, data generation and reconstruction, hybrid optimization strategies, etc., effectively balancing the contradiction between feature generality and classification task specificity.
[0026] However, source domain data does not always meet the basic assumptions of traditional UDA due to reasons such as privacy policies and memory limitations of small devices. For example, the European Union has promulgated the General Data Protection Regulation (GDPR) to restrict the use and transmission of personal data. Many industries such as military and hospitals will refuse to share data information with the outside world due to privacy reasons. On the other hand, source domain data sets are often huge. The huge ImageNet data set contains 14 million images and occupies nearly 170GB of memory, which is unbearable for small storage devices. If only the pre-trained model on the source domain is transmitted to replace the source domain data, the above problems can be solved, that is, the novel method of source-free domain adaptation is adopted.
[0027] In view of the above problems, the present invention provides a training method for a classification model based on a rotation network. The target domain images are subjected to data augmentation and rotation to obtain corresponding augmented images and rotated images, and pseudo-labels of target features are generated through a classification network and a clustering algorithm. The learning ability of the model for geometric transformation is improved through a rotation classification network and a mapping network. The parameters of the feature extraction network and the rotation classification network are updated by calculating the final total loss. The finally obtained image classification model consists of a feature extraction network and a classification network. In this way, the robustness of target domain image classification is improved, and the data features are fully mined through multi-task collaborative optimization, and the adaptability and generalization performance of the image classification model to target domain samples are improved, and the training of the image classification model for source-free domain data can be realized.
[0028] Please refer to Figure 1 , the training method of the classification model based on the rotation network includes the following steps:
[0029] S1. Obtain a target domain image set, an augmented image set, a rotated image set, and a true label set; wherein, the augmented images are obtained by data augmentation of the target domain images, the rotated images are generated by rotating the augmented images, and the true labels are the rotation angles for rotating the augmented images to the corresponding rotated images.
[0030] As Figure 2 shown, the target domain images refer to unlabeled images for feature extraction and pseudo-label generation. Data augmentation is performed on each target domain image x to obtain a corresponding augmented image x i , where the data augmentation methods include but are not limited to random cropping, Gaussian blur, image color jitter, etc. And for all target domain images, different data augmentation methods can be used, and at least two different data augmentations are performed on the same target domain image (such as x i and x j , i and j are different data augmentation methods). The image after at least one data augmentation (such as x j ) is rotated at a preset angle to obtain a rotated image Rotate x i and to form an augmented sample pair. In one embodiment, the rotation angles of the images are respectively θ ∈ {0°, 90°, 180°, 270°}, and the corresponding true labels are y r ∈ {0, 1, 2, 3}, and each true label corresponds to a rotation angle.
[0031] S2. Input the augmented image set, the rotated image set, and the target domain image set into the feature extraction network of the classification model to extract corresponding augmented features, rotated features, and target features.
[0032] Specifically, in an embodiment of the present invention, the enhanced image includes image enhancement in a first manner and image enhancement in a non-first manner for the target domain image. For each target domain image, the process of generating enhanced features, rotation features, and target features through the enhanced image and the rotated image includes:
[0033] Combining the enhanced image obtained by image enhancement in the first manner and the rotated image corresponding to the non-first manner of image enhancement to form an enhanced sample pair;
[0034] Inputting the enhanced sample pair and the corresponding target domain image into the feature extraction network to extract enhanced features, selected features, and target features correspondingly.
[0035] For each target domain image, the following processing is performed: image enhancement in at least two different manners is performed respectively to obtain corresponding enhanced images. The enhanced images obtained by the two different enhancement manners can be respectively denoted as x i and x j . Exemplarily, the enhanced image obtained by performing Gaussian blur (such as the first manner) on the target domain image is x i , and the enhanced image obtained by performing color jitter (such as the non-first manner) on the same target domain image is x j . For the enhanced image x j obtained by the non-first manner, it is rotated by a preset angle to generate a rotated image It can be understood that the enhanced image obtained by the non-first manner can be rotated by only one preset angle to obtain a rotated image. In order to increase the amount of data, the enhanced image x j obtained by the non-first manner can also be rotated by all preset angles (such as rotated by 0°, 90°, 180°, 270° respectively to obtain corresponding rotated images). The specific manner is not limited herein. For each target domain image: combining the enhanced image x i obtained by the first manner and the rotated image of the enhanced image x j obtained by the non-first manner to form an enhanced sample pair, denoted as Inputting the target domain image and the corresponding enhanced sample pair into the feature extraction network of the classification model to extract the target feature of the target domain image x, the enhanced feature of the enhanced image x i obtained by the first manner, and the rotation feature of the rotated image .
[0036] It can be understood that the classification model is pre-trained through source domain images. The pre-trained model includes a feature extraction network and a classification network. The present invention can directly obtain the classification model pre-trained on the source domain without the need to obtain the source domain images for self-training, and only transmit the model pre-trained on the source domain, thereby replacing the direct transmission of the source domain images and effectively avoiding the privacy leakage problem that may be caused by using the source domain images. In addition, due to the large volume of the source domain images, it greatly occupies the memory space, which is unbearable for small storage devices. By directly obtaining the pre-trained classification model, the present invention not only solves the storage limitation problem, but also significantly reduces the training time and resource consumption, and improves the efficiency of model deployment. Therefore, when the capacity of the storage device is limited, the classification model pre-trained on the source domain can be directly used without obtaining the source domain data for local training, thus taking into account both privacy protection and efficiency.
[0037] It can be further understood that this solution can also train a classification model including a feature generation network and a classification network by using the source domain images. The pre-training process should be repeated until the model converges and reaches the preset recognition accuracy on the source domain. For the labeled images on the source domain, using the conventional cross-entropy function as the loss function can achieve the training effect.
[0038] The classification model can be any model that recognizes image features and classifies them, including but not limited to convolutional neural networks, residual networks, etc. In order to extract the complex features of the images and facilitate the accurate recognition of the image types, the classification model in the present invention is ResNet-101, and its network consists of an initial convolutional layer (conv1), 3 conv2_x modules, 4 conv3_x modules, 23 conv4_x modules, and 3 conv5_x modules. Each module contains 3 convolutional layers, and the number of convolutional kernels is 64, 128, 256, and 512 in sequence, and a skip layer is introduced through a residual connection. In addition, the network contains an average pooling layer and a fully connected layer. It can be understood that after obtaining the classification model pre-trained on the source domain or training the classification model with the source domain data, the source domain data is no longer accessed thereafter, and only the unlabeled data in the target domain is used for subsequent model parameter updates.
[0039] S3. Input the target features into the classification network of the classification model to generate a prediction result, and calculate the classification difference degree based on the prediction result and the corresponding pseudo-label; wherein, the pseudo-label is obtained by clustering all the target features input into the classification network.
[0040] Input the target features obtained from the target domain image through the feature extraction network into the classification network to generate corresponding prediction results, and cluster based on all the target features, and assign pseudo-labels according to the similarity between the target features and the cluster centers. Compare the prediction results of the classification network with the pseudo-labels, calculate the cross-entropy loss, and weight the cross-entropy loss according to the confidence weights of each target domain image to obtain the classification difference degree.
[0041] In an embodiment of the present invention, the process of generating pseudo-labels includes:
[0042] Cluster all the target features according to the corresponding prediction results to obtain multiple category clusters;
[0043] For each target feature input to the classification network:
[0044] Calculate the distances between the target feature and each category cluster;
[0045] Select the category cluster with the closest distance to the target feature, and use the corresponding category as the pseudo-label of the target feature.
[0046] For the current training batch: First, pass the target domain images of the current batch through the feature extraction network G to obtain the target features of each target domain image, and input each target feature into the classification network C to obtain the corresponding prediction results (t is any target domain image in the current batch), and use the prediction result of each target feature as the initial pseudo-label of the corresponding target feature. Cluster all the target features according to the predicted prediction results to generate multiple category clusters. During the clustering process, the initial cluster center is calculated as shown in formula (1):
[0047]
[0048] where represents the initial cluster center of the k-th category cluster, N t represents the number of target domain samples in the current training batch, represents the prediction result generated by the classification network, K represents the K-th category, and 1(*) is an indicator function, which takes the value of 1 when the prediction result of sample i belongs to category K, otherwise 0. Through the above process, the initial cluster center of each category cluster can be calculated. Calculate the distances between each target feature and the initial cluster centers of each category cluster, and by calculating the distances between the target feature and the cluster centers of each category cluster, select the category with the closest distance to the target feature as the first pseudo-label corresponding to the target domain image, as shown in formula (2):
[0049]
[0050] Among them, is the first pseudo-label generated after the first clustering iteration of the target feature. D(a, b) represents a distance measurement function. Exemplarily, it can be a cosine similarity function. After obtaining the first pseudo-label, clustering processing is performed according to all the first pseudo-labels in the current batch, and the cluster centers of each category cluster are recalculated according to formula (3):
[0051]
[0052] Among them, is the cluster center of the updated category cluster. By updating the center of each category cluster in the above manner, a new round of distance calculation is performed on the updated cluster center and the target feature according to formula (2), and the category of the category cluster closest to the target feature is used as the final pseudo-label of the corresponding target domain image. In this way, pseudo-labels with less noise can be obtained
[0053] In an embodiment of the present invention, inputting the target feature into the classification network of the classification model to generate a prediction result, and calculating a classification difference degree based on the prediction result and the corresponding pseudo-label, including:
[0054] Inputting the target feature into the classification network to generate a prediction result;
[0055] Calculating the cross-entropy loss between the prediction result and the corresponding pseudo-label;
[0056] Obtaining a classification weight according to the prediction result and the total number of categories predicted by the classification network;
[0057] Weighting the cross-entropy loss based on the classification weight to obtain a classification difference degree.
[0058] Input the target feature into the classification network C of the classification model, and the classification network generates a prediction result based on the target feature And normalize the prediction result through the Softmax function to obtain the final prediction result Denoted as Among them, G(x) is the target feature of the target domain image x, G is the feature extraction network, C(G(x)) is the output of the classification network C, and δ(*) is the Softmax function. In order to further enhance the optimization ability of the model for high-confidence samples, the present invention also calculates the corresponding classification weight according to the prediction result, and the classification weight is calculated based on the entropy value of the prediction result, as shown in formula (4):
[0059]
[0060] Among them, K represents the total number of categories, ε represents a stable coefficient constant, w i is the classification weight corresponding to the target domain image, is the final prediction result corresponding to the target domain image. Through the classification weights, the prediction confidence of the classification network for each sample can be reflected. The higher the weight, the more credible the network's prediction for the sample. After calculating the cross-entropy loss between the prediction result and the corresponding pseudo-label, the classification weights and the corresponding cross-entropy loss are weighted, as shown in formula (5), to obtain the classification loss L cls :
[0061]
[0062] The present invention adjusts the weight of each sample by measuring the uncertainty of the prediction distribution. The more certain the prediction (the lower the entropy value), the greater the weight, so that the finally obtained model has higher accuracy.
[0063] S4. Input the enhanced feature and the rotated feature into the rotation classification network of the classification model, and calculate the rotation prediction difference degree based on the predicted rotation angle generated by the rotation classification network and the corresponding true label.
[0064] The enhanced image x obtained by performing data augmentation on the target domain image x in the first manner i , is input into the feature extraction network to obtain the corresponding enhanced feature, denoted as f i . The enhanced image x obtained by performing data augmentation on the same target domain image x in a non-first manner j , is rotated to obtain a rotated image The rotated image is input into the feature extraction network to obtain the corresponding rotated feature, denoted as f j . Since the enhanced image only uses the data augmentation method and does not involve image rotation, the enhanced feature f i and the rotated feature f j are input into the rotation classification network, and the predicted rotation angle from the enhanced image x i rotated to the rotated image x j can be obtained. Calculate the rotation prediction difference degree L rot between the predicted rotation angle and the corresponding true label, as shown in formula (6):
[0065]
[0066] where is the true rotation label, G represents the feature generation network, C represents the classification network, is the combination of the enhanced feature of the target domain image x and the rotated feature of the rotated image , N is the number of enhanced sample pairs in the current batch, K is the total number of categories in the rotation classification task. Exemplarily, K is 4, corresponding to rotation angles of 0°, 90°, 180°, and 270°.
[0067] S5. Input the enhanced feature and the rotation feature into the mapping network of the classification model, and calculate the contrast learning difference degree based on the projection vector of the enhanced feature and the projection vector of the rotation feature obtained from the mapping network.
[0068] The present invention trains a model in combination with a contrast learning framework. In order to further improve the effect of contrast learning, the present invention introduces specific projection processing in the feature space, making the data more discriminative during the contrast learning process. Specifically, the enhanced feature and the rotation feature are input into the mapping network, and the mapping network can be an MLP, a fully connected layer, etc. Preferably, for better contrast learning, the mapping network in this embodiment is an MLP. Through the mapping network, the corresponding enhanced projection vector and rotation projection vector are obtained, and the contrast learning difference degree L between the two projection vectors is calculated. con as shown in formula (7):
[0069]
[0070] where z i is the projection vector of the enhanced feature, is the projection vector of the rotation feature, sim(*) is a similarity calculation function, and exemplarily, it can be a cosine similarity. t is a preset temperature parameter (such as 0.5). Contrast learning constructs positive sample pairs of enhanced features and rotation features, narrows the distance between different enhancements of the same view, and at the same time maximizes the discrimination between positive samples and other negative samples, improving the representation ability of the model in the target domain. In addition, during the contrast learning process, the performance of rotation prediction is directly related to the learning of enhanced view features. Therefore, on the basis of contrast adaptation, the present invention introduces the sensitivity modeling of the model to the rotation view, that is, by learning the rotation view of specific data augmentation, further improving the ability of contrast learning to represent the features of target domain samples. All data augmentations including the rotation angle are included in the positive sample pairs, and these sample pairs are input into the classical contrast learning formula for optimization training. Through the contrast learning in each training epoch, the model gradually narrows the distance between positive sample pairs and improves the accuracy of rotation training; on the other hand, the positive sample pairs after data augmentation further improve the robustness of rotation training, thus significantly improving the adaptability of the model to target domain data. It should be noted that the positive sample pairs in the present invention are composed of enhanced features and rotation features generated by the same target domain image through different enhancement methods. Since these features come from the same image, they can be regarded as having the same semantic information. The negative sample pairs are composed of enhanced features and rotation features of different target domain images. Since these features come from different images, they can be regarded as having different semantic information.
[0071] S6. Weight the rotation prediction difference degree, classification difference degree, and contrast learning difference degree to obtain the total difference degree, and update the parameters of the rotation classification network and the feature extraction network based on the total difference degree to obtain a trained classification model, and extract the image classification model composed of the feature extraction network and the classification network therefrom.
[0072] Weight the rotation prediction difference degree, classification difference degree, and contrast learning difference degree by the preset weights and sum them up to obtain the total difference degree L target , as shown in formula (8):
[0073] L target =αL cls +βL rot +λL con (8)
[0074] where α, β, λ are preset weight balance parameters. By continuously updating the parameters of the rotation classification network and the feature extraction network and keeping the parameters of the fixed mapping network and the classification network unchanged, when the total difference degree is less than the preset threshold, the entire classification model is considered trained. At this time, the feature extraction network and the classification network are used to form an image classification model.
[0075] In another embodiment of the present invention, a classification method is further provided. The recognition method includes:
[0076] Obtain the target image to be recognized;
[0077] Input the target image into the feature extraction network of the trained image classification model to extract the target features of the target image;
[0078] Input the target features into the classification network of the image classification model to obtain the classification result of the target image.
[0079] For the new target domain image, input it into the trained image classification model. After being processed by the feature extraction network and the classification network, the model can accurately predict the category to which the target domain image belongs according to the feature representation of the image. Through the learning and adaptation ability of the model on the target domain data, the output prediction result is the category label of the target domain image, realizing the automatic classification and efficient recognition of the target domain image.
[0080] As Figure 3 shown, input the target domain image and the augmented sample pair into the feature extraction network of the classification model to obtain the target features, augmented features, and rotation features. Input the target features into the classification network C of the classification model to obtain the classification loss according to the corresponding pseudo-label; input the augmented features and rotation features into the mapping network MLP to calculate the contrast learning loss between them; input the augmented features and rotation features into the rotation classification network C r, the rotation angle is obtained, and the rotation prediction loss is obtained based on the rotation angle and the true label. These three losses are weighted to obtain the total loss, and the parameters of the feature extraction network and the rotation classification network are updated until a trained classification model is obtained.
[0081] Please refer to Figure 4 , the training device 100 of the classification model based on the rotation network includes: a data acquisition module 110, a feature extraction module 120, a classification module 130, a rotation angle prediction module 140, a contrast learning module 150, and a parameter update module 160. The above-mentioned data acquisition module 110 is used to acquire a target domain image set, an augmented image set, a rotated image set, and a true label set; wherein, the augmented image is obtained by augmenting the target domain image data, the rotated image is generated by rotating the augmented image, and the true label is the rotation angle of the augmented image rotated to the corresponding rotated image. The feature extraction module 120 is used to input the augmented image set, the rotated image set, and the target domain image set into the feature extraction network of the classification model to extract the corresponding augmented features, rotated features, and target features. The classification module 130 is used to input the target features into the classification network of the classification model to generate a prediction result, and calculate the classification difference degree based on the prediction result and the corresponding pseudo-label; wherein, the pseudo-label is obtained by clustering all the target features input into the classification network. The rotation angle prediction module 140 is used to input the augmented features and the rotated features into the rotation classification network of the classification model, and calculate the rotation prediction difference degree based on the predicted rotation angle generated by the rotation classification network and the corresponding true label. The contrast learning module 150 is used to input the augmented features and the rotated features into the mapping network of the classification model, and calculate the contrast learning difference degree based on the projection vector of the augmented features and the projection vector of the rotated features obtained by the mapping network. The parameter update module 160 is used to perform weighted processing on the rotation prediction difference degree, the classification difference degree, and the contrast learning difference degree to obtain the total difference degree, and update the parameters of the rotation network and the feature extraction network based on the total difference degree to obtain a trained classification model, and extract the feature extraction network and the classification network from it to form an image classification model.
[0082] For the specific limitations of the training device of the classification model based on the rotation network, reference can be made to the limitations of the training method of the image classification model in the above text, which will not be elaborated here. Each module in the above-mentioned training device of the classification model based on the rotation network can be implemented in whole or in part by software, hardware, and their combination. The above-mentioned modules can be embedded in or independent of the processor in the computer device in a hardware format, or stored in the memory of the computer device in a software format, so that the processor can call the corresponding operations of the above-mentioned modules.
[0083] It should be noted that, in order to highlight the innovative part of the present invention, modules that are not closely related to solving the technical problems proposed by the present invention are not introduced in this embodiment. However, this does not mean that there are no other modules in this embodiment.
[0084] Please refer to Figure 5 , the electronic device 1 may include a memory 12, a processor 13, and a bus, and may also include a computer program stored in the memory 12 and executable on the processor 13, such as a training program for a classification model based on a rotation network.
[0085] Among them, the memory 12 includes at least one type of readable storage medium, and the readable storage medium includes flash memory, mobile hard disk, multimedia card, card-type memory (such as: SD or DX memory, etc.), magnetic memory, magnetic disk, optical disk, etc. The memory 12 may be an internal storage unit of the electronic device 1 in some embodiments, such as the mobile hard disk of the electronic device 1. The memory 12 may also be an external storage device of the electronic device 1 in other embodiments, such as a plug-in mobile hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. equipped on the electronic device 1. Further, the memory 12 may include both an internal storage unit and an external storage device of the electronic device 1. The memory 12 can be used not only to store application software installed in the electronic device 1 and various types of data, such as the code for training a classification model based on a rotation network, etc., but also to temporarily store data that has been output or will be output.
[0086] The processor 13 may be composed of integrated circuits in some embodiments. For example, it may be composed of a single packaged integrated circuit, or may be composed of multiple integrated circuits with the same or different functions, including a combination of one or more Central Processing Units (CPUs), microprocessors, digital processing chips, graphics processors, and various control chips. The processor 13 is the control core (Control Unit) of the electronic device 1, connecting various components of the entire electronic device 1 through various interfaces and lines, and by running or executing programs or modules stored in the memory 12 (such as a training program for a classification model based on a rotation network, etc.), and calling data stored in the memory 12, to execute various functions of the electronic device 1 and process data.
[0087] The processor 13 executes the operating system of the electronic device 1 and various installed application programs. The processor 13 executes the application program to implement the steps in the above-mentioned training method for a classification model based on a rotation network.
[0088] Exemplarily, the computer program may be divided into one or more modules, which are stored in the memory 12 and executed by the processor 13 to complete the present application. The one or more modules may be a series of computer program instruction segments capable of performing specific functions, and these instruction segments are used to describe the execution process of the computer program in the electronic device 1. For example, the computer program may be divided into a data acquisition module 110, a feature extraction module 120, a classification module 130, a rotation angle prediction module 140, a contrast learning module 150, and a parameter update module 160.
[0089] The integrated units implemented in the form of software function modules as described above may be stored in a computer-readable storage medium, which may be non-volatile or volatile. The above software function modules are stored in a storage medium and include several instructions for causing a computer device (which may be a personal computer, a computer device, or a network device, etc.) or a processor to execute part of the functions of the training method of the image classification model according to various embodiments of the present application.
[0090] In summary, for a training and classification method and related device of a classification model based on a rotation network disclosed in the present invention, the target domain images are subjected to data augmentation and rotation, and the augmented images and rotated images are correspondingly obtained. The pseudo-labels of the target features are generated through a classification network and a clustering algorithm. The learning ability of the model for geometric transformations is improved through a rotation classification network and a mapping network. The parameters of the feature extraction network and the rotation classification network are updated by calculating the final total loss. The finally obtained image classification model is composed of a feature extraction network and a classification network. In this way, the robustness of target domain image classification is improved, and the data features are fully mined through multi-task collaborative optimization, enhancing the adaptability and generalization performance of the image classification model to target domain samples, and enabling the training of an image classification model for source-free domain data. The present invention addresses the challenges faced by the current technology, such as the inability to obtain a large number of source domain labeled samples in real time and the inability to access source domain data during the training process, which leads to problems such as ineffective training of detectors. It solves the problem of difficult acquisition of source domain data due to factors such as privacy policies and data transmission restrictions, and completes the domain adaptation image classification technology for source-free domains. Therefore, the present invention effectively overcomes various drawbacks in the prior art and has high industrial utilization value.
[0091] The above embodiments are only illustrative of the principles and effects of the present invention and are not intended to limit the present invention. Any person familiar with this technology can modify or change the above embodiments without departing from the spirit and scope of the present invention. Therefore, all equivalent modifications or changes made by those with ordinary knowledge in the technical field without departing from the spirit and technical idea disclosed by the present invention should still be covered by the claims of the present invention.
Claims
1. A training method for a classification model based on a rotation network, characterized in that, The training method includes: Obtain a target domain image set, an enhanced image set, a rotated image set, and a true label set; wherein, the enhanced images are obtained by data augmentation of the target domain images, the rotated images are generated by rotating the enhanced images, and the true labels are the rotation angles at which the enhanced images are rotated to the corresponding rotated images; Input the enhanced image set, the rotated image set, and the target domain image set into the feature extraction network of the classification model to extract corresponding enhanced features, rotated features, and target features; Input the target features into the classification network of the classification model to generate a prediction result, and calculate the classification difference degree based on the prediction result and the corresponding pseudo label; wherein, the pseudo label is obtained by clustering all the target features input into the classification network; Input the enhanced features and the rotated features into the rotation classification network of the classification model, and calculate the rotation prediction difference degree based on the predicted rotation angle generated by the rotation classification network and the corresponding true label; Input the enhanced features and the rotated features into the mapping network of the classification model, and calculate the contrast learning difference degree based on the projection vectors of the enhanced features and the projection vectors of the rotated features obtained by the mapping network; Perform weighted processing on the rotation prediction difference degree, the classification difference degree, and the contrast learning difference degree to obtain the total difference degree, and update the parameters of the rotation network and the feature extraction network based on the total difference degree to obtain a trained classification model, and extract the feature extraction network and the classification network from it to form an image classification model.
2. The training method of the classification model based on the rotation network according to claim 1, wherein The generation process of the pseudo label includes: Cluster all the target features according to the corresponding prediction results to obtain multiple category clusters; For each target feature input into the classification network: Calculate the distances between the target feature and each category cluster; Select the category cluster with the closest distance to the target feature, and use the corresponding category as the pseudo label of the target feature.
3. The training method of the classification model based on the rotation network according to claim 1, wherein The enhanced images include image enhancement in the first way and image enhancement not in the first way for the target domain images. For each target domain image, the process of generating enhanced features, rotated features, and target features through the enhanced images and the rotated images includes: Form enhanced sample pairs by combining the enhanced images obtained by image enhancement in the first way and the rotated images corresponding to the image enhancement not in the first way; Input the enhanced sample pairs and the corresponding target domain images into the feature extraction network to extract enhanced features, selected features, and target features correspondingly.
4. The training method of the classification model based on the rotation network according to claim 1, wherein The step of inputting the target features into the classification network of the classification model to generate a prediction result and calculating the classification difference degree based on the prediction result and the corresponding pseudo label includes: Input the target features into the classification network to generate a prediction result; Calculate the cross-entropy loss between the prediction result and the corresponding pseudo label; Obtain the classification weights based on the prediction result and the total number of categories predicted by the classification network; Perform weighted processing on the cross-entropy loss based on the classification weights to obtain the classification difference degree.
5. The training method of the classification model based on the rotation network according to claim 4, characterized in that, The classification weight is where K is the total number of categories, ε is a preset stability coefficient constant, and w i is the classification weight corresponding to the target domain image, and is the prediction result corresponding to the target domain image.
6. The training method of the classification model based on the rotation network according to claim 1, characterized in that The classification model is ResNet-101.
7. A classification method, characterized in that, The classification method includes: Obtain a target image to be classified; Input the target image into the feature extraction network of the trained image classification model to extract the target features of the target image; Input the target feature into the classification network of the image classification model to obtain the classification result of the target image.
8. A training device for a classification model based on a rotation network, characterized in that, The device includes: A data acquisition module, configured to acquire a target domain image set, an enhanced image set, a rotated image set, and a true label set; wherein, the enhanced image is obtained by data augmentation of the target domain image, the rotated image is generated by rotating the enhanced image, and the true label is the rotation angle of the enhanced image rotated to the corresponding rotated image. A feature extraction module, configured to input the enhanced image set, the rotated image set, and the target domain image set into the feature extraction network of the classification model to extract corresponding enhanced features, rotated features, and target features. A classification module, configured to input the target feature into the classification network of the classification model to generate a prediction result, and calculate a classification difference degree based on the prediction result and the corresponding pseudo label; wherein, the pseudo label is obtained by clustering all the target features input into the classification network. A rotation angle prediction module, configured to input the enhanced feature and the rotated feature into the rotation classification network of the classification model, and calculate a rotation prediction difference degree based on the predicted rotation angle generated by the rotation classification network and the corresponding true label. A contrastive learning module, configured to input the enhanced feature and the rotated feature into the mapping network of the classification model, and calculate a contrastive learning difference degree based on the projection vector of the enhanced feature and the projection vector of the rotated feature obtained by the mapping network. A parameter update module, configured to perform weighted processing on the rotation prediction difference degree, the classification difference degree, and the contrastive learning difference degree to obtain a total difference degree, and update the parameters of the rotation network and the feature extraction network based on the total difference degree to obtain a trained classification model, and extract the feature extraction network and the classification network from it to form an image classification model.
9. An electronic device, characterized in that, The electronic device includes: One or more processors; A storage device, configured to store one or more programs, which, when executed by the one or more processors, cause the electronic device to implement the training method of the classification model based on the rotation network according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, A computer program is stored thereon, which, when executed by the processor of the computer, causes the computer to execute the training method of the classification model based on the rotation network according to any one of claims 1 to 7.