A Long-Tail Image Classification Method Based on Multi-Expert Dynamic Collaboration
By constructing a long-tail image classification model with dynamic collaboration with multiple experts, using dynamic adaptive learning and heterogeneous knowledge transfer learning, the problem of category imbalance in long-tail visual recognition is solved, and the model's recognition and generalization ability of all categories is improved.
Patent Information
- Application Number
- CN202411125310.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-16
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2044-08-16
AI Technical Summary
When dealing with category imbalance, existing long-tail visual recognition methods often ignore tail categories, resulting in the model overfitting common categories, unable to effectively identify uncommon categories, and unable to adapt to the dynamic changes in data distribution, limiting the model's ability to generalize new categories or rare categories.
Using a multi-expert dynamic collaboration method, a classification model including all-domain and domain experts is constructed. Through dynamic adaptive learning and heterogeneous knowledge transfer learning, the loss function and weight are dynamically adjusted, and collaboration among different experts is achieved to improve the recognition ability of all categories.
The model's ability to identify all categories is improved, the ability to discriminate against unusual categories is enhanced, the model's generalization ability and classification accuracy is improved, and prediction errors and uncertainties are reduced.
Smart Images

Figure CN119006917B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and particularly to a long-tail image classification method based on multi-expert dynamic collaboration. Background Art
[0002] Long-tail visual recognition is a challenging task in the fields of computer vision and image processing, which focuses on how to accurately recognize various visual objects with extremely unbalanced quantity distributions in the real world. In long-tail datasets, a small number of common categories occupy most of the samples, while a large number of uncommon categories have few samples. This imbalance poses a huge challenge to the image classification task because the model often overfits to the common categories and has insufficient recognition ability for the uncommon categories. However, these uncommon categories may be more important than the recognition of head categories, such as disease classification and dangerous driving recognition, etc. Therefore, solving the long-tail problem plays an important role in promoting the processing of visual tasks.
[0003] To address the imbalance problem in long-tail recognition, existing methods usually adopt rebalancing strategies, such as resampling and reweighting techniques, to alleviate the negative impact of data imbalance. These methods attempt to improve the recognition performance of the model for tail categories by adjusting the class weights or sample distributions during training. In addition, some methods also adopt a multi-expert collaboration framework to improve the overall recognition ability by constructing multiple expert models focusing on different data subsets.
[0004] However, existing methods still have significant limitations when dealing with the long-tail problem. Traditional methods often simply divide the dataset based on class frequencies without fully considering the complex relationships between classes and the intrinsic characteristics of samples. This results in some important tail categories being possibly ignored, and the overrepresentation of head categories may further exacerbate the imbalance problem of the model. In addition, simple division methods cannot adapt to the dynamic changes in data distribution, limiting the generalization ability of the model for new or rare categories. Summary of the Invention
[0005] The present invention proposes a long-tail image classification method based on multi-expert dynamic collaboration, which solves the problem of class imbalance in existing long-tail visual recognition and improves the recognition ability of the model for all classes.
[0006] To achieve the above object, the present invention provides a long-tail image classification method based on multi-expert dynamic collaboration, including the following steps:
[0007] Step S1: Construct a classification model including multiple expert sub-networks, perform frequency analysis on the image categories in the long-tail image training set, and divide the image categories into head, middle, and tail categories according to the frequency distribution of the image categories;
[0008] Step S2: Use the long-tail image training set as the input of the first expert sub-network, use the long-tail images corresponding to the middle and tail categories as the input of the second expert sub-network, and use the long-tail images corresponding to the tail category as the input of the third expert sub-network;
[0009] Step S3: Calculate the loss function of the classification model according to the outputs of each expert sub-network, and iteratively optimize the parameters of each expert sub-network according to the loss function until the set maximum number of iterations is reached;
[0010] Step S4: Input the long-tail image to be classified into the trained classification model to obtain the image classification result.
[0011] Preferably, the frequencies of the head, middle, and tail categories in Step S1 satisfy the following expression:
[0012] ;
[0013] In the formula, represents the frequency of the image category with the lowest frequency among the head, middle, and tail categories, where represent the head, middle, and tail categories respectively; represents the frequency of the image category with the highest frequency in the long-tail image training set.
[0014] Preferably, the loss function is constructed in the following steps in Step S3:
[0015] Step S31: Calculate the loss values of multiple expert sub-networks, and use dynamic adaptive weights to adjust each loss value to obtain the dynamic adaptive learning loss of the classification model;
[0016] Step S32: Perform weighted average on the outputs of all expert sub-networks to obtain a fusion vector, and calculate the KL divergence between the fusion vector and the predicted vectors of the outputs of each expert sub-network to obtain the KL divergence loss of the classification model;
[0017] Step S33: Combine the dynamic adaptive learning loss and the KL divergence loss to obtain the loss function of the classification model.
[0018] Preferably, the expression for calculating the loss values of multiple expert sub-networks in Step S31 is:
[0019] ;
[0020] ;
[0021] In the formula, is the loss value of the expert sub-network; is the i th image category; is the total number of image categories; is the i th probability value output by the expert sub-network; is the i th prediction vector output by the expert sub-network.
[0022] Preferably, the expression of the dynamic adaptive learning loss in step S31 is:
[0023] ;
[0024] In the formula, is the dynamic adaptive learning loss of the classification model; is the dynamic adaptive weight; , , are the loss values of the first, second, and third expert sub-networks respectively.
[0025] Preferably, the expression of the fusion vector in step S32 is:
[0026] ;
[0027] ;
[0028] ;
[0029] ;
[0030] In the above formula, is the fusion vector; , and are the terms representing the head, middle, and tail categories respectively; is the output vector of the first expert sub-network for the head category, where i represents the index number of the head category; is the weight vector of the first expert sub-network for the head category; represents the second norm; , are the output vectors of the first expert sub-network and the second expert sub-network for the middle category respectively, where j represents the index number of the middle category; , are the weight vectors of the first expert sub-network and the second expert sub-network for the middle category respectively; is the weight matrix of the middle category; , , They are the output vectors of the first expert sub-network, the second expert sub-network, and the third expert sub-network for the tail category, where k represents the index number of the tail category; , , are the weight vectors of the first expert sub-network, the second expert sub-network, and the third expert sub-network for the tail category respectively; is the weight matrix of the tail category.
[0031] Preferably, the expression of the KL divergence loss in step S32 is:
[0032] ;
[0033] ;
[0034] ;
[0035] ;
[0036] ;
[0037] In the formula, is the KL divergence loss of the classification function; is the KL divergence function; is the ideal probability distribution of the expert sub-network; is the actual probability distribution of the expert sub-network.
[0038] Preferably, the expression of the loss function of the classification model in step S33 is:
[0039] ;
[0040] In the formula, is the dynamic adaptive learning loss of the classification model; is the KL divergence loss of the classification function; is the weight of the KL divergence loss.
[0041] Preferably, when performing iterative optimization in step S3, the update method of the dynamic adaptive weight is:
[0042] ;
[0043] In the formula, is the current iteration number; is the set maximum iteration number.
[0044] Preferably, in step S1, the classification model fuses the prediction results of different expert sub-networks through a weighted average algorithm during both the training and inference phases.
[0045] The advantages of the present invention at least include:
[0046] 1. By enabling the first expert sub-network to simultaneously learn the head, middle, and tail categories, it ensures the comprehensive coverage of the model for all categories, thereby improving the recognition ability for common and uncommon categories; by enabling the second expert sub-network to simultaneously learn the middle and tail categories, the model can focus more on differentiating those categories that are similar in frequency but difficult to distinguish, enhancing the discriminative ability of the model for these categories; by enabling the third expert sub-network to focus on the tail categories, it helps to improve the recognition ability of the model for categories with a small sample size;
[0047] 2. By calculating the loss function for each expert sub-network and performing iterative optimization, it is possible to balance the impact of different categories on the model performance, improve the overall generalization ability of the model. By iteratively optimizing until reaching the maximum number of iterations, the model can continuously adjust and improve, gradually reducing the prediction error and improving the classification accuracy. Description of the Drawings
[0048] Figure 1 Schematic diagram for comparison of KL distances between backbone networks trained with cross-entropy and resampling;
[0049] Figure 2 Schematic diagram of the structure of an existing multi-expert model;
[0050] Figure 3 Schematic diagram of the structure of the classification model in the embodiment of the present invention;
[0051] Figure 4 Schematic diagram of the method flow in the embodiment of the present invention;
[0052] Figure 5 Comparison diagram of the existing dataset partitioning method and the dataset partitioning method in the embodiment of the present invention;
[0053] Figure 6 Schematic diagram of the trend of the dynamic adaptive weight changing with the increase of the number of iterations in the embodiment of the present invention;
[0054] Figure 7 Schematic diagram of fusing the outputs of different expert sub-networks in the embodiment of the present invention. Detailed Embodiments
[0055] Combined with the accompanying drawings in the embodiments of the present invention, the technical solutions in the embodiments of the present invention are clearly and completely described. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0056] Although adopting a multi-expert model is the optimal method to solve long-tail classification recognition at present, how to solve the problem of the decline in feature extraction ability of the multi-expert model in long-tail image classification recognition and at the same time reduce the uncertainty of integrated prediction is still an urgent problem to be solved at present:
[0057] First, there is uncertainty in the learning process of the backbone network of the multi-expert model. Even under the same network structure and training conditions, the prediction results of different networks for tail samples may vary significantly, which will affect the recognition accuracy of the multi-expert model for tail categories. To deeply understand this problem, experiments can be set up to explore the influence of feature representation and KL distance on the multi-expert model. As Figure 1 shown, two backbone networks with the same structure are trained using resampling (RS) and cross-entropy (CE) methods respectively, and the classifier parameters of the expert model are fixed. After training, the KL distance between the two experts is calculated based on the entire test set, and the average results of each category are statistically analyzed. The experimental results show that the shared backbone network can provide stronger feature representation ability, which helps to improve the consistency of predictions between different expert models, thereby reducing the uncertainty in the learning process. As can be seen from Figure 1 it, the backbone network trained by CE shows better feature representation ability, with a smaller KL distance and correspondingly reduced prediction uncertainty. Therefore, in the training stage of the multi-expert model, restoring and strengthening the feature representation ability of the backbone network is the key to alleviating prediction uncertainty.
[0058] In addition, the existing models increase the uncertainty and complexity of the model when implementing multi-expert prediction integration in the inference stage. The existing multi-expert models usually use different data to train each branch in the training stage, but in the inference stage, a fusion strategy needs to be designed to integrate the prediction results of different branches. As Figure 2 shown, different fusion methods, such as summation, concatenation or weighted average, may lead to different prediction results, and even some fusion methods may reduce the performance of the model and produce unsatisfactory predictions. To obtain more accurate prediction results, different fusion methods can be tried, but the selection of the best fusion strategy is still unknown. Therefore, how to design an effective fusion strategy to reduce the uncertainty of the multi-expert model in the inference stage is another important problem to be solved in the current research.
[0059] In view of the problems encountered by existing multi-expert models in long-tailed image recognition, the embodiments of the present invention design a multi-expert dynamic collaborative learning model (DCHKT) based on heterogeneous knowledge transfer, which can enhance the discriminative ability of each expert and the collaborative ability among multiple experts simultaneously. As Figure 3 shown, the DCHKT model includes 3 experts with different domain knowledge: 1 global expert and 2 domain experts. The model consists of two key components: dynamic adaptive learning and heterogeneous knowledge transfer learning. The dynamic adaptive learning mechanism can dynamically adjust the weights between feature learning and classifier learning, which can mitigate the possible damage to feature representation caused by the multi-expert model during the balanced training process. Heterogeneous knowledge transfer learning fuses the prediction results from experts in different domains through a weighted average algorithm, ensuring the consistency of model training and prediction, and realizing information transmission among multiple experts through heterogeneous knowledge transfer learning to enhance the collaboration between different experts.
[0060] Based on the constructed DCHKT model, the embodiments of the present invention provide a long-tailed image classification method based on multi-expert dynamic collaboration, as Figure 4 shown, including the following steps:
[0061] Step S1: Obtain a long-tailed image training set and partition the long-tailed image training set.
[0062] As Figure 5 shown, existing multi-expert models usually use a linear function as the partitioning criterion for the data set. However, due to the limitation of the number of training samples, the experts responsible for medium and low frequencies and the experts responsible for low-frequency categories may not be able to fully learn the decision boundaries of different categories. To solve this problem, the embodiments of the present invention use a square function as the partitioning basis for the data set. This partitioning method based on the square function can more accurately reflect the imbalance between categories, so that experts in each domain can obtain sufficient training samples to learn more accurate classifiers. This data set partitioning method based on the square function can improve the learning effect of experts on the decision boundaries of different categories.
[0063] Specifically, obtain the training set , where X represents an image and Y represents the category label corresponding to the image. Divide all the image categories C included in the training set into three subsets , where and , for the i th subset , the image category j will be assigned to it:
[0064] ;
[0065] In the formula, Represents the frequency of the image category with the highest frequency in the training set; Represents the j frequency of the th image category; K is the total number of subsets, and in the embodiments of the present invention
[0066] The three experts are represented as , that is, the first expert sub-network , the second expert sub-network and the third expert sub-network , where each expert can be assigned a subset . For the randomly sampled mini-batch training data , the expert can be trained on . According to the dataset division, the first expert sub-network is called the global expert, which can classify all categories, and the second and third expert sub-networks are called domain experts, which can identify some specific classes. These domain experts are usually dominated by intermediate or tail categories. For tail categories, multiple groups of experts will participate in the decision-making together to eliminate the bias of the global expert towards head categories. Through the division, the domain experts will output -dimensional prediction vectors, including n categories within its responsible range and one other class. Through data division, the intermediate and tail categories include sufficient training samples, and the corresponding experts can be fully trained.
[0067] Step S2: Use the divided long-tail image training set to train a preset classification model, that is, the DCHKT model, and obtain the output results of each domain expert.
[0068] In the training of the multi-expert model, it is crucial to achieve balance between different branches. However, traditional models often simply assign the same weights to each branch or use a residual mechanism for fusion, which ignores the differences in the contributions of different experts to the model. To solve this problem, enhancing the feature representation ability of the shared backbone network is the key, which can not only enhance the consistency of predictions among experts but also reduce the uncertainty in the model learning process. In the embodiments of the present invention, by designing a dynamic adaptive learning mechanism, instead of adjusting features, the loss values output by each expert are dynamically adjusted to train the backbone network and classifier of the model, balance the feature learning and classifier training from the three branches, and improve the overall performance of the model when dealing with long-tail distribution data.
[0069] In the dynamic adaptive learning mechanism, the global expert is responsible for classifying all categories, with a particular focus on the head categories. Its core advantage lies in the learning of feature representations, which can capture a wide range of feature information and provide a solid foundation for the model. The domain expert focuses on specific groups of categories, learning the mid-low frequency categories and the tail categories separately. By focusing on these categories, the domain expert can learn more deeply how to distinguish rare categories, thereby improving the model's recognition ability for these categories.
[0070] After the training set is input into the classification model, the global expert The output probability And the loss value Can be expressed as:
[0071] ;
[0072] ;
[0073] In the above formula, Is the Th i Prediction vector output by the global expert; , is the Th i Image category in; Is the Total number of all image categories in.
[0074] Similarly, the probability Output by the domain expert And the loss value Can be expressed as:
[0075] ;
[0076] ;
[0077] In the above formula, Is the Th i Prediction vector output by the domain expert; , is the Th i Image category in; Is the Total number of all image categories in.
[0078] The probability Output by the domain expert And the loss value Can be expressed as:
[0079] ;
[0080] ;
[0081] In the above formula, is the domain expert The i th predicted vector output; , is The i th image category in; is The number of all image categories in.
[0082] At this stage, the dynamic adaptive learning loss of the classification model can be expressed as:
[0083] ;
[0084] where is the dynamic adaptive weight, which is automatically adjusted during training.
[0085] Set the maximum number of iterations for model training to , for the T-th iteration, The update method of is:
[0086] .
[0087] As Figure 6 shown is The trend of changing with the increase of the number of training times. It can be seen from the figure that if , then , and the loss of the global expert in training dominates the total loss of the model. In this case, the model will focus on general classifier training and feature representation learning in the initial stage.
[0088] If or , then , and the losses generated by the two domain experts in training dominate the total loss. At this stage, the learning focus of the model gradually shifts from the global expert to the domain experts, and at the same time, it will also promote the domain experts' classifiers to recognize intermediate samples and tail samples.
[0089] If , then , and shows a decreasing trend. At this stage, the loss of the global expert again dominates the total loss of the model. At the same time, the model gradually shifts its attention from the domain experts to the global expert, thereby restoring the model's feature representation ability and balancing the learning of features and classifiers.
[0090] Step S3: Fuse the prediction results of experts in different knowledge domains through a weighted average algorithm based on weight norm to obtain a fusion vector.
[0091] Specifically, fusing the outputs of different branches during the inference phase is crucial for a multi-expert model because the model needs to be able to handle images of unknown classes, whether they are head, middle, or tail classes. However, many existing multi-expert models decouple the fusion of predictions from their training process, which may lead to some unexpected results. For example, simply averaging the outputs of different branches may amplify the bias towards tail classes.
[0092] To improve the cooperation of experts in different domains in terms of consistent predictions, the embodiments of the present invention propose a heterogeneous knowledge transfer learning method to achieve effective message passing between experts in different domains. First, use a weighted average algorithm based on weight norm to fuse the prediction results of multiple experts to form a fusion vector. Then, normalize the prediction vectors of experts in different domains and calculate the KL divergence between the fusion vector and each expert's normalized vector to perform message passing. This method not only improves the efficiency of the training phase, but also during the inference phase, the model will also use the same fusion method to generate a fusion vector for classification and recognition, maintaining the consistency between the training and inference processes.
[0093] As Figure 7 shown, the weighted average algorithm based on weight norm normalizes and evaluates the prediction results of different experts through the classifier weights. Each expert will only participate in the final decision within their own capabilities, thus eliminating the model's bias towards tail-class samples through group collaborative decision-making. Among them, the head classes are mainly determined by the global experts decide, the middle classes are determined by the global experts and the domain experts decide, and the tail classes will be determined by all experts , and jointly decide.
[0094] After training, the outputs of the classifiers of each expert can be expressed as:
[0095]
[0096] Among them, , and represent the input features of the classifier, the weight vector of the classifier, and the bias respectively.
[0097] The embodiments of the present invention perform norm weighting on the weights of each classifier, and the weighted term representing the head class Can be expressed as:
[0098] ;
[0099] In the formula, is the output vector of the global expert for the head category, where i represents the index number of the head category; is the weight vector of the global expert for the head category; ||·|| represents the second norm.
[0100] The weighted term representing the middle category Can be expressed as:
[0101] ;
[0102] In the formula, , are respectively the output vectors of the global expert and the domain expert for the middle category, where j represents the index number of the middle category; , are respectively the output vectors of the global expert and the domain expert for the middle category; is the weight matrix of the middle category, and its size is .
[0103] The weighted term representing the tail category Can be expressed as:
[0104] ;
[0105] In the formula, , , are respectively the output vectors of the global expert , the domain expert and for the tail category, where k represents the index number of the tail category; , , are respectively the output vectors of the global expert , the domain expert and for the tail category; is the weight matrix of the tail category, and its size is .
[0106] Finally, by concatenating the prediction items representing the head, middle, and tail parts mentioned above, the final fusion vector can be obtained. , the model will calculate the final probability vector based on the fusion vector for class prediction.
[0107] By analyzing the prediction vectors of different branches, it can be found that the items representing the middle category in the prediction vector of the global expert are generally smaller than the corresponding items in the prediction vector of the domain expert because the global expert needs to identify the three categories of head, middle, and tail, while the domain expert mainly identifies the middle and tail categories. Therefore, weights , are set in the embodiments of the present invention.
[0108] In the heterogeneous knowledge transfer learning stage, the KL divergence loss of the classification model can be expressed as:
[0109] ;
[0110] ;
[0111]
[0112] ;
[0113] ;
[0114] ;
[0115] In the above formula, is the KL divergence function; is the ideal probability distribution of expert , ; is the actual probability distribution of expert .
[0116] Step S4: Calculate the final classification and recognition result of the image according to the fusion vector, and at the same time construct the loss function of the classification model, and iteratively optimize the parameters of the classification model according to the loss function until the set maximum number of iterations is reached.
[0117] Specifically, the total loss function of the classification model is constructed as:
[0118] ;
[0119] Among them, is to balance The weight of the contribution, in the embodiments of the present invention .
[0120] Step S5: Input the long-tail image to be classified into the trained classification model to obtain the image classification result.
[0121] A long-tail image classification method based on multi-expert dynamic collaboration proposed by the present invention realizes the collaborative learning of the head, middle, and tail sample classifiers by constructing a multi-expert collaboration model including a global expert and domain experts. The global expert is responsible for learning the features of all categories, while the domain experts focus on specific category groups, such as medium-frequency and low-frequency categories. This design aims to improve the classification ability of various category samples in the long-tail distribution through the collaboration of different experts.
[0122] Aiming at the problem of possible decline in feature expression ability when the existing multi-expert model is trained on an imbalanced dataset, the present invention designs a dynamic adaptive learning model training method. This method can simultaneously meet the requirements of the model in extracting image feature expression ability and training multi-expert classifiers, ensuring that the model can effectively learn and express image features.
[0123] At the same time, in order to solve the problem of prediction uncertainty that may be caused by the additional design of different branch fusion methods in the prediction and inference stage of the existing multi-expert model, as well as the problem of lack of collaborative cooperation between different experts, the present invention proposes a heterogeneous knowledge transfer learning method. By using the weighted average algorithm to fuse the prediction results from different domain experts, this method not only maintains the consistency of model training and prediction, but also realizes the information transfer between multiple experts through heterogeneous knowledge transfer learning, enhancing the collaboration between different experts.
[0124] By adopting the same fusion method in the prediction stage as in the training stage, the need for additional design of fusion strategies is avoided, thereby reducing the uncertainty in the prediction process. At the same time, through heterogeneous knowledge transfer learning, the model can transfer information between different experts, improving the collaboration efficiency and classification performance of the entire system.
[0125] In summary, the multi-expert dynamic collaboration long-tail image classification method of the present invention effectively solves the imbalance problem in long-tail image classification and improves the classification accuracy and robustness of the model through innovative dynamic adaptive learning and heterogeneous knowledge transfer learning.
[0126] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. Only the preferred embodiments of the present invention are expressed. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the patent of the present invention. As long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.
[0127] It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the patent of the present invention should be subject to the appended claims.
Claims
1. A long-tail image classification method based on multi-expert dynamic collaboration, characterized in that It includes the following steps: Step S1: Construct a classification model including multiple expert sub-networks, perform frequency analysis on the image categories in the long-tailed image training set, and divide the image categories into head, middle, and tail categories according to the frequency distribution of the image categories; Step S2: Use the long-tailed image training set as the input of the first expert sub-network, use the long-tailed images corresponding to the middle and tail categories as the input of the second expert sub-network, and use the long-tailed images corresponding to the tail category as the input of the third expert sub-network; Step S3: Calculate the loss function of the classification model according to the outputs of each expert sub-network, and iteratively optimize the parameters of each expert sub-network according to the loss function until the set maximum number of iterations is reached; Construct the loss function through the following steps: Step S31: Calculate the loss values of multiple expert sub-networks, adjust each loss value using dynamic adaptive weights, and obtain the dynamic adaptive learning loss of the classification model; Step S32: Perform weighted average on the outputs of all expert sub-networks to obtain a fusion vector, and calculate the KL divergence loss of the classification model according to the fusion vector; The expression of the fusion vector is: In the above formula, is the fusion vector; and are terms representing the head, middle, and tail categories respectively; is the output vector of the first expert sub-network for the head category, where a represents the index number of the head category; is the weight vector of the first expert sub-network for the head category; ||·|| represents the second norm; are the output vectors of the first expert sub-network and the second expert sub-network for the middle category respectively, where b represents the index number of the middle category; are the weight vectors of the first expert sub-network and the second expert sub-network for the middle category respectively; W M is the weight matrix of the middle category; are the output vectors of the first expert sub-network, the second expert sub-network, and the third expert sub-network for the tail category respectively, where c represents the index number of the tail category; are the weight vectors of the first expert sub-network, the second expert sub-network, and the third expert sub-network for the tail category respectively; W T is the weight matrix of the tail category; The expression of the KL divergence loss is: Where, L KL is the KL divergence loss of the classification function; KL() is the KL divergence function; is the ideal probability distribution of the expert sub-network e k ; is the actual probability distribution of the expert sub-network e k ; Step S33: Combine the dynamic adaptive learning loss and the KL divergence loss to obtain the loss function of the classification model; Step S4: Input the long-tailed image to be classified into the trained classification model to obtain the image classification result.
2. The long-tail image classification method based on multi-expert dynamic collaboration according to claim 1, wherein: The frequencies of the head, middle, and tail categories in Step S1 satisfy the following expression: Where N i represents the frequency of the image category with the lowest frequency among the head, middle, and tail categories, where i = 1, 2, and 3 represent the head, middle, and tail categories respectively; N max represents the frequency of the image category with the highest frequency in the long-tailed image training set.
3. A long-tail image classification method based on multi-expert dynamic collaboration according to claim 1, characterized in that: The expression for calculating the loss values of multiple expert sub-networks in Step S31 is: where L is the loss value of the expert sub-network; y i is the i-th image category; n is the total number of image categories; P i is the i-th probability value output by the expert sub-network; z i is the i-th prediction vector output by the expert sub-network.
4. The long-tail image classification method based on multi-expert dynamic collaboration according to claim 1, characterized in that: The expression of the dynamic adaptive learning loss in Step S31 is: L DAL = α(L2 + L3) + (1 - α)L1; Where L DAL is the dynamic adaptive learning loss of the classification model; α is the dynamic adaptive weight; L1, L2, and L3 are the loss values of the first, second, and third expert sub-networks, respectively.
5. A long-tail image classification method based on multi-expert dynamic collaboration according to claim 1, characterized in that: The expression of the loss function of the classification model in Step S33 is: L all = L DAL + λL KL ; where, L DAL is the dynamic adaptive learning loss of the classification model; L KL is the KL divergence loss of the classification function; λ is the weight of the KL divergence loss.
6. The long-tail image classification method based on multi-expert dynamic collaboration according to claim 1, characterized in that: When performing iterative optimization in Step S3, the update method of the dynamic adaptive weight α is: where T is the current iteration number; T max is the set maximum number of iterations.
7. A long-tail image classification method based on multi-expert dynamic collaboration according to claim 1, characterized in that: In Step S1, the classification model fuses the prediction results of different expert sub-networks through a weighted average algorithm during both the training and inference stages.
Citation Information
Patent Citations
Egg quality measurement method, system and equipment based on machine vision and medium
CN117745661A