A small-sample object detection method and device based on mixture of experts
Through the hybrid expert network combined with repeated upsampling and data enhancement, the problems of backbone network freezing and category imbalance in small sample object detection are solved, achieving better classification effect and robustness.
Patent Information
- Application Number
- CN202510020152.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-07
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-01-07
AI Technical Summary
In the existing small sample object detection method, the backbone network freeze results in the failure to further optimize the features, the flexibly utilize different depth features, and the problem of category imbalance has not been effectively solved, resulting in the model's detection performance in small sample categories.
A hybrid expert network is adopted, combining repeated upsampling and data enhancement, and through the shared shallow feature network, unique deep feature network, feature fusion network and object detection network, the knowledge distillation loss function is used for parameter optimization to achieve fusion and learning of different deep features.
It improves the classification effect and robustness of the model in small sample object detection, alleviates the overfitting problem, and enhances the model's adaptability and detection performance to complex environments.
Smart Images

Figure CN119418041B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of object detection, and particularly relates to a few-shot object detection method and device based on mixture of experts. Background Art
[0002] Object detection is a fundamental task in the field of computer vision, aiming to identify various things in an image and classify them. Traditional object detection methods often rely on a large amount of labeled data to train models. However, the cost of collecting a large amount of high-quality labeled data is high. Therefore, few-shot object detection (FSOD) has emerged. Few-shot object detection aims to solve the problem of the performance degradation of traditional object detection models in the case of scarce labeled data, narrowing the gap between machine vision models and the human visual system, so that the model can accurately detect objects of new categories even with only a very small number of labeled samples.
[0003] Current few-shot object detection methods mainly improve the detection performance of the model from three aspects. First, improving the forgetfulness of the model: The model needs to rely on the basic category knowledge learned during the pre-training period to assist in the learning of few-shot categories. Therefore, it is extremely important not to forget the knowledge learned previously during the learning of few-shot categories. Second, improving the confusion of the model: The model does not want to be confused with the basic category knowledge learned during the pre-training period when learning few-shot categories, resulting in the inability to distinguish between basic categories and few-shot categories. Third, handling class imbalance: Since the number of basic category images during the pre-training period of the model is huge, it is particularly important to balance the huge difference between the number of few-shot category images and basic category images.
[0004] Existing few-shot object detection methods are usually divided into two stages. In stage 1, pre-training is performed on basic category images with a large amount of data annotation. Then, in stage 2, the weights of the backbone network (ResNet-101 is used in this paper) are frozen and various schemes are used for fine-tuning on few-shot categories and basic categories. The disadvantage of such a method is that the features of the backbone network have been frozen during fine-tuning and cannot be further optimized. At the same time, only the deepest features are used for learning, and the learned features sometimes cannot well distinguish between basic categories and new categories. Although there are already methods such as Feature Pyramid Network (FPN) to deal with the problem of only using the deepest network, FPN is not flexible enough in processing features of different depths. Therefore, the problem of combining features of different depths is still worthy of exploration. In addition, when dealing with class imbalance in few-shot object detection tasks, simple upsampling or downsampling methods are often used, thus ignoring more traditional data augmentation methods. Summary of the Invention
[0005] In view of the above, the object of the present invention is to provide a few-shot object detection method and device based on mixture of experts. By introducing multiple mixture of experts, not only can their respective deep features be obtained, but also shallow features with different depths can be obtained separately. By combining deep features and shallow features and using the method of knowledge distillation, each mixture of experts can learn from each other, so that the object detection model can capture the key points of classification from different depth features and achieve better classification effects.
[0006] To achieve the above object of the invention, an embodiment provides a few-shot object method based on mixture of experts, including the following steps:
[0007] Use the base-class images to preliminarily train the object detection model;
[0008] Use repeated upsampling and data augmentation methods to augment the new-class small-sample images to obtain new-class augmented images;
[0009] Construct a few-shot object detection model based on mixture of experts, which includes a shared shallow feature network, a unique deep feature network for each mixture of experts, a feature fusion network, and an object detection network corresponding to each mixture of experts. The shared shallow feature network is used to extract shallow features with different depths from the input image, the unique deep feature network is used to extract their respective deep features based on the shallow features, the feature fusion network is used to fuse the shallow features with different depths and their respective deep features obtained by each mixture of experts to obtain their respective fused features, and the object detection network is used to predict the object detection results corresponding to each mixture of experts based on the fused features; wherein, the backbone network of the preliminarily trained object detection model is used to initialize the shared shallow feature network and the unique deep feature network for each mixture of experts;
[0010] Use the base-class images and the new-class augmented images to optimize the parameters of the above few-shot object detection model based on mixture of experts. The loss function used during parameter optimization includes the object detection losses corresponding to each mixture of experts and the distillation losses between each mixture of experts;
[0011] Use the few-shot object detection model with optimized parameters to perform few-shot object detection, and select the final object detection result from the object detection results of multiple mixture of experts.
[0012] Preferably, the object detection model uses the Faster RCNN network. When it is preliminarily trained, the cross-entropy loss is used for the classification results included in the object detection results, and the bounding box regression loss is used for the bounding box regression results included in the object detection results.
[0013] Preferably, the new category small sample images are enhanced by adopting repeated upsampling and data augmentation methods to obtain new category enhanced images, including:
[0014] Firstly, repeated upsampling is performed on the new category small sample images to increase the number of images, and then data augmentation is performed on each set of small sample images with increased quantity. The data augmentation methods include contrast adjustment, saturation adjustment, hue adjustment, addition of Gaussian noise and salt-and-pepper noise.
[0015] Preferably, the feature fusion network fuses the shallow features with different depths and their respective deep features obtained by each mixture of experts to obtain their respective fused features, including:
[0016] Firstly, the shallow features with different depths obtained by each mixture of experts are downsampled to adjust to the same size as the deep features, and then the downsampled shallow features and the deep features are fused to obtain the fused features. , where represents the shallow features obtained by the i -th mixture of experts, represents the deep features corresponding to the i -th mixture of experts, The symbol represents the Hadamard product, represents the downsampling process, represents the calculation of taking the logarithm.
[0017] Preferably, the object detection network corresponding to each mixture of experts includes a region proposal network and an ROI Head module that provides class prediction and bounding box regression prediction for each candidate region. The classification result and the bounding box regression result output by the ROI Head module form the object detection result corresponding to each mixture of experts.
[0018] Preferably, the loss function adopted during parameter optimization includes the object detection loss corresponding to each mixture of experts and the distillation loss between each mixture of experts
[0019] ;
[0020] ;
[0021] ;
[0022] where represents the sum of the classification losses corresponding to all mixtures of experts, represents the sum of the logistic regression losses corresponding to all mixtures of experts, and denotes hyperparameters, and denotes the index of the mixture of experts, m denotes the number of the mixture of experts, and respectively denote the i th mixture of experts and the j th classification result of the mixture of experts, denotes and the KL divergence loss between them.
[0023] Preferably, screening the final object detection result from the object detection results of multiple mixtures of experts includes:
[0024] Screening the object detection result corresponding to the mixture of experts with the highest class score from the object detection results of multiple mixtures of experts as the final object detection result.
[0025] To achieve the above object of the invention, an embodiment of the present invention further provides a small sample object detection device based on a mixture of experts, including:
[0026] A preliminary training module, which is used to preliminarily train the object detection model using basic class images;
[0027] A data augmentation module, which is used to augment the small sample images of new classes by adopting repeated upsampling and data augmentation methods to obtain new class augmented images;
[0028] A model construction module, which is used to construct a small sample object detection model based on a mixture of experts, which includes a shared shallow feature network, a unique deep feature network for each mixture of experts, a feature fusion network, and an object detection network corresponding to each mixture of experts, wherein the shared shallow feature network is used to extract shallow features of different depths from the input image, the unique deep feature network is used to extract their respective deep features based on the shallow features, the feature fusion network is used to fuse the shallow features of different depths and their respective deep features obtained by each mixture of experts to obtain their respective fused features, and the object detection network is used to predict the object detection results corresponding to each mixture of experts based on the fused features; wherein, the backbone network of the preliminarily trained object detection model is used to initialize the shared shallow feature network and the unique deep feature network for each mixture of experts;
[0029] A model training module, which is used to optimize the parameters of the above small sample object detection model based on a mixture of experts using basic class images and new class augmented images, and the loss function used in parameter optimization includes the object detection loss corresponding to each mixture of experts and the distillation loss between each mixture of experts;
[0030] A target detection module, which is used to perform few-shot object detection using a parameter-optimized few-shot object detection model and screen the final object detection result from the object detection results of multiple mixture-of-experts.
[0031] To achieve the above-mentioned invention purpose, the embodiment also provides a computing device, including a memory and one or more processors. An executable code is stored in the memory. When the one or more processors execute the executable code, it is used to implement the above-mentioned few-shot object detection method based on mixture-of-experts.
[0032] To achieve the above-mentioned invention purpose, the embodiment also provides a computer-readable storage medium, on which a program is stored. When the program is executed by a processor, the above-mentioned few-shot object detection method based on mixture-of-experts is implemented.
[0033] Compared with the prior art, the beneficial effects of the present invention at least include:
[0034] The method of the present invention is the first method to apply mixture-of-experts in the field of few-shot object detection. It not only avoids the negative impact brought by freezing the parameters of the backbone network, but also can more reasonably combine features of different depths to more comprehensively analyze images and optimize the image classification process;
[0035] The method of the present invention is the first method to combine repeated upsampling and data augmentation in the field of few-shot object detection. It not only alleviates the overfitting problem that may be brought by the simple upsampling process, but also increases the data volume of few-shot categories, enabling the model to handle various image detection problems and improving the generalization of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for description in the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0037] Figure 1 is a flowchart of the few-shot object detection method based on mixture-of-experts provided by the embodiment;
[0038] Figure 2 is a schematic structural diagram of the preliminary training of the few-shot object detection model provided by the embodiment;
[0039] Figure 3 is a schematic diagram of data augmentation provided by the embodiment;
[0040] Figure 4 is a schematic structural diagram of the few-shot object detection model based on mixture-of-experts provided by the embodiment;
[0041] Figure 5 It is a schematic structural diagram of the few-shot object detection device based on mixture of experts provided by the embodiment. Specific embodiments
[0042] To make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and do not limit the protection scope of the present invention.
[0043] The inventive concept of the present invention is as follows: The embodiments of the present invention provide a few-shot object detection method and device. Aiming at the technical problems in existing few-shot object detection, such as freezing the backbone network and being unable to flexibly use different-depth features to distinguish base classes and new classes, by introducing multiple mixture of experts, not only can deep features of each be obtained, but also shallow features with different depths can be obtained respectively. By combining deep features and shallow features and using the method of knowledge distillation to let each mixture of experts learn from each other, the object detection model can capture the key points of classification from different-depth features and achieve better classification effects. Aiming at the technical problem of class imbalance in few-shot learning, repeated upsampling and data augmentation methods are used to expand the data and improve the robustness of the object detection model.
[0044] As Figure 1 shown, a few-shot object detection method based on mixture of experts provided by the embodiment includes the following steps:
[0045] S110, initially train the object detection model using base class images.
[0046] In the embodiment, as Figure 2 shown, the few-shot object detection model can adopt the Faster RCNN network, which includes a backbone network for feature extraction, a Region Proposal Network (RPN module) for selecting candidate regions, and an ROI Head module for providing class prediction and bounding box regression prediction for each candidate region. The ROI Head module includes ROI pooling, a class classifier, and bounding box regression. This object detection model is used for object detection of base class images, and the object detection model is initially trained using a large number of labeled base class images. Specifically, during training, the cross-entropy loss is used for the classification results included in the object detection results, and the bounding box regression loss is used for the bounding box regression results included in the object detection results.
[0047] S120, use repeated upsampling and data augmentation methods to enhance the few-shot images of new classes to obtain new class enhanced images.
[0048] In the data augmentation step of the embodiment, three traditional techniques are applied to enrich the diversity of training data and prevent the model from overfitting, including horizontal flipping, rotation, and brightness adjustment. Horizontal flipping increases the sample size by symmetrically flipping the image along its vertical central axis; rotation enhances the model's recognition ability for multi-view objects by randomly rotating the image by a fixed angle; brightness adjustment enhances the model's adaptability to various lighting conditions by randomly changing the image brightness. To further improve the diversity of data and the robustness of the model, the embodiment also introduces five other augmentation techniques, including contrast adjustment, saturation adjustment, hue adjustment, addition of Gaussian noise, and addition of salt-and-pepper noise. Contrast adjustment highlights the key features in the image by enhancing the light-dark difference in the image; saturation adjustment makes the colors in the image more vivid, helping the model learn more color features; hue adjustment simulates the image changes under different color temperatures by changing the hue of the image; Gaussian noise improves the model's adaptability to noisy data by simulating noise interference in different environments; salt-and-pepper noise enhances the model's anti-interference ability to image noise by adding noise points. The effects of each are as Figure 3 shown. In practical applications, each augmentation scheme will be applied to the specified small-sample category images with a certain probability, and the augmentation process can be superimposed.
[0049] Regarding the class imbalance problem, there is a simple but effective Repeated Factor Sampling (RFS) method in past inventions to upsample the new class samples. RFS alleviates the extreme imbalance in the class data distribution and thus improves the model's learning ability for few-shot classes by repeatedly sampling the data of certain classes (i.e., increasing the occurrence times of class samples). In the method of the present invention, the idea of upsampling is adopted, and the data augmentation strategy is applied to the repeated sampling. This can not only solve the problem of unbalanced data distribution and alleviate the data forgetfulness, but also these data augmentation operations not only improve the training effect and generalization ability of the model, but also enhance the model's adaptability to complex environments, providing a solid foundation for subsequent model training and optimization. The application of data augmentation technology effectively improves the recognition accuracy of the model under various perspectives, lighting conditions, and noise interferences, demonstrating its key role in enhancing the robustness and performance of the model.
[0050] Based on this, when performing data augmentation in the embodiment, the small-sample images of the new class are first repeatedly upsampled to increase the number of images, and then data augmentation is performed on each set of upsampled small-sample images.
[0051] S130. Construct a small-sample object detection model based on mixture of experts, which includes a shared shallow feature network, a unique deep feature network for each mixture of experts, a feature fusion network, and an object detection network corresponding to each mixture of experts. Among them, the backbone network of the preliminarily trained object detection model is used to initialize the shared shallow feature network and the unique deep feature network for each mixture of experts.
[0052] Regarding the above-mentioned technical problems of freezing the weights of the backbone network and being unable to reasonably utilize features of different depths, the embodiments of the present invention propose to use a mixture of experts network to solve this technical problem. First, since in dealing with the class imbalance problem, a method of repeated upsampling and data augmentation is proposed, which can obtain different images of many new categories. Therefore, simply freezing the weights of the backbone network will instead make it impossible to well train this part of the images. Therefore, it is imperative to restore the trainability of the backbone network. At the same time, in order to enable the network to learn more comprehensive features, the embodiments of the present invention adopt a mixture of experts method to also fuse shallow features into deep features for joint learning.
[0053] As Figure 4 shown, the small-sample object detection model constructed based on mixture of experts includes a shared shallow feature network, a unique deep feature network for each mixture of experts, a feature fusion network, and an object detection network corresponding to each mixture of experts. Among them, the shared shallow feature network is used to extract shallow features of different depths from the input image, the unique deep feature network is used to extract their respective deep features based on the shallow features, the feature fusion network is used to fuse the shallow features of different depths and their respective deep features obtained by each mixture of experts to obtain their respective fused features, and the object detection network is used to predict the object detection results corresponding to each mixture of experts based on the fused features.
[0054] In the above model, let each mixture of experts share the first few layers of the backbone network (such as ResNet-101) in the object detection model as the shared shallow feature network, and independently own the last few layers as the unique deep feature network. The parameters of both the shared part and the independent part are initialized using the parameters of the backbone network in the pre-training stage. This is because the model of the present invention is mainly optimized for small-sample categories. In contrast, the attention to basic categories will decrease. Therefore, if the multi-expert network is trained completely from scratch, it will instead affect the basic categories. Therefore, using the weights of the backbone network pre-trained on basic categories for initialization can effectively avoid this problem. Subsequently, feature fusion is performed through the feature fusion network. Specifically, each mixture of experts will obtain shallow features of different depths from different parts of the shared shallow feature network and perform deep fusion with the deep features obtained by the unique deep feature network where represents the expert number. Since the sizes of the shallow features and the deep features are not the same and cannot be directly concatenated or computed, a convolutional layer of corresponding size is designed after each shallow feature for downsampling to achieve the alignment of the shallow features and the deep features. Then, the Hadamard product is used to fuse the shallow features and the deep features, and the feature is transformed into a logarithm to prevent problems such as gradient explosion caused by excessive data, obtaining the final fused feature of each mixture of experts. , where denotes the shallow feature obtained by the i -th mixture of experts, denotes the deep feature corresponding to the i -th mixture of experts, The symbol denotes downsampling, denotes the operation of taking the logarithm. Since the fused feature obtained by each mixture of experts is the fusion of shallow features of different depths and their respective deep features, the subsequent object detection networks of each mixture of experts are also exclusive. The object detection network corresponding to each mixture of experts includes a Region Proposal Network (RPN) and an ROI Head module that provides class prediction and bounding box regression prediction for each candidate region. The classification result and the bounding box regression result output by the ROI Head module constitute the object detection result corresponding to each mixture of experts.
[0055] S140. Use the base class images and the new class enhanced images to optimize the parameters of the above small-sample object detection model based on the mixture of experts. The loss function used for parameter optimization includes the object detection loss corresponding to each mixture of experts and the distillation loss between each mixture of experts.
[0056] In the embodiment, in order to better utilize features at different depths to facilitate category prediction, knowledge distillation is performed between any two mixture-of-experts (MoEs) so that they can learn from each other. Since each MoE has the same architecture at the deepest position of the network, it can be ensured that each MoE can play the role of a teacher or a student, and mutual knowledge distillation can be carried out between any two MoEs, providing a perfect opportunity for each expert to aggregate knowledge at different depths. During the training of the few-shot object detection model, the KL divergence is used as the distillation loss to enable mutual knowledge distillation learning between different MoEs. The KL divergence can quantify the difference between probability distributions and help adjust the model parameters to make the distribution generated by the model closer to the target distribution. The KL divergence can prompt each MoE to learn knowledge at different depths. Since each MoE learns features at different depths and obtains more comprehensive information, the difference between the base classes and the new classes will be clearer. Therefore, the learning method of MoEs can improve the confusion between classes. In addition, each MoE also needs to learn the ground truth of the real label, which can ensure that each MoE does not deviate from the original goal. Here, the traditional object detection loss can be maintained. The loss function used for parameter optimization includes the object detection loss corresponding to each MoE and the distillation loss between each pair of MoEs , which are respectively expressed as:
[0057] ;
[0058] ;
[0059] ;
[0060] where represents the sum of the classification losses corresponding to all MoEs. The classification loss can use the cross-entropy loss, represents the sum of the logistic regression losses corresponding to all MoEs, and represent hyperparameters, and represent the MoE indices, m represents the number of MoEs, and respectively represent the classification results of the i -th MoE and the j -th MoE, represents and the KL divergence loss between them.
[0061] Through the above loss function, the model can fully learn features of different depths.
[0062] S150, performing small sample target detection using a small sample target detection model with optimized parameters, and screening a final target detection result from target detection results of multiple mixed experts.
[0063] In the inference stage, multiple hybrid experts in the small sample target detection model are used to simultaneously detect targets on the image. Since multiple hybrid experts can make predictions in parallel, the inference time will not be increased too much. When the target detection results of multiple experts are obtained, the target detection results include logistic regression results and classification results. ,in k is the total number of base categories and new categories, Represents the classification result of the i-th hybrid expert for the k-th category. The target detection result of the hybrid expert with the highest category score is selected as the final target result. This is because the higher the score means the higher the confidence and the closer the image feature is to the center of the category. Therefore, this further strengthens the compactness within the class and alleviates the forgetfulness in small sample target detection. The specific expression is as follows:
[0064] ;
[0065] Among them, max refers to the maximum value, argmax refers to the position of the maximum value, and id refers to the number of the hybrid expert finally selected.
[0066] The small sample target detection method based on hybrid experts provided in the above embodiment uses hybrid experts to cleverly combine features of different depths, learn knowledge of different depths, and improve the detection performance of the model in the small sample target detection task; at the same time, a method combining repeated upsampling and data enhancement is proposed, which not only effectively alleviates the overfitting problem caused by the upsampling process, but also enhances the diversity of small sample data and improves the robustness of the model.
[0067] The core of the small sample target detection method based on hybrid experts provided in the above embodiment is to use the hybrid expert method and combine the repeated upsampling and data enhancement methods to solve a series of problems in the field of small sample target detection. The present invention is not limited to the above-mentioned backbone network, target detection loss and distillation loss. Simply replacing the backbone network, simply changing the KL divergence and cross entropy loss function, or simply increasing or deleting the number of hybrid experts are all within the protection scope of the present invention.
[0068] like Figure 5As shown in the figure, the embodiment also provides a few-shot object detection device 500 based on mixture of experts, including: a preliminary training module 510, a data augmentation module 520, a model construction module 530, a model training module 540, and an object detection module 550. Among them, the preliminary training module 510 is used to preliminarily train an object detection model using base-class images; the data augmentation module 520 is used to enhance few-shot images of new classes by adopting repeated upsampling and data augmentation methods to obtain new-class enhanced images; the model construction module 530 is used to construct a few-shot object detection model based on mixture of experts, which includes a shared shallow feature network, a unique deep feature network for each mixture of experts, a feature fusion network, and an object detection network corresponding to each mixture of experts. Among them, the shared shallow feature network is used to extract shallow features of different depths from the input image, the unique deep feature network is used to extract their respective deep features based on the shallow features, the feature fusion network is used to fuse the shallow features of different depths and their respective deep features obtained by each mixture of experts to obtain their respective fused features, and the object detection network is used to predict the object detection results corresponding to each mixture of experts based on the fused features; among them, the backbone network of the preliminarily trained object detection model is used to initialize the shared shallow feature network and the unique deep feature network of each mixture of experts; the model training module 540 is used to optimize the parameters of the above-mentioned few-shot object detection model based on mixture of experts using base-class images and new-class enhanced images. The loss function used during parameter optimization includes the object detection losses corresponding to each mixture of experts and the distillation losses between each mixture of experts; the object detection module 550 is used to perform few-shot object detection using the few-shot object detection model with optimized parameters and screen the final object detection results from the object detection results of multiple mixtures of experts.
[0069] It should be noted that when the above-mentioned few-shot object detection device based on mixture of experts performs few-shot object detection, the above-mentioned division of each functional module should be used for illustration. The above functions can be allocated to different functional modules according to needs, that is, the internal structure of the terminal or server is divided into different functional modules to complete all or part of the functions described above. In addition, the above-mentioned few-shot object detection device based on mixture of experts and the few-shot object detection method embodiment based on mixture of experts belong to the same concept. For the specific implementation process, please refer to the few-shot object detection method embodiment based on mixture of experts, which will not be elaborated here.
[0070] Based on the same inventive concept, the embodiment also provides a computing device, including a memory and one or more processors. An executable code is stored in the memory. When the one or more processors execute the executable code, it is used to implement the above-mentioned few-shot object detection method based on mixture of experts, specifically including the following steps:
[0071] S110. Initially train the object detection model using the basic category images;
[0072] S120. Enhance the small-sample images of the new category by using repeated upsampling and data augmentation to obtain the enhanced images of the new category;
[0073] S130. Construct a small-sample object detection model based on mixture of experts, which includes a shared shallow feature network, unique deep feature networks for each mixture of experts, a feature fusion network, and object detection networks corresponding to each mixture of experts. The shared shallow feature network is used to extract shallow features of different depths from the input image. The unique deep feature networks are used to extract their respective deep features based on the shallow features. The feature fusion network is used to fuse the shallow features of different depths and their respective deep features obtained by each mixture of experts to obtain their respective fused features. The object detection networks are used to predict the object detection results corresponding to each mixture of experts based on the fused features. Among them, initialize the shared shallow feature network and the unique deep feature networks for each mixture of experts using the backbone network of the initially trained object detection model;
[0074] S140. Optimize the parameters of the above-mentioned small-sample object detection model based on mixture of experts using the basic category images and the enhanced images of the new category. The loss function used during parameter optimization includes the object detection losses corresponding to each mixture of experts and the distillation losses between each mixture of experts;
[0075] S150. Perform small-sample object detection using the small-sample object detection model with optimized parameters, and screen the final object detection results from the object detection results of multiple mixtures of experts.
[0076] The computing device provided in the embodiment, at the hardware level, in addition to including a processor and a memory, also includes other hardware required for other services such as an internal bus, a network interface, and a memory. The memory is a non-volatile memory. The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it to implement the small-sample object detection method based on mixture of experts described in S110-S150 above. Of course, in addition to the software implementation method, the present invention does not exclude other implementation methods, such as logic devices or a combination of software and hardware, etc. That is to say, the execution subject of the following processing flow is not limited to each logic unit, and can also be hardware or a logic device.
[0077] Based on the same inventive concept, the embodiment also provides a computer-readable storage medium, on which a program is stored. When the program is executed by a processor, it implements the small-sample object detection method based on mixture of experts, specifically including the following steps:
[0078] S110. Initially train the object detection model using the basic category images;
[0079] S120, perform upsampling and data augmentation repeatedly on the small-sample images of the new category to obtain enhanced images of the new category;
[0080] S130, construct a small-sample object detection model based on mixture of experts, which includes a shared shallow feature network, unique deep feature networks for each mixture of experts, a feature fusion network, and object detection networks corresponding to each mixture of experts. The shared shallow feature network is used to extract shallow features of different depths from the input image, the unique deep feature networks are used to extract their respective deep features based on the shallow features, the feature fusion network is used to fuse the shallow features of different depths and their respective deep features obtained by each mixture of experts to obtain their respective fused features, and the object detection networks are used to predict the object detection results corresponding to each mixture of experts based on the fused features. Among them, initialize the shared shallow feature network and the unique deep feature networks for each mixture of experts using the backbone network of the preliminarily trained object detection model;
[0081] S140, optimize the parameters of the above small-sample object detection model based on mixture of experts using the base-category images and the enhanced images of the new category. The loss function used during parameter optimization includes the object detection losses corresponding to each mixture of experts and the distillation losses between each mixture of experts;
[0082] S150, perform small-sample object detection using the small-sample object detection model with optimized parameters, and screen the final object detection results from the object detection results of multiple mixtures of experts.
[0083] In the embodiment, the computer-readable medium includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data.
[0084] The specific embodiments described above have elaborated on the technical solutions and beneficial effects of the present invention. It should be understood that the above are only the most preferred embodiments of the present invention and are not used to limit the present invention. Any modifications, supplements, equivalent replacements, etc. made within the scope of the principles of the present invention shall be included within the protection scope of the present invention.
Claims
1. A small-sample object detection method based on mixture of experts, characterized in that, It includes the following steps: Use the basic category images to preliminarily train the object detection model; Adopt repeated upsampling and data augmentation methods to enhance the small-sample images of the new category to obtain new-category enhanced images, including: first, perform repeated sampling on the small-sample images of the new category, that is, increase the occurrence times of the small-sample images of the new category to increase the number of images, and then perform data augmentation on each set of small-sample images with increased quantity; Build a few-shot object detection model based on mixture of experts, which includes a shared shallow feature network, a unique deep feature network for each mixture of experts, a feature fusion network, and an object detection network corresponding to each mixture of experts. The shared shallow feature network is used to extract shallow features of different depths from the input image; the unique deep feature network is used to extract their respective deep features based on the shallow features; the feature fusion network is used to fuse the shallow features of different depths and their respective deep features obtained by each mixture of experts to obtain their respective fused features. Specifically, it includes: first, downsampling the shallow features of different depths obtained by each mixture of experts to adjust them to the same size as the deep features, and then fusing the downsampled shallow features with the deep features to obtain the fused features Among them, represents the shallow features obtained by the i-th mixture of experts, represents the deep features corresponding to the i-th mixture of experts, The symbol represents the Hadamard product, and ω(·) represents the downsampling process, represents the calculation of taking the logarithm; the object detection network is used to predict the object detection results corresponding to each mixture of experts based on the fused features. Among them, the backbone network of the preliminarily trained object detection model is used to initialize the shared shallow feature network and the unique deep feature network of each mixture of experts, so that each mixture of experts shares the first few layers of the backbone network in the object detection model as the shared shallow feature network and independently owns the last few layers as the unique deep feature network to restore the trainability of the backbone network; Use the basic category images and the new-category enhanced images to optimize the parameters of the above-mentioned small-sample object detection model based on mixture of experts. The loss function used in parameter optimization includes the object detection losses corresponding to each mixture of experts and the distillation loss between each mixture of experts. Use the KL divergence as the distillation loss to enable mutual knowledge distillation learning between different mixtures of experts; Perform few-shot object detection using a few-shot object detection model optimized with parameters, and screen the final object detection result from the object detection results of multiple mixture experts, where the object detection result includes a logistic regression result and a classification result where k refers to the total number of base classes and new classes represents the classification result of the i-th mixture expert for the k-th class, and select the object detection result of the mixture expert with the highest class score as the final object result 2. The small-sample object detection method based on a mixture of experts according to claim 1, wherein The object detection model uses the Faster RCNN network. When performing preliminary training on it, the cross-entropy loss is used for the classification results included in the object detection results, and the bounding box regression loss is used for the bounding box regression results included in the object detection results.
3. The small-sample object detection method based on a mixture of experts according to claim 1, characterized in that The data augmentation methods include: horizontal flipping, rotation, brightness adjustment, contrast adjustment, saturation adjustment, hue adjustment, addition of Gaussian noise and salt-and-pepper noise.
4. The small-sample object detection method based on a mixture of experts according to claim 1, characterized in that The object detection network corresponding to each mixture of experts includes a region proposal network and an ROI Head module that provides class prediction and bounding box regression prediction for each candidate region. The classification results and bounding box regression results output by the ROI Head module form the object detection results corresponding to each mixture of experts.
5. The small-sample object detection method based on a mixture of experts according to claim 1, wherein The loss function l used during parameter optimization total includes the object detection loss l corresponding to each mixture of experts td and the distillation loss l between each mixture of experts mu , which are respectively expressed as: l total =αl td +βl mu l td =l cls +l reg Among them, l cls represents the sum of the classification losses corresponding to all mixture experts, l reg represents the sum of the logistic regression losses corresponding to all mixture experts, α and β represent hyperparameters, i and j represent mixture expert indices, m represents the number of mixture experts, p i and p j represent the classification results of the i-th mixture expert and the j-th mixture expert respectively, KL(p i |p j ) represents the KL divergence loss between p i and p j .
6. A small-sample object detection device based on a mixture of experts, characterized in that, It includes: A preliminary training module, which is used to preliminarily train the object detection model using the basic category images; A data augmentation module, which is used to enhance the small-sample images of the new category by adopting repeated upsampling and data augmentation methods to obtain new-category enhanced images, including: first, perform repeated sampling on the small-sample images of the new category, that is, increase the occurrence times of the small-sample images of the new category to increase the number of images, and then perform data augmentation on each set of small-sample images with increased quantity; The model construction module is used to construct a few-shot object detection model based on mixture of experts, which includes a shared shallow feature network, the unique deep feature network of each mixture of experts, a feature fusion network, and the object detection network corresponding to each mixture of experts. The shared shallow feature network is used to extract shallow features of different depths from the input image; the unique deep feature network is used to extract their respective deep features based on the shallow features; the feature fusion network is used to fuse the shallow features of different depths and their respective deep features obtained by each mixture of experts to obtain their respective fused features, specifically including: first, downsampling the shallow features of different depths obtained by each mixture of experts to adjust them to the same size as the deep features, and then fusing the downsampled shallow features with the deep features to obtain the fused features Among them, represents the shallow features obtained by the i-th mixture of experts, represents the deep features corresponding to the i-th mixture of experts, The symbol represents the Hadamard product, and ω(·) represents the downsampling process, represents the operation of taking the logarithm; the object detection network is used to predict the object detection results corresponding to each mixture of experts based on the fused features; among them, the backbone network of the preliminarily trained object detection model is used to initialize the shared shallow feature network and the unique deep feature network of each mixture of experts, so that each mixture of experts shares the first few layers of the backbone network in the object detection model as the shared shallow feature network and independently owns the last few layers as the unique deep feature network to restore the trainability of the backbone network; A model training module, which is used to optimize the parameters of the above-mentioned small-sample object detection model based on mixture of experts using the basic category images and the new-category enhanced images. The loss function used in parameter optimization includes the object detection losses corresponding to each mixture of experts and the distillation loss between each mixture of experts. Use the KL divergence as the distillation loss to enable mutual knowledge distillation learning between different mixtures of experts; A target detection module, which is used to perform few-shot target detection using a few-shot target detection model with optimized parameters, and screen the final target detection result from the target detection results of multiple mixture of experts, where the target detection result includes a logistic regression result and a classification result where k refers to the total number of base classes and new classes represents the classification result of the i-th mixture of experts for the k-th class, and the target detection result of the mixture of experts with the highest class score is selected as the final target result 7. A computing device, comprising a memory and one or more processors, wherein executable code is stored in the memory, characterized in that, When the one or more processors execute the executable code, it is used to implement the small-sample object detection method based on mixture of experts according to any one of claims 1-5.
8. A computer-readable storage medium, characterized in that, A program is stored thereon, and when the program is executed by a processor, it implements the small-sample object detection method based on mixture of experts according to any one of claims 1-5.
Citation Information
Patent Citations
Small sample target detection method based on target suggestion box increment
CN115439645A
Long-tail image recognition method based on neural network multi-expert hierarchical logic fusion
CN118711029A