Passive field adaptive target detection model generation method based on class prototype alignment and target detection method
By adopting a class prototype alignment method in passive field adaptive object detection, the class prototype is iteratively updated to align the characteristics and prediction results of student model and teacher model, the model training collapse problem caused by pseudo-label noise is solved, and the accuracy of the object detection results is improved.
Patent Information
- Application Number
- CN202510141780.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-08
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-02-08
AI Technical Summary
In the existing adaptive object detection method in the passive field, the selection of pseudo-labels depends on the predicted filtering results of the teacher model, resulting in noise in the pseudo-labels, causing incorrect supervision signals, resulting in model training collapse, affecting the accuracy of the target detection results.
Using a method based on class prototype alignment, the features extracted by the student model and teacher model are further extracted through the object detection box, and the first and second types of prototypes belonging to the same target category are iteratively updated. The alignment loss of the prototype is used to construct the prototype to align the feature extraction results and prediction results of the student model and teacher model.
Through iterative update of class prototypes, historical knowledge is included in the training process to avoid incorrect supervision signals brought by pseudo-labels, and improve the stability of model training and the accuracy of target detection results.
Smart Images

Figure CN120147606A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and particularly to a method for generating a source-free domain adaptation object detection model based on class prototype alignment and an object detection method. Background Art
[0002] Due to the availability of widely labeled datasets for model training, object detection technology has made great progress. To reduce the need for manual data annotation, there is unsupervised domain adaptation (UDA) technology for object detection. UDA usually assumes that both the source domain and target domain data can be accessed simultaneously for effective adaptation. However, in the real world, due to data privacy, unreliable data transmission, or data security reasons, the source data may not be accessible. For example, the data used to train large language models (LLMs) is usually huge and highly valuable, making it unsuitable for open sharing. To overcome the problem of unavailable source domain data, source-free object detection (SFOD) technology has emerged, which relies only on the model weights trained in the source domain and unlabeled data in the target domain to achieve model adaptation from the source domain to the target domain and obtain good object detection results in the target domain.
[0003] Most existing SFOD methods adopt the mean teacher framework, in which both the teacher and student models are initialized with a pre-trained source domain model, and the unlabeled data in the target domain is used to update the student and teacher models. In this framework, the prediction results of the teacher model are filtered and used as pseudo-labels to train the student model, and then the parameters of the student model are used to update the teacher model.
[0004] In the prior art, for a mean teacher framework, the selection of pseudo-labels depends on the prediction filtering results of the teacher model, which makes there be a certain degree of noise in the pseudo-labels, thus bringing incorrect supervision signals, leading to model training collapse, and ultimately affecting the accuracy of the object detection results of the model based on training completion. Summary of the Invention
[0005] The present invention provides a method for generating a source-free domain adaptation object detection model based on class prototype alignment and an object detection method, so as to solve the defect of low accuracy of object detection results of the object detection model in the prior art and achieve improvement in the accuracy of object detection results of the object detection model.
[0006] The present invention provides a method for generating a source-free domain adaptation object detection model based on class prototype alignment, including: Input the target domain sample image into the first feature extraction module in the teacher model to obtain the first extracted feature. Input the target domain sample image into the second feature extraction module of the student model to obtain the second extracted feature. The initial parameters of the teacher model and the student model are trained based on the source domain data, and the source domain data includes source domain sample images and the corresponding object detection labels of the source domain sample images; Extract the first target feature and the second target feature from the first extracted feature and the second extracted feature respectively based on the object detection bounding box. The object detection bounding box is obtained based on the result of the teacher model performing object detection on the target domain sample image. Input the first target feature into the first target category prediction module of the teacher model to obtain the first target category prediction result. Input the second target feature into the second target category prediction module of the student model to obtain the second target category prediction result; Iteratively update the first type of prototype of the preset target category based on each of the first target features and the first target category prediction results of the same preset target category. Iteratively update the second type of prototype of the preset target category based on each of the second target features and the second target category prediction results corresponding to the same target preset category; Obtain the alignment loss between the first type of prototype and the second type of prototype. Determine the training loss based on the alignment loss. Update the parameters of the student model based on the training loss. Update the parameters of the teacher model based on the updated parameters of the student model. Obtain the teacher model after multiple parameter updates as the object detection model. The alignment loss reflects the similarity between the first type of prototype and the second type of prototype and the class prototype knowledge distillation from the teacher model to the student model.
[0007] According to a method for generating a domain - free adaptive object detection model based on class prototype alignment provided by the present invention, before inputting the target domain sample image into the first feature extraction module in the teacher model, it includes: Perform the first data augmentation process on the target domain sample image; Before inputting the target domain sample image into the second feature extraction module of the student model, it includes: Perform the second data augmentation process on the target domain sample image; Wherein, the data augmentation intensity of the second data augmentation process is greater than that of the first data augmentation process.
[0008] According to a method for generating a domain - free adaptive object detection model based on class prototype alignment provided by the present invention, before extracting the first target feature and the second target feature from the first extracted feature and the second extracted feature respectively based on the object detection bounding box, it includes: Determine the target detection box based on the intersection over union between the detection boxes obtained by performing target detection on the target domain sample image output by the teacher model.
[0009] According to a method for generating a domain - adaptive target detection model without source based on class prototype alignment provided by the present invention, the iteration of the first class prototype of the preset target category based on each of the first target features and the first target category prediction results of the same preset target category includes: Normalize the first target category prediction result corresponding to the preset target category in the current target domain sample image to obtain a normalized first target category prediction result, and determine the current first - type feature based on the normalized first target category prediction result and the first target feature corresponding to the preset target category in the current target domain sample image; Iteratively update the current first - type prototype based on the first - type feature; Wherein, the initial first - type prototype is the first - type feature corresponding to the target and sample image input to the teacher model for the first time; The iteration and update of the second - type prototype of the preset target category based on each of the second target features and the second target category prediction results corresponding to the same target preset category includes: Normalize the second target category prediction result corresponding to the preset target category in the current target domain sample image to obtain a normalized second target category prediction result, and determine the current second - type feature based on the normalized second target category prediction result and the second target feature corresponding to the preset target category in the current target domain sample image; Iteratively update the current first - type prototype based on the second - type feature; Wherein, the initial second - type prototype is the second - type feature corresponding to the target and sample image input to the teacher model for the first time.
[0010] According to a method for generating a domain - adaptive target detection model without source based on class prototype alignment provided by the present invention, the obtaining of the alignment loss between the first - type prototype and the second - type prototype includes: Determine a prototype contrast loss based on the similarity between the first - type prototype and the second - type prototype; Input the first - type prototype into the teacher model and the student model respectively to obtain a first - type prediction result and a second - type prediction result, input the second - type prototype into the teacher model and the student model respectively to obtain a third - type prediction result and a fourth - type prediction result; Determine the prototype distillation loss based on the distribution differences between any two of the first type of prediction results, the second type of prediction results, the third type of prediction results, and the fourth type of prediction results; Determine the alignment loss based on the prototype contrast loss and the prototype distillation loss.
[0011] The present invention also provides an object detection method based on the method for generating a domain - adaptive object detection model with class prototype alignment according to any one of the above, including: Obtain an image to be detected, and input the image to be detected into the object detection model; Obtain the object detection result of the object detection model for the image to be detected.
[0012] The present invention also provides a device for generating a domain - adaptive object detection model with class prototype alignment, including: A feature extraction module, configured to input a target - domain sample image into a first feature extraction module in a teacher model to obtain a first extracted feature, and input the target - domain sample image into a second feature extraction module in a student model to obtain a second extracted feature. The initial parameters of the teacher model and the student model are trained based on source - domain data, and the source - domain data includes source - domain sample images and the object detection labels corresponding to the source - domain sample images; A class prediction module, configured to extract a first target feature and a second target feature from the first extracted feature and the second extracted feature respectively based on object detection frames. The object detection frames are obtained based on the object detection results of the teacher model for the target - domain sample images. Input the first target feature into a first target class prediction module of the teacher model to obtain a first target class prediction result, and input the second target feature into a second target class prediction module of the student model to obtain a second target class prediction result; A class prototype update module, configured to iteratively update a first - type prototype of a preset target class based on each of the first target features and the first target class prediction results of the same preset target class, and iteratively update a second - type prototype of the preset target class based on each of the second target features and the second target class prediction results corresponding to the same target preset class; A parameter update module, configured to obtain the alignment loss between the first - type prototype and the second - type prototype, determine a training loss based on the alignment loss, update the parameters of the student model based on the training loss, update the parameters of the teacher model based on the updated parameters of the student model, and obtain the teacher model after multiple parameter updates as the object detection model. The alignment loss reflects the similarity between the first - type prototype and the second - type prototype and the class prototype knowledge distillation from the teacher model to the student model.
[0013] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the above-mentioned method for generating a passive domain adaptive object detection model based on class prototype alignment is implemented.
[0014] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the above-mentioned method for generating a passive domain adaptive object detection model based on class prototype alignment is implemented.
[0015] The present invention also provides a computer program product, including a computer program. When the computer program is executed by a processor, the above-mentioned method for generating a passive domain adaptive object detection model based on class prototype alignment is implemented.
[0016] In the method for generating a passive domain adaptive object detection model based on class prototype alignment and the object detection method provided by the present invention, during the training process using a teacher model and a student model, the object detection boxes are used to further extract object features from the features extracted by the student model and the teacher model. Based on the features of the extracted object detection boxes and the object category prediction results output by the student model and the teacher model based on the features of the object detection boxes, the first class prototype and the second class prototype belonging to the same object category are iteratively updated. The alignment loss of the prototypes is constructed using the class prototypes to align the feature extraction results and prediction results of the student model and the teacher model. In this way, through the iterative update of the class prototypes, historical knowledge is incorporated into the training process, thereby avoiding the problem that the incorrect supervision signal brought by the pseudo labels causes the collapse of the model training, improving the performance of the trained object detection model, and thus improving the accuracy of the object detection results output by the object detection model. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0018] Figure 1 is a schematic flowchart of the method for generating a passive domain adaptive object detection model based on class prototype alignment provided by the present invention.
[0019] Figure 2 is a schematic overall architecture diagram of the method for generating a passive domain adaptive object detection model based on class prototype alignment provided by the present invention.
[0020] Figure 3 It is a schematic diagram of model training in the method for generating a passive domain adaptive object detection model based on class prototype alignment provided by the present invention.
[0021] Figure 4 It is a schematic diagram of model testing in the method for generating a passive domain adaptive object detection model based on class prototype alignment provided by the present invention.
[0022] Figure 5 It is a schematic structural diagram of the device for generating a passive domain adaptive object detection model based on class prototype alignment provided by the present invention.
[0023] Figure 6 It is a schematic structural diagram of the electronic device provided by the present invention. Detailed implementation manners
[0024] To make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below with reference to the accompanying drawings in the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art without making creative efforts based on the embodiments in the present invention belong to the scope of protection of the present invention.
[0025] The following combines Figures 1-4 to describe the method for generating a passive domain adaptive object detection model based on class prototype alignment provided by the present invention. As Figure 1 shown, the method includes the steps: S110. Input the target domain sample image into the first feature extraction module in the teacher model to obtain the first extracted feature, and input the target domain sample image into the second feature extraction module in the student model to obtain the second extracted feature. The initial parameters of the teacher model and the student model are trained based on the source domain data, and the source domain data includes the source domain sample image and the target detection label corresponding to the source domain sample image; S120. Extract the first target feature and the second target feature from the first extracted feature and the second extracted feature respectively based on the target detection box. The target detection box is obtained based on the result of the teacher model performing object detection on the target domain sample image. Input the first target feature into the first target category prediction module in the teacher model to obtain the first target category prediction result, and input the second target feature into the second target category prediction module in the student model to obtain the second target category prediction result; S130. Determine the first class prototype of the preset target category based on each first target feature corresponding to the same preset target category in the first target category prediction result, and determine the second class prototype of the preset target category based on each second target feature corresponding to the same preset target category in the second target category prediction result; S140. Obtain the alignment loss between the first type of prototypes and the second type of prototypes, determine the training loss based on the alignment loss, update the parameters of the student model based on the training loss, update the parameters of the teacher model based on the parameters of the updated student model, obtain the teacher model after multiple parameter updates as the object detection model, and its loss reflects the similarity between the first type of prototypes and the second type of prototypes as well as the class prototype knowledge distillation from the teacher model to the student model.
[0026] The method provided by the present invention realizes incorporating historical knowledge into the training process through iterative update of class prototypes, thereby avoiding the problem that incorrect supervision signals brought by pseudo-labels lead to the collapse of model training, improving the performance of the trained object detection model, and thus improving the accuracy of the object detection results output by the object detection model.
[0027] The initial parameters of the teacher model and the student model are trained based on the labeled source domain data. In the unsupervised domain adaptation object detection task, the samples for training the model include two parts, namely the source domain labeled samples and the target domain unlabeled samples , denoted as , where represents the sample, represents the corresponding label, including the target box and the category. represents the total number of labeled samples. , where represents the unlabeled sample, represents the total number of unlabeled samples. The source domain model weights are trained using the labeled samples in the source domain, and then the student model and the teacher model are initialized with the source domain model weights in the target domain, and only the target domain unlabeled samples are used to train the student and teacher models.
[0028] As Figure 3 shown, during the training process of the student model and the teacher model, the sample data in the target domain are respectively input into the teacher model and the student model, and the parameters of the student model and the teacher model are updated based on the outputs of the teacher model and the student model. After training is completed, the teacher model is used as the object detection model to perform object detection on the data in the target domain, as Figure 4 shown. The training processes of the teacher model and the student model are specifically described below.
[0029] As Figure 2 shown, before the sample images in the target domain are respectively input into the teacher model and the student model, they can be enhanced, that is, before the target domain sample images are input into the first feature extraction module in the teacher model, it includes: Performing the first data enhancement process on the target domain sample images; Before the target domain sample image is input into the second feature extraction module of the student model, it includes: Perform second data augmentation processing on the target domain sample image.
[0030] Among them, the data augmentation intensity of the second data augmentation processing is greater than that of the first data augmentation processing. That is to say, after weakly augmenting the data of the target domain sample image, it is input into the teacher model, and after strongly augmenting the data of the target domain sample image, it is input into the student model.
[0031] Data Augmentation generates new training samples by transforming and augmenting existing data, thereby enhancing the diversity and quantity of the dataset. These transformations can be geometric transformations, color transformations, noise addition, etc., enabling the model to see more types of data during the training process, thus improving the generalization ability and robustness of the model.
[0032] For the target domain sample image after data augmentation, it is sequentially input into the student model and the teacher model for processing. Each time the target domain sample image is input into the student model and the teacher model, the class prototype is iteratively updated once based on the outputs of the student model and the teacher model. Specifically, both the student model and the teacher model have a feature extraction module and a target class prediction module. The feature extraction module is used to extract features and output target detection boxes based on the extracted features. For example, it can adopt Figure 2 the RPN model architecture shown in Figure 2 and Figure 3 the RCNN model architecture shown in. It can be understood that the architectures of the student model and the teacher model are not limited to Figure 2 and Figure 3 the combination of RPN + RCNN shown in, and can also be other neural network model architectures.
[0033] In object detection, target detection boxes with high localization quality often exist in dense candidate boxes. In the method provided by the present invention, a large number of target detection boxes generated by the teacher model are screened, and target detection boxes with high localization quality are selected, and then the features extracted by the feature extraction module are further extracted based on the selected target detection boxes.
[0034] Specifically, before extracting the first target feature and the second target feature from the first extracted feature and the second extracted feature respectively based on the target detection box, it includes: Determine the target detection box based on the intersection over union between the detection boxes obtained by performing object detection on the target domain sample image output by the teacher model.
[0035] The process of obtaining the target detection boxes by selecting fewer detection boxes generated by the teacher model can be expressed by the formula: ; ; where, represents the set intersection over union threshold, represents the i-th detection box, and N is the total number of detection boxes generated by the teacher model.
[0036] After selecting the target detection boxes with higher localization quality from the detection boxes, these target detection boxes are used to iteratively update the class prototypes, so as to store the historical knowledge of the teacher model and the student model.
[0037] By screening out the target detection boxes with high localization quality, it is not only beneficial to construct more representative class prototypes, but also can improve the localization quality of the pseudo-labels.
[0038] Iteratively updating the first-class prototype of the preset target category based on each first target feature and the first target category prediction result of the same preset target category includes: Normalizing the first target category prediction result corresponding to the preset target category in the current target domain sample image to obtain the normalized first target category prediction result, and determining the current first-class feature based on the normalized first target category prediction result and the first target feature corresponding to the preset target category in the current target domain sample image; Iteratively updating the current first-class prototype based on the first-class feature; where, the initial first-class prototype is the first-class feature corresponding to the target and sample image input to the teacher model for the first time; Iteratively updating the second-class prototype of the preset target category based on each second target feature and the second target category prediction result corresponding to the same target preset category includes: Normalizing the second target category prediction result corresponding to the preset target category in the current target domain sample image to obtain the normalized second target category prediction result, and determining the current second-class feature based on the normalized second target category prediction result and the second target feature corresponding to the preset target category in the current target domain sample image; Iteratively updating the current first-class prototype based on the second-class feature; where, the initial second-class prototype is the second-class feature corresponding to the target and sample image input to the teacher model for the first time.
[0039] The iterative update processes of the first-class prototype of the teacher model and the second-class prototype of the student model are similar. Here, the iterative update process of the second-class prototype of the student model is taken as an example for specific description.
[0040] After inputting the current target domain sample image into the student model, extract the features corresponding to the target detection box from the second extracted features of the features of the feature extraction module of the student model to obtain the second target feature. to obtain the second target feature.
[0041] In a possible implementation, when extracting the second target feature from the second extracted features, first further extract the corresponding features from the second extracted features based on the position of the target detection box, and then perform average pooling on the further extracted features. This process can be expressed by the formula: ; where represents the second target feature, represents the strong augmentation of the data, represents extracting the features corresponding to the candidate box in the feature map of the student model, represents using average pooling, and average pooling can achieve dimensionality reduction of the features.
[0042] Input the obtained second target feature into the second target category prediction module in the student model to obtain the prediction result, which can be expressed by the formula: , represents the processing of the second target category prediction model, is the second target category prediction result, and the second target category prediction result is a feature vector, and the elements therein include the probabilities belonging to different preset target categories.
[0043] Perform normalization processing on the second target category prediction results corresponding to the same class in the target detection box to obtain the normalized second target category prediction result , which can be expressed by the formula: ; where represents the second target category prediction result corresponding to the i-th target detection box, and M is the number of target detection boxes.
[0044] Through the above processing, the normalized second target category prediction results of each preset target category and the corresponding second target features can be obtained. Combine the normalized second target category prediction results of each preset target category and the corresponding second target features to obtain the normalized second prediction result matrix and the second target feature matrix , take as the second type of feature.
[0045] For each target domain sample image, after inputting it into the student model, a second type of feature can be obtained. The first second type of feature is used as the initial value of the second type of prototype. After that, each time a second type of feature is obtained, the second type of prototype is iteratively updated based on this second type of feature. The iterative update can adopt the momentum update (EMA) method. The iterative update process of the second type of prototype can be expressed by the formula: ; where t represents the number of iterations, is the momentum update coefficient of EMA, represents the updated second type of prototype at the t-th iteration.
[0046] For all C + 1 preset target categories (including the background category), using the teacher model and the student model respectively, the first type of prototype and the second type of prototype can be obtained, where C is the number of target domain categories, is the class feature dimension of a single class, .
[0047] As can be seen from the previous description, the first type of prototype and the second type of prototype retain the historical knowledge of the target domain sample images in the teacher model and the student model. The method provided by the present invention further adds the alignment loss between the first type of prototype and the second type of prototype to the total loss of model training, and uses the class prototypes to construct the prototype contrast loss and the prototype distillation loss to align the representations and prediction results of the student model and the teacher model, so as to promote the knowledge transfer from the source domain to the target domain, overcome the problem that the model training collapses due to the incorrect decrease of the model parameter gradient caused by the pseudo-label noise, and improve the performance of the model in the target domain.
[0048] Specifically, obtaining the alignment loss between the first type of prototype and the second type of prototype includes: Determining the prototype contrast loss based on the similarity between the first type of prototype and the second type of prototype; Inputting the first type of prototype into the teacher model and the student model respectively to obtain the first type of prediction result and the second type of prediction result, and inputting the second type of prototype into the teacher model and the student model respectively to obtain the third type of prediction result and the fourth type of prediction result; Determining the prototype distillation loss based on the distribution differences between any two of the first type of prediction result, the second type of prediction result, the third type of prediction result and the fourth type of prediction result; Determining the alignment loss based on the prototype contrast loss and the prototype distillation loss.
[0049] The prototype contrast loss between the first type of prototype and the second type of prototype can be determined by evaluating the similarity between the prototypes of each two categories through cosine similarity. Specifically, it can be expressed by the formula: ; ; Among them, represents the prototype contrast loss, represents S calculated based on the data corresponding to the i-th category in the first type of prototype and the second type of prototype, represents S calculated based on the data corresponding to the i-th category in the first type of prototype and the data corresponding to the k-th category in the second type of prototype. T represents the contrast loss temperature parameter, which is used to adjust the emphasis relationship of the contrast loss function.
[0050] Furthermore, in the method provided by the present invention, a prototype distillation loss is also added to the alignment loss of the class prototypes to promote knowledge distillation from the teacher model to the student model and align the prediction results of different class prototypes. The specific calculation formula of the prototype distillation loss can be: ; ; Among them, and represent the processing of the second target category prediction module in the student model and the first target category prediction module in the teacher model respectively, represents the KL (Kullback-Leibler) divergence between x and y.
[0051] After obtaining the alignment loss between the first type of prototype and the second type of prototype, it can be summed with the loss obtained based on the pseudo-label to obtain the total training loss for updating the parameters of the student model. That is , represents the total training loss, represents the loss obtained based on the pseudo-label.
[0052] The loss obtained based on the pseudo-label can be constructed by using the existing loss construction method based on the pseudo-label in the source-free domain adaptation object detection task. For example: ; Among them, represents the loss obtained based on the pseudo-label for the target domain sample image obtained, represents the pseudo-label, represents the classification loss of RPN, represents the regression loss of RPN, represents the classification loss of ROI (region of interest), represents the regression loss of ROI.
[0053] Through the training loss After updating the parameters of the student model, the parameters of the teacher model are updated based on the updated parameters of the student model. Specifically, in each round, the parameters of the student model can be updated by the training loss obtained from each batch in turn (a round includes multiple batches, and a batch includes a target domain sample image). After each round, the teacher model is updated using the parameters of the student model. After training for multiple rounds (e.g., 20 rounds), the training ends.
[0054] The update of the student model parameters can be expressed by the formula: The update of the teacher model parameters can be expressed by the formula: where, are the parameters of the student model, are the parameters of the teacher model, is the learning rate, is the EMA update coefficient.
[0055] After the training ends, the teacher model is used as the object detection model for testing and actual object detection tasks, as Figure 4 shown.
[0056] Based on the method for generating a domain - adaptive object detection model without source domain based on class prototype alignment provided by the present invention, the present invention also provides an object detection method, including: Obtain the image to be detected, and input the image to be detected into the object detection model; Obtain the object detection result of the object detection model for the image to be detected.
[0057] wherein, the object detection model is the teacher model after training in the method for generating a domain - adaptive object detection model without source domain based on class prototype alignment provided by the present invention.
[0058] Next, the apparatus for generating a domain - adaptive object detection model without source domain based on class prototype alignment provided by the present invention is described. The apparatus for generating a domain - adaptive object detection model without source domain based on class prototype alignment described below can be correspondingly referred to the method for generating a domain - adaptive object detection model without source domain based on class prototype alignment described above. As Figure 5 shown, the apparatus for generating a domain - adaptive object detection model without source domain based on class prototype alignment provided by the present invention includes: A feature extraction module 510, configured to input the target domain sample image into the first feature extraction module in the teacher model to obtain the first extracted feature, input the target domain sample image into the second feature extraction module in the student model to obtain the second extracted feature. The initial parameters of the teacher model and the student model are trained based on source domain data, and the source domain data includes source domain sample images and the corresponding object detection labels of the source domain sample images; The category prediction module 520 is configured to extract a first target feature and a second target feature from the first extracted feature and the second extracted feature respectively based on the target detection box, where the target detection box is obtained based on the result of performing target detection on the target domain sample image by the teacher model. The first target feature is input into the first target category prediction module of the teacher model to obtain a first target category prediction result, and the second target feature is input into the second target category prediction module of the student model to obtain a second target category prediction result; The class prototype update module 530 is configured to iteratively update the first class prototype of the preset target class based on each first target feature and the first target category prediction result of the same preset target class, and iteratively update the second class prototype of the preset target class based on each second target feature and the second target category prediction result corresponding to the same target preset class; The parameter update module 540 is configured to obtain the alignment loss between the first class prototype and the second class prototype, determine the training loss based on the alignment loss, update the parameters of the student model based on the training loss, update the parameters of the teacher model based on the parameters of the updated student model, and obtain the teacher model after multiple parameter updates as the target detection model. The alignment loss reflects the similarity between the first class prototype and the second class prototype and the class prototype knowledge distillation from the teacher model to the student model.
[0059] Figure 6 An example of the entity structure diagram of an electronic device is shown in Figure 6As shown in the figure, the electronic device may include: a processor 610, a communications interface 620, a memory 630, and a communication bus 640. Among them, the processor 610, the communications interface 620, and the memory 630 complete communication with each other through the communication bus 640. The processor 610 may call logic instructions in the memory 630 to execute a method for generating a passive domain adaptive object detection model based on class prototype alignment and / or an object detection method. The method for generating a passive domain adaptive object detection model based on class prototype alignment includes: inputting a target domain sample image into a first feature extraction module in a teacher model to obtain a first extracted feature, inputting the target domain sample image into a second feature extraction module of a student model to obtain a second extracted feature. The initial parameters of the teacher model and the student model are trained based on source domain data, and the source domain data includes source domain sample images and object detection labels corresponding to the source domain sample images; extracting a first target feature and a second target feature from the first extracted feature and the second extracted feature respectively based on object detection boxes, where the object detection boxes are obtained based on the object detection results of the teacher model for the target domain sample images, inputting the first target feature into a first target class prediction module of the teacher model to obtain a first target class prediction result, and inputting the second target feature into a second target class prediction module of the student model to obtain a second target class prediction result; iteratively updating a first class prototype of a preset target class based on each first target feature and the first target class prediction result of the same preset target class, and iteratively updating a second class prototype of the preset target class based on each second target feature and the second target class prediction result corresponding to the same target preset class; obtaining an alignment loss between the first class prototype and the second class prototype, determining a training loss based on the alignment loss, updating the parameters of the student model based on the training loss, updating the parameters of the teacher model based on the updated parameters of the student model, and obtaining the teacher model after multiple parameter updates as the object detection model. The alignment loss reflects the similarity between the first class prototype and the second class prototype and the class prototype knowledge distillation from the teacher model to the student model. The object detection method includes: obtaining an image to be detected, and inputting the image to be detected into the object detection model; obtaining an object detection result of the object detection model for the image to be detected, where the object detection model is the teacher model after training in the method for generating a passive domain adaptive object detection model based on class prototype alignment provided by the present invention.
[0060] In addition, when the logical instructions in the above-mentioned memory 630 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs.
[0061] On the other hand, the present invention also provides a computer program product, which includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the method for generating a source-free domain adaptation object detection model based on class prototype alignment and / or the object detection method provided by the above-mentioned various methods. The method for generating a source-free domain adaptation object detection model based on class prototype alignment includes: inputting a target domain sample image into a first feature extraction module in a teacher model to obtain a first extracted feature, and inputting the target domain sample image into a second feature extraction module of a student model to obtain a second extracted feature. The initial parameters of the teacher model and the student model are trained based on source domain data, and the source domain data includes source domain sample images and object detection labels corresponding to the source domain sample images; extracting a first target feature and a second target feature from the first extracted feature and the second extracted feature respectively based on object detection boxes, where the object detection boxes are obtained based on the object detection results of the teacher model for the target domain sample images; inputting the first target feature into a first target category prediction module of the teacher model to obtain a first target category prediction result, and inputting the second target feature into a second target category prediction module of the student model to obtain a second target category prediction result; iteratively updating a first class prototype of a preset target category based on each first target feature and the first target category prediction result of the same preset target category, and iteratively updating a second class prototype of the preset target category based on each second target feature and the second target category prediction result corresponding to the same target preset category; obtaining an alignment loss between the first class prototype and the second class prototype, determining a training loss based on the alignment loss, updating the parameters of the student model based on the training loss, updating the parameters of the teacher model based on the updated parameters of the student model, and obtaining the teacher model after multiple parameter updates as the object detection model. The alignment loss reflects the similarity between the first class prototype and the second class prototype and the class prototype knowledge distillation from the teacher model to the student model. The object detection method includes: obtaining an image to be detected and inputting the image to be detected into the object detection model; obtaining the object detection result of the object detection model for the image to be detected, where the object detection model is the teacher model after training in the method for generating a source-free domain adaptation object detection model based on class prototype alignment provided by the present invention.
[0062] In another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the method for generating a passive domain adaptive object detection model based on class prototype alignment and / or the object detection method provided by the above-mentioned various methods. The method for generating a passive domain adaptive object detection model based on class prototype alignment includes: inputting a target domain sample image into a first feature extraction module in a teacher model to obtain a first extracted feature, inputting the target domain sample image into a second feature extraction module of a student model to obtain a second extracted feature. The initial parameters of the teacher model and the student model are trained based on source domain data, and the source domain data includes source domain sample images and object detection labels corresponding to the source domain sample images; extracting a first target feature and a second target feature from the first extracted feature and the second extracted feature respectively based on object detection bounding boxes, where the object detection bounding boxes are obtained based on the object detection results of the teacher model for the target domain sample images, inputting the first target feature into a first target class prediction module of the teacher model to obtain a first target class prediction result, and inputting the second target feature into a second target class prediction module of the student model to obtain a second target class prediction result; iteratively updating a first class prototype of a preset target class based on each first target feature and the first target class prediction result of the same preset target class, and iteratively updating a second class prototype of the preset target class based on each second target feature and the second target class prediction result corresponding to the same target preset class; obtaining an alignment loss between the first class prototype and the second class prototype, determining a training loss based on the alignment loss, updating the parameters of the student model based on the training loss, updating the parameters of the teacher model based on the updated parameters of the student model, and obtaining the teacher model after multiple parameter updates as the object detection model. The alignment loss reflects the similarity between the first class prototype and the second class prototype and the class prototype knowledge distillation from the teacher model to the student model. The object detection method includes: obtaining an image to be detected, and inputting the image to be detected into the object detection model; obtaining the object detection result of the image to be detected output by the object detection model, where the object detection model is the teacher model after training in the method for generating a passive domain adaptive object detection model based on class prototype alignment provided by the present invention.
[0063] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.
[0064] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0065] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than limiting it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features therein; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A passive domain adaptive target detection model generation method based on class prototype alignment, characterized in that: include: Inputting a target domain sample image into a first feature extraction module in a teacher model to obtain a first extracted feature, inputting the target domain sample image into a second feature extraction module in a student model to obtain a second extracted feature, wherein initial parameters of the teacher model and the student model are obtained by training based on source domain data, wherein the source domain data includes a source domain sample image and a target detection label corresponding to the source domain sample image; Extracting a first target feature and a second target feature from the first extracted feature and the second extracted feature respectively based on a target detection frame, wherein the target detection frame is obtained based on a result of target detection performed by the teacher model on the target domain sample image, inputting the first target feature into a first target category prediction module of the teacher model to obtain a first target category prediction result, and inputting the second target feature into a second target category prediction module of the student model to obtain a second target category prediction result; Iterate the first type prototype of the preset target category based on each of the first target features and the first target category prediction result of the same preset target category, and iterate and update the second type prototype of the preset target category based on each of the second target features and the second target category prediction result corresponding to the same preset target category; Obtain an alignment loss between the first class prototype and the second class prototype, determine a training loss based on the alignment loss, update parameters of the student model based on the training loss, update parameters of the teacher model based on the updated parameters of the student model, and obtain the teacher model after multiple parameter updates as a target detection model, wherein the alignment loss reflects the similarity between the first class prototype and the second class prototype and the class prototype knowledge distillation from the teacher model to the student model.
2. The method for generating a passive domain adaptive target detection model based on class prototype alignment according to claim 1, characterized in that: Before the target domain sample image is input into the first feature extraction module in the teacher model, the method includes: Performing a first data enhancement process on the target domain sample image; Before the target domain sample image is input into the second feature extraction module in the student model, it includes: Performing a second data enhancement process on the target domain sample image; The data enhancement intensity of the second data enhancement processing is greater than that of the first data enhancement processing.
3. The method for generating a passive domain adaptive target detection model based on class prototype alignment according to claim 1, characterized in that: Before extracting the first target feature and the second target feature from the first extracted features and the second extracted features respectively based on the target detection frame, the method includes: The target detection frame is determined based on the intersection-over-union ratio between detection frames obtained by performing target detection on the target domain sample image output by the teacher model.
4. The method for generating a passive domain adaptive target detection model based on class prototype alignment according to claim 1, characterized in that: The iterating the first type prototype of the preset target category based on each of the first target features and the first target category prediction result of the same preset target category comprises: Normalizing the first target category prediction result corresponding to the preset target category in the current target domain sample image to obtain a normalized first target category prediction result, and determining a current first category feature based on the normalized first target category prediction result and the first target feature corresponding to the preset target category in the current target domain sample image; Iteratively updating the current first-category prototype based on the first-category features; The initial first-category prototype is the first-category feature corresponding to the target and sample image first input to the teacher model; The iterative updating of the second type prototype of the preset target category based on each of the second target features corresponding to the same preset target category and the second target category prediction result comprises: Normalizing the second target category prediction result corresponding to the preset target category in the current target domain sample image to obtain a normalized second target category prediction result, and determining a current second category feature based on the normalized second target category prediction result and the second target feature corresponding to the preset target category in the current target domain sample image; Iteratively updating the current first-category prototype based on the second-category features; The initial second-category prototype is the second-category feature corresponding to the target and sample image first input into the teacher model.
5. The method for generating a passive domain adaptive target detection model based on class prototype alignment according to claim 1, characterized in that: The obtaining of the alignment loss between the first type of prototype and the second type of prototype includes: Determining a prototype contrast loss based on the similarity between the first type of prototype and the second type of prototype; Inputting the first type of prototypes into the teacher model and the student model respectively to obtain first type of prediction results and second type of prediction results, inputting the second type of prototypes into the teacher model and the student model respectively to obtain third type of prediction results and fourth type of prediction results; Determine a prototype distillation loss based on distribution differences between any two of the first category prediction results, the second category prediction results, the third category prediction results, and the fourth category prediction results; The alignment loss is determined based on the prototype contrast loss and the prototype distillation loss.
6. A target detection method based on the passive domain adaptive target detection model generation method based on class prototype alignment according to any one of claims 1 to 5, characterized in that: include: Acquire an image to be detected, and input the image to be detected into the target detection model; Obtain a target detection result output by the target detection model for the image to be detected.
7. A passive domain adaptive target detection model generation device based on class prototype alignment, characterized in that: The device comprises: A feature extraction module, used to input a target domain sample image into a first feature extraction module in a teacher model to obtain a first extracted feature, and input the target domain sample image into a second feature extraction module in a student model to obtain a second extracted feature, wherein the initial parameters of the teacher model and the student model are obtained by training based on source domain data, and the source domain data includes a source domain sample image and a target detection label corresponding to the source domain sample image; a category prediction module, for extracting a first target feature and a second target feature from the first extracted feature and the second extracted feature respectively based on a target detection frame, wherein the target detection frame is obtained based on a result of target detection performed by the teacher model on the target domain sample image, the first target feature is input into a first target category prediction module of the teacher model to obtain a first target category prediction result, and the second target feature is input into a second target category prediction module of the student model to obtain a second target category prediction result; A class prototype updating module, configured to iterate the first class prototype of the preset target category based on each of the first target features and the first target category prediction result of the same preset target category, and iteratively update the second class prototype of the preset target category based on each of the second target features and the second target category prediction result corresponding to the same preset target category; A parameter updating module is used to obtain the alignment loss between the first type of prototype and the second type of prototype, determine the training loss based on the alignment loss, update the parameters of the student model based on the training loss, update the parameters of the teacher model based on the updated parameters of the student model, and obtain the teacher model after multiple parameter updates as the target detection model, wherein the alignment loss reflects the similarity between the first type of prototype and the second type of prototype and the class prototype knowledge distillation from the teacher model to the student model.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the method for generating a passive domain adaptive target detection model based on class prototype alignment as described in any one of claims 1 to 5 and / or the target detection method as described in claim 6 is implemented.
9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for generating a passive domain adaptive target detection model based on class prototype alignment as described in any one of claims 1 to 5 and / or the target detection method as described in claim 6 is implemented.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the method for generating a passive domain adaptive target detection model based on class prototype alignment as described in any one of claims 1 to 5 and / or the target detection method as described in claim 6 is implemented.
Citation Information
Patent Citations
Passive field adaptive target detection method
CN112861616A
Passive domain adaptive target detection method and device
CN117636086A
Target detection model training method and device, target detection method and device, equipment and medium
CN118155011A
Remote sensing image unsupervised domain adaptation method based on comparative learning and multi-prototype alignment
CN119251646A
Systems and methods for training machine learning model based on cross-domain data
US20220198339A1
Cited By
Passenger abnormal behavior recognition method and device in elevator monitoring night vision mode
CN120452068A
Method and device for identifying abnormal passenger behavior in elevator monitoring night vision mode
CN120452068B
Domain adaptive target detection method and system based on adversarial double-teacher knowledge distillation
CN121437856A
A Domain-Adaptive Target Detection Method and System Based on Adversarial Dual-Teacher Knowledge Distillation
CN121437856B
Small sample defect detection method based on class awareness prototype migration
CN121616602A