Image-processing device, image-processing method, and image-processing program

The image processing apparatus addresses domain shift issues in DETR by generating pseudo-labels and optimizing the model during inference, enhancing object detection performance and practicality.

WO2025158679A1PCT designated stage Publication Date: 2025-07-31NT T INC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2024/002533
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-01-26
Publication Date
2025-07-31

AI Technical Summary

Technical Problem

Conventional image processing models, such as CNN, are vulnerable to domain shift, leading to significant performance deterioration, and existing test-time adaptation methods are not applicable to advanced object detection models like DETR, limiting their practical application.

Method used

An image processing apparatus that includes a pseudo-label generation unit to create pseudo-labels with a confidence threshold and an adaptation unit to optimize the DETR model using target domain data and pseudo-labels, enabling domain adaptation during inference without prior data collection.

Benefits of technology

The proposed method enhances the ability of DETR to adapt to domain shift, improving object detection performance and expanding the applicability of high-performance models in various environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2024002533_31072025_PF_FP_ABST
    Figure JP2024002533_31072025_PF_FP_ABST
Patent Text Reader

Abstract

An image-processing device (10) has a pseudo label generation unit (101) and an adaptation unit (102). The pseudo label generation unit (101) generates, as a pseudo label, an object detection result obtained by inputting target domain data to a teacher model generated on the basis of a student model, which is an object detection model trained using source domain data, the object detection result having a confidence level equal to or greater than a threshold value. The adaptation unit (102) uses the target domain data and the pseudo label to optimize the student model.
Need to check novelty before this filing date? Find Prior Art

Description

Image processing device, image processing method, and image processing program

[0001] The present invention relates to an image processing device, an image processing method, and an image processing program.

[0002] Conventionally, convolutional neural networks (CNNs) have been known as machine learning techniques for identifying, detecting, and segmenting objects in images, etc. For example, CNNs contribute to the automation of visual inspection processes in business.

[0003] When promoting the automation of visual inspection processes in business by processing captured images, it is desirable for image processing models such as CNN to be in line with the intuition of humans who are originally performing the visual inspection processes.

[0004] On the other hand, image processing models such as CNNs are vulnerable to changes known as domain shifts, and it is known that the occurrence of domain shifts significantly degrades processing accuracy. CNNs' vulnerability to domain shifts implies that image processing models must be used in the same environment during training and inference, which can pose a major practical challenge. Therefore, a technique known as domain adaptation is being actively researched to enable CNNs to perform appropriate image processing even when domain shifts occur.

[0005] In domain adaptation in image processing using CNNs, a methodology called unsupervised domain adaptation is most commonly adopted. Unsupervised domain adaptation assumes the existence of supervised data (source domain data) in an environment where no domain shift has occurred and unsupervised data (target domain data) in which a domain shift has occurred during CNN training, and aims to adapt the CNN model to the domain shift using the unsupervised target domain data.

[0006] Unsupervised domain adaptation is currently being researched as a promising method to combat domain shift. However, unsupervised domain adaptation has the limitation that it requires prior knowledge of the environment at the time of inference, since it assumes the existence of target domain data. Since CNNs are used for a variety of applications and environments, it is not realistic to know all of the environments at the time of inference in advance.

[0007] To alleviate this constraint, a methodology called test-time adaptation (see, for example, Non-Patent Document 1 and Non-Patent Document 2) has recently attracted attention. In test-time adaptation, domain adaptation is performed at the time of inference using inference data acquired in the environment at the time of inference. This eliminates the need to collect target domain data before inference, and enables domain adaptation in any environment. For example, Non-Patent Document 3 describes a test-time adaptation method for object detection tasks.

[0008] Y. Sun et al., “Test-Time Training with Self-Supervision for Generalization under Distribution Shifts”, 2020. S. Goyal et al., “Test-Time Adaptation via Conjugate Pseudo-labels”, 2022. V. VS et al., “Towards Online Domain Adaptive Object Detection”, 2022. S. Carion et al., “End-to-End Object Detection with Transformers”, 2020.

[0009] However, conventional techniques may not be able to adapt high-performance object detection models to domain shifts.

[0010] Conventional test-time adaptation methods are often targeted at image classification tasks, and may not be applicable to tasks such as object detection, which are important in practical applications.

[0011] Non-Patent Document 3 describes test-time adaptation specialized for a long-standing object detection model called Faster RCNN. However, the test-time adaptation described in Non-Patent Document 3 cannot be applied to DETR (see, for example, Non-Patent Document 4), which is a cutting-edge object detection model.

[0012] DETR is known to have high object detection performance. Many successor models to DETR have been proposed. Therefore, if test-time adaptation can be applied to DETR, it is expected that the range of models capable of high-accuracy object detection even when domain shift occurs will expand, thereby enhancing the practicality of deep learning models.

[0013] In order to solve the above-mentioned problems and achieve the objectives, the image processing device is characterized by having a pseudo label generation unit that generates pseudo labels from object detection results obtained by inputting target domain data into a second model generated based on a first model, which is an object detection model that has been trained using source domain data, for those results that have a confidence level equal to or greater than a threshold, and an adaptation unit that optimizes the first model using the target domain data and the pseudo labels.

[0014] According to the present invention, a high performance object detection model can be adapted to domain shifts.

[0015] FIG. 1 is a diagram illustrating an example of the configuration of an image processing apparatus according to a first embodiment. FIG. 2 is a flowchart illustrating a processing flow of the image processing apparatus according to the first embodiment. FIG. 3 is a flowchart illustrating a processing flow of a pseudo label generation unit. FIG. 4 is a diagram illustrating generation of pseudo labels. FIG. 5 is a diagram illustrating the influence of a threshold on pseudo labels. FIG. 6 is a diagram illustrating learning for low label quality. FIG. 7 is a flowchart illustrating a processing flow of an adaptation unit. FIG. 8 is a flowchart illustrating a processing flow of an inference unit. FIG. 9 is a flowchart illustrating a processing flow of a merging unit. FIG. 10 is a diagram illustrating experimental results. FIG. 11 is a diagram illustrating experimental results. FIG. 12 is a diagram illustrating an example of a computer that executes an image processing program.

[0016] DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS An image processing apparatus, an image processing method, and an image processing program according to the present invention will be described in detail below with reference to the accompanying drawings. However, the present invention is not limited to the following embodiments.

[0017] Here, DETR described in Non-Patent Document 4 is an object detection model that combines CNN and Transformer (Reference Document 1). As mentioned above, the performance of DETR significantly deteriorates when a domain shift occurs. A domain shift is a phenomenon in which the distribution of data during training (source domain) changes from the distribution of data during testing (operation) (target domain).

[0018] Reference 1: A. Vaswani, et al.: “Attention is All you Need”, NIPS, 2017

[0019] Unsupervised domain adaptation is a well-known method for maintaining model performance even during domain shifts (References 2 and 3). Unsupervised domain adaptation uses both labeled source domain data and unlabeled target domain data to achieve domain adaptation. However, unsupervised domain adaptation has the drawback of requiring the collection of target domain data in advance.

[0020] Reference 2: W. Wang, et al.: “Exploring Sequence Feature Alignment for Domain Adaptive Detection Transformers”, ACMMM, 2021.

[0021] Reference 3: K. Gong, et al.: “Improving Transferability for Domain Adaptive Detection Transformers”, ACMMM, 2022.

[0022] Test-Time Adaptation (TTA) is a known method that can solve the problems of unsupervised domain adaptation (see Non-Patent Document 1, Non-Patent Document 2, Reference 4). TTA is a new domain adaptation paradigm that uses only test target data obtained during testing, and does not require prior collection of target domain data.

[0023] Reference 4: D. Wang, et al.: “Tent: Fully Test-Time Adaptation by Entropy Minimization”, ICLR, 2021.

[0024] TTA is primarily specialized for image classification (class prediction) tasks or object detection tasks using older object detection models such as Faster RCNN (Non-Patent Document 3), and cannot be applied to modern object detection tasks such as DETR. Note that an image classification task is a task of identifying the type of object appearing in an image. Also, an object detection class is a task of detecting the position and type of an object appearing in an image.

[0025] [First Embodiment] Therefore, in this embodiment, a TTA applicable to DETR is proposed. First, the configuration of an image processing apparatus according to the first embodiment will be described with reference to Fig. 1. Fig. 1 is a diagram showing an example of the configuration of an image processing apparatus according to the first embodiment.

[0026] As shown in FIG. 1, the image processing device 10 includes a pseudo label generation unit 101, an adaptation unit 102, a merging unit 103, an inference unit 104, an inference image storage unit 111, and an object detection parameter storage unit 112.

[0027] The image processing device 10 performs object detection on an input inference image using a trained object detection model. The input inference image is target domain data, and a domain shift may have occurred with respect to the source domain data used to train the object detection model. The object detection model in this embodiment is DETR.

[0028] The inference image storage unit 111 stores the inference image. The object detection parameter storage unit 112 stores the DETR parameters. For example, the DETR parameters are weights and biases of a neural network.

[0029] The flow of processing by the image processing device 10 will be described with reference to Fig. 2. Fig. 2 is a flowchart showing the flow of processing by the image processing device according to the first embodiment.

[0030] As shown in FIG. 2, first, the image processing device 10 sends parameters for a desired object detection task stored in the object detection parameter storage unit 112 to the adaptation unit 102, the merging unit 103, and the pseudo label generation unit 101 (step S11).

[0031] Next, the image processing device 10 sends the inference image acquired from the inference image storage unit 111 to the adaptation unit 102, the pseudo label generation unit 101, and the inference unit 104 (step S12).

[0032] Here, the image processing device 10 generates pseudo teacher labels from the inference image in the pseudo label generation unit 101 and sends them to the adaptation unit 102 (step S13).

[0033] Next, the image processing device 10 executes a process of adapting the parameters of the object detection model using the inference image and pseudo labels in the adaptation unit 102, and sends the adapted parameters to the merging unit 103 and the inference unit 104 (step S14).

[0034] Then, the image processing device 10 performs object detection on the inference image using the adapted parameters in the inference unit 104, and outputs the result as an inference result (step S15).

[0035] Furthermore, the image processing device 10 uses the adapted parameters in the merge unit 103 to calculate new parameters to be used in the pseudo label generation unit 101, and sends the new parameters to the pseudo label generation unit 101 as updated teaching parameters (step S16).

[0036] If the inference has ended (step S17, Yes), the image processing device 10 ends the process. If the inference has not ended (step S17, No), the image processing device 10 returns to step S12 and repeats the process.

[0037] The image processing device 10 determines that the inference is complete when, for example, there are no unprocessed inference images. After executing step S16 once, the image processing device 10 does not perform step S11, but performs adaptation (step S14) using the new parameters calculated in step S16.

[0038] The processing flow of the pseudo label generation unit 101 will be described with reference to Fig. 3. Fig. 3 is a flowchart showing the processing flow of the pseudo label generation unit.

[0039] First, as shown in Fig. 3, when inference has started (step S101, Yes), the pseudo label generation unit 101 receives object detection parameters from the object detection parameter storage unit 112 (step S102). The inference start is the state before new parameters are calculated in step S16 in Fig. 2. In other words, when the repetition of steps S12 to S16 in Fig. 2 enters the second cycle, the inference start state is no longer reached. Note that the object detection parameters are parameters of DETR that have been trained using source domain data.

[0040] If inference has not started (No at step S101), the pseudo label generation unit 101 receives updated teaching parameters from the merge unit 103 (step S103).

[0041] Next, the pseudo label generation unit 101 receives the inferred image (step S104). The pseudo label generation unit 101 applies weak data augmentation to the inferred image (step S105). The weak data augmentation is, for example, horizontally flipping the image. Note that the inferred image is target domain data.

[0042] The pseudo label generation unit 101 performs inference using the received parameters on the data-augmented inference image (step S106). The pseudo label generation unit 101 performs threshold processing on the inference result and sends it to the adaptation unit 102 (step S107). That is, the pseudo label generation unit 101 generates recall-focused pseudo labels and sends them to the adaptation unit 102.

[0043] The pseudo label generation unit 101 also updates the parameters through self-supervised learning based on the inference results (step S108). Step S108 improves the performance of the teacher model. However, this process may be omitted.

[0044] In this way, the pseudo label generation unit 101 generates pseudo labels from the object detection results obtained by inputting target domain data into a teacher model generated based on a student model, which is an object detection model that has been trained using source domain data, and those object detection results whose confidence level is above a threshold.

[0045] The following describes in detail the processing of the pseudo label generation unit 101. Fig. 4 is a diagram for explaining the generation of pseudo labels.

[0046] 4 is an inference image. Usually, an inference image does not have a label. The pseudo label generator 101 generates pseudo labels for test adaptation.

[0047] The student model is DETR constructed from the object detection parameters received by the pseudo label generation unit 101 from the object detection parameter storage unit 112 in step S102. The teacher model is DETR constructed from the teacher parameters received by the pseudo label generation unit 101 in step S103. The teacher parameters are updated teacher parameters calculated by the merging unit 103. If the updated teacher parameters have not been calculated by the merging unit 103, the teacher parameters may be the same as the object detection parameters.

[0048] The pseudo label generation unit 101 inputs test data into the teacher model and causes it to perform object detection. The pseudo label generation unit 101 determines, among the detection results (position and object type), those with a certainty level equal to or greater than a threshold as recall-focused pseudo labels.

[0049] The pseudo label generating unit 101 sets a threshold value for the confidence level that is smaller (for example, 0.2 to 0.5) than the threshold value (for example, 0.7) in the conventional self-training methods described in References 5 and 6.

[0050] Reference 5: K. Sohn, et al.: “FixMatch: Simplifying Semi-Supervised Learning with Consistency and Confidence”, NeruIPS, 2020.

[0051] Reference 6: P. Wang, et al.: “Omni-DETR: Omni-Supervised Object Detection with Transformers”, CVPR, 2022.

[0052] Because Transformer-based DETR has data-hungry characteristics, if pseudo-labeling is insufficient, overfitting can cause a significant performance degradation during TTA. In contrast, the pseudo-label generation unit 101 can obtain many pseudo-labels (recall-oriented pseudo-labels) by using a smaller threshold than in conventional techniques, thereby avoiding overfitting.

[0053] Figure 5 illustrates the effect of the threshold on pseudo labels. When the threshold is large, the quality of the pseudo labels obtained is high, but the number is small, so performing TTA induces overfitting of DETR. On the other hand, when the threshold is small, the quality of the labels is low, but the number is large (high recall), so overfitting of DETR is avoided.

[0054] 6 is a diagram illustrating low-label quality training. Even when low-quality recall-focused pseudo labels are used in TTA, appropriate training is performed by low-label quality training. The low-label quality training is performed by the adaptation unit 102.

[0055] Confidence Refinement corresponds to step S108. At this time, a combination of test data or extended test data and pseudo labels is used as training data for updating (learning) the parameters of the training model.

[0056] Confidence refinement is a process for improving the quality of pseudo labels obtained from a teacher by minimizing the entropy of the teacher model output. This is based on the fact that there is a positive correlation between entropy and class prediction performance (Reference 3). In confidence refinement, the pseudo label generation unit 101 updates the parameters of the teacher model as shown in equation (1).

[0057]

[0058] Θ teach (i) is a parameter of the teacher model. γ is a predetermined learning rate. L H is the entropy. i is the i-th test data.

[0059] This improves the predictive performance of the teacher model, and the next time test data is input, the quality of the recall-focused pseudo-labels obtained from the teacher model will be higher.

[0060] The processing flow of the adaptation unit 102 will be described with reference to Fig. 7. Fig. 7 is a flowchart showing the processing flow of the adaptation unit.

[0061] 7, first, when inference has started (step S201, Yes), the adaptation unit 102 receives parameters for object detection (step S202). When inference has not started (step S201, No), the adaptation unit 102 receives an inference image (step S203).

[0062] The adaptation unit 102 acquires pseudo labels (recall-weighted pseudo labels) from the pseudo label generation unit 101 (step S204). The adaptation unit 102 applies strong data augmentation to the inference image (step S205). Strong data augmentation is, for example, adding noise and blur. Alternatively, strong data augmentation may be the same process as weak data augmentation.

[0063] The adaptation unit 102 applies object detection using the received parameters to the inference image (step S206). The adaptation unit 102 calculates the loss with respect to the pseudo label only for the object detection results with high confidence, and updates the parameters (step S207). The adaptation unit 102 outputs the updated parameters as adapted parameters (step S208).

[0064] The adaptation unit 102 updates the parameters of the student model using unconfident filtering. At this time, a combination of target domain data (test data or augmented test data) and pseudo labels is used as training data for updating (learning) the parameters of the student model.

[0065] Unconfident filtering skips certain loss calculations when the confidence value in the output of the student model is low. In other words, the adaptation unit 102 reflects only detection results with a confidence value equal to or greater than a threshold in the calculation of some losses. This reduces the impact of low-quality pseudo-labels on updating the parameters of the student model. This is based on the hypothesis that the lower the confidence, the more susceptible the model is to the influence of low-quality labels.

[0066] Loss L using Unconfident Filtering UF The calculation is expressed as in equation (2).

[0067]

[0068] N is the number of outputs (combinations of classes and box positions) of DETR. n s is the confidence of each output. If the confidence is not in the range from T1 to T2, the loss L n cl is calculated. N / C is a scaling factor. L n box is a loss calculated independently of confidence.

[0069] The adaptation unit 102 calculates the loss L UF Update the parameters of the student model so that is small.

[0070] In this way, the adaptation unit 102 optimizes the student model using a loss that takes a positive value when the confidence level of the object detection results obtained by inputting target domain data into the student model is within a specified range, and takes a value of 0 when the confidence level is not within the specified range.

[0071] The flow of processing by the inference unit 104 will be described with reference to Fig. 8. Fig. 8 is a flowchart showing the flow of processing by the inference unit.

[0072] 8, first, the inference unit 104 acquires the adapted parameters from the adaptation unit 102 (step S301). The adapted parameters are parameters of the updated student model.

[0073] The inference unit 104 receives the inference image (step S302). Then, the inference unit 104 performs object detection on the inference image using the adapted parameters and outputs the result as an inference result (step S303). That is, the inference unit 104 performs inference using the DETR on which TTA has been performed.

[0074] The flow of processing by the merge unit 103 will be described with reference to Fig. 9. Fig. 9 is a flowchart showing the flow of processing by the merge unit.

[0075] 9, first, when inference has started (step S401, Yes), the merging unit 103 receives object detection parameters (step S402). When inference has not started (step S401, No), the merging unit 103 acquires adapted parameters from the adaptation unit 102 (step S403).

[0076] The merging unit 103 merges the object detection parameters and the adapted parameters and outputs the merged parameters as updated teacher parameters (step S404).

[0077] For example, the merging unit 103 performs merging using an exponential moving average that linearly combines the object detection parameters and the adapted parameters. The merging unit 103 may perform merging using a method that does not lose information about the object detection parameters.

[0078] In this way, the merging unit 103 generates a teacher model by linearly combining (for example, moving average) the student model before being optimized by the adaptation unit 102 and the student model after being optimized by the adaptation unit 102. For example, the value P 1 However, after adaptation, the parameters are 2 If the parameter has been updated to αP 1 +(1-α)P 2 It is calculated as follows.

[0079] [Experiment] An experiment conducted to confirm the effects of the embodiment will be described. The processing flow in the experiment is the same as that in the first embodiment. The evaluation was based on the object detection accuracy when test data was input again to DETR (student model) constructed based on the adapted parameters output by the adaptation unit 102.

[0080] The datasets used in the experiment are as follows: Source domain data: Cityscapes (Reference 7) Target domain data: FoggyCityscapes (artificially fogged based on depth information) (Reference 8)

[0081] Reference 7: M. Cordts, et al.: “The Cityscapes Dataset for Semantic Urban Scene Understanding”, CVPR, 2016.

[0082] Reference 8: C. Sakaridis, et al.: “Semantic Foggy Scene Understanding with Synthetic Data”, IJCV, 2018.

[0083] The models used in the experiments were Deform-DETR (Reference 9), which is an improved version of DETR, and DN-DETR (Reference 10).

[0084] Reference 9: X. Zhu, et al.: “Deformable DETR: Deformable Transformers for End-to-End Object Detection”, ICLR, 2021.

[0085] Reference 10: F. Li, et al.: “DN-DETR: Accelerate detr training by introducing query denoising”, CVPR, 2022.

[0086] The optimizer used in the experiment was Adam, and the learning rate was 1.0 × 10 -6 and the weight decay is 0.

[0087] Fig. 10 shows the experimental results, which show the object detection performance (mAP) when varying the confidence threshold for generating recall-weighted pseudo labels (the larger the value, the better the performance).

[0088] As shown in Figure 11, when the threshold is set to a small value as in the first embodiment, performance is improved compared to when the threshold is set to 0.6 or 0.7. This suggests that overlearning is suppressed by the first embodiment, improving model performance. Note that 0.6 and 0.7 are thresholds used in the conventional self-training methods described in References 5 and 6.

[0089] Fig. 11 shows the experimental results. Fig. 11 shows the comparison results of the performance when using source domain data (Source), the performance when using the conventional technology TENT (Reference 4), and the performance when using the first embodiment. "w / o" means that training for low label quality was not performed.

[0090] FIG. 11 shows that the performance of the model according to the first embodiment is improved compared to the prior art, and that the performance is further improved by introducing training for low label quality.

[0091] [System Configuration, etc.] The components of each device shown in the figure are conceptual functional units and do not necessarily have to be physically configured as shown. In other words, the specific form of distribution and integration of each device is not limited to that shown in the figure, and all or part of the devices can be functionally or physically distributed or integrated in any unit depending on various loads, usage conditions, etc. Furthermore, all or any part of the processing functions performed by each device can be realized by a CPU (Central Processing Unit) and a program analyzed and executed by the CPU, or can be realized as hardware using wired logic. The program may be executed not only by the CPU but also by other processors such as a GPU.

[0092] Furthermore, among the processes described in this embodiment, all or part of the processes described as being performed automatically can be performed manually, or all or part of the processes described as being performed manually can be performed automatically using a known method.In addition, the information including the processing procedures, control procedures, specific names, various data and parameters shown in the above documents and drawings can be changed as desired unless otherwise specified.

[0093] [Program] In one embodiment, the image processing device 10 can be implemented by installing an image processing program that executes the above-described processes as package software or online software on a desired computer. For example, by executing the image processing program on an information processing device, the information processing device can function as the image processing device 10. The information processing device referred to here includes desktop and notebook personal computers. Other information processing devices also include mobile communication terminals such as smartphones, mobile phones, and PHS (Personal Handyphone Systems), as well as slate terminals such as PDAs (Personal Digital Assistants).

[0094] The image processing device 10 may also be implemented as a server device that provides services related to the above processing to a client terminal device used by a user. For example, the server device may be implemented as a server device that provides a service that outputs an object detection result. In this case, the server device may be implemented as a web server or as a cloud that provides services related to the above processing through outsourcing.

[0095] 12 is a diagram showing an example of a computer that executes an image processing program. The computer 1000 includes, for example, a memory 1010 and a CPU 1020. The computer 1000 also includes a hard disk drive interface 1030, a disk drive interface 1040, a serial port interface 1050, a video adapter 1060, and a network interface 1070. These components are connected by a bus 1080.

[0096] The memory 1010 includes a read-only memory (ROM) 1011 and a random access memory (RAM) 1012. The ROM 1011 stores a boot program such as a basic input / output system (BIOS). The hard disk drive interface 1030 is connected to a hard disk drive 1090. The disk drive interface 1040 is connected to a disk drive 1100. A removable storage medium such as a magnetic disk or optical disk is inserted into the disk drive 1100. The serial port interface 1050 is connected to a mouse 1110 and a keyboard 1120, for example. The video adapter 1060 is connected to a display 1130, for example.

[0097] The hard disk drive 1090 stores, for example, an OS 1091, an application program 1092, a program module 1093, and program data 1094. That is, the programs that define each process of the image processing device 10 are implemented as program modules 1093 in which computer-executable code is written. The program modules 1093 are stored, for example, in the hard disk drive 1090. For example, the program modules 1093 for executing processes similar to those of the functional configuration of the image processing device 10 are stored in the hard disk drive 1090. Note that the hard disk drive 1090 may be replaced with an SSD (Solid State Drive).

[0098] Furthermore, setting data used in the processing of the above-described embodiment is stored as program data 1094, for example, in the memory 1010 or the hard disk drive 1090. Then, the CPU 1020 reads the program module 1093 or the program data 1094 stored in the memory 1010 or the hard disk drive 1090 into the RAM 1012 as necessary, and executes the processing of the above-described embodiment.

[0099] The program module 1093 and program data 1094 may not necessarily be stored in the hard disk drive 1090, but may also be stored in a removable storage medium and read by the CPU 1020 via the disk drive 1100 or the like. Alternatively, the program module 1093 and program data 1094 may be stored in another computer connected via a network (such as a local area network (LAN) or a wide area network (WAN)). The program module 1093 and program data 1094 may then be read by the CPU 1020 from the other computer via the network interface 1070.

[0100] The following additional notes are provided regarding the above-described embodiments.

[0101] (Supplementary Item 1) An image processing device including: a memory; and at least one processor connected to the memory, wherein the processor generates, as pseudo labels, object detection results having a confidence level equal to or greater than a threshold, obtained by inputting target domain data into a second model generated based on a first model, which is an object detection model trained using source domain data; and optimizes the first model using the target domain data and the pseudo labels.

[0102] (Supplementary Item 2) A non-transitory storage medium storing a program executable by a computer to perform image processing, wherein the image processing comprises: inputting target domain data into a second model generated based on a first model, which is an object detection model trained using source domain data, and generating pseudo labels from among the object detection results obtained, those with a confidence level equal to or greater than a threshold; and optimizing the first model using the target domain data and the pseudo labels.

[0103] REFERENCE SIGNS LIST 10 Image processing device 101 Pseudo label generation unit 102 Adaptation unit 103 Merging unit 104 Inference unit 111 Inference image storage unit 112 Object detection parameter storage unit

Claims

1. A pseudo-label generation unit that generates, as pseudo-labels, those with a confidence level equal to or higher than a threshold value among object detection results obtained by inputting target domain data to a second model generated based on a first model, which is an object detection model trained using source domain data; and an adaptation unit that optimizes the first model using the target domain data and the pseudo-labels. An image processing apparatus characterized by comprising the above.

2. The image processing apparatus according to claim 1, further comprising a merge unit that generates the second model by linearly combining the first model before being optimized by the adaptation unit and the first model after being optimized by the adaptation unit.

3. The adaptation unit optimizes the first model using a loss that takes a positive value when the confidence level is within a predetermined range among object detection results obtained by inputting the target domain data to the first model, and takes 0 when the confidence level is not within the predetermined range. The image processing apparatus according to claim 1, characterized by the above.

4. An image processing method executed by an image processing apparatus, the method including: a pseudo-label generation step of generating, as pseudo-labels, those with a confidence level equal to or higher than a threshold value among object detection results obtained by inputting target domain data to a second model generated based on a first model, which is an object detection model trained using source domain data; and an adaptation step of optimizing the first model using the target domain data and the pseudo-labels. An image processing method characterized by including the above.

5. A pseudo-label generation step of generating, as pseudo-labels, those with a confidence level equal to or higher than a threshold value among object detection results obtained by inputting target domain data to a second model generated based on a first model, which is an object detection model trained using source domain data; and an adaptation step of optimizing the first model using the target domain data and the pseudo-labels. An image processing program characterized by causing a computer to execute the above.

Citation Information

Patent Citations

  • Pedestrian detection method and system based on unsupervised domain adaptation and storage medium

    CN112052818A