Underwater domain adaptation-oriented style perception target detection method

By constructing the image pair attention guidance module and the two-stage teacher-student area proposed alignment method, the problem of limited model detection accuracy in underwater environments is solved, and efficient object detection in different underwater environments is achieved.

CN120564019APending Publication Date: 2025-08-29JIANGSU OCEAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510531851.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-25
Publication Date
2025-08-29

Smart Images

  • Figure CN120564019A_ABST
    Figure CN120564019A_ABST
Patent Text Reader

Abstract

The invention discloses a style perception target detection method for underwater domain adaptation, and the method comprises the steps: constructing an image pair attention guidance (IAG) module based on wavelet transformation, enabling the object content of a source domain and the style of a target domain to form an image pair, taking the image pair as a group of complementary information, and carrying out the high-order feature modeling, a cross-attention module and discrete wavelet transform are combined, differential content perception is performed on four sub-bands after wavelet decomposition, high-order domain invariant features are extracted, and a student model is guided to learn cross-domain invariant information from an image pair. Furthermore, a two-stage teacher-student regional proposal alignment (TTRPA) strategy is provided, an instance-level feature alignment guide model is utilized to allocate higher attention weights for more effective regions, and consistency loss is constructed to realize sub-optimal supervision, so that the generalization ability of the model is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of image processing technology, and in particular relates to a style-aware target detection method for underwater domain adaptation. Background Art

[0002] In underwater environments, different water quality types serving as backgrounds share certain similarities, and cross-domain targets also share a significant degree of similarity. This provides a fundamental cross-domain transferability foundation for domain adaptation in diverse water object detection tasks. Due to the significant domain differences between the source and target domains, and in addition to the effects of domain shift, images captured in underwater environments are subject to multiple factors, including turbid water, varying illumination, underwater refraction, and scattering. Distinguishing background and target features in complex environments is difficult, making the teacher model prone to missed or false detections. Under the Mean Teacher framework, the model's prediction accuracy is limited by the quality of the pseudo-labels. Given the characteristics of underwater environments, the quality of pseudo-labels is primarily affected by two factors. First, the model may inaccurately predict certain samples; second, the model may learn errors in the pseudo-labels, which accumulate in subsequent iterations, leading to performance degradation.

[0003] Because data augmentation increases the size of the low-probability subset neighborhood of the generalized target domain data, resulting in a larger expansion factor and thus higher accuracy bounds, using more aggressive data augmentation to regularize input stability can significantly improve representation quality. The source domain image provides clear target features, while the target domain style image provides features and interference information in the target environment. The combination of the two can more comprehensively describe the target, facilitating the detection and recognition of difficult examples. Combining the Mean Teacher method with feature alignment methods helps align the data distribution of the target and source domains, enhancing the model's generalization performance. Summary of the Invention

[0004] This paper proposes a style-aware object detection method for underwater domain adaptation. It constructs a wavelet-based Image Pair Attention Guidance (IAG) module and proposes a two-stage Teacher-Student Region Proposal Alignment (TTRPA) method. The attention module perceives the stylistic features of the target domain, enabling the model to simultaneously focus on the feature representations before and after stylization. Instance-level feature alignment guides the model to assign higher attention weights to more effective regions, and a consistency loss is constructed to achieve suboptimal supervision.

[0005] In order to achieve the above object, the technical solution adopted by the present invention is as follows:

[0006] A style-aware object detection method for underwater domain adaptation is characterized by:

[0007] S1: Pre-training the object detection model using data and annotations from the source domain

[0008] Using the source domain image I S ∈R H×W×3 As the content image, the target domain image I T ∈R H×W×3 As a style image, each content image I S Combined with the randomly selected style image, the style transfer model is used to generate an output image I with the target domain style offline. S_Tstyle ∈R H×W×3 , and with I S Constitute image pair I P =(I S ,I S_Tstyle ), the two images in an image pair share the same annotation;

[0009] S2: Image pair I P =(I S ,I S_Tstyle ) Feature extraction is performed through ResNet50, and then the IAG module is used to enable the model to focus on the feature representation before and after stylization;

[0010] S3: The extracted features are regressed and classified through the Encoder layer and the Decoder layer to obtain the classification loss of the model. The classification loss is defined as:

[0011]

[0012] Among them, L S box is the bounding box regression loss of the model on the source domain, L S giou is the GIoU loss of the model, L S cls is the classification loss of the model, x s and y s are the prediction results and true values ​​of the model in the source domain respectively;

[0013] S4: The obtained model is used as the student model, and a model with the same network structure as the teacher model is defined. The weight of the teacher model is updated from the weight of the student network through the exponential moving average EMA, and the model is trained through supervised learning using classification loss.

[0014] S5: The teacher model and the student model are aligned at the instance level for the data in the source and target domains through the TTRPA mechanism;

[0015] S6: Construct the total target loss function and train the network to obtain the final target detection result.

[0016] Furthermore, the IAG module in S2 is an image-to-attention guidance module based on wavelet transform, and its construction method is as follows:

[0017] S2.1: Constructing hue-based style transfer image pairs I P =(I S ,I S_Tstyle );

[0018] S2.2: Apply wavelet downsampling to high-order feature maps, I P =(I S ,I S_Tstyle ) After Resnet50 feature extraction, three feature maps of different scales are obtained Using Haar wavelet Downsampling is performed to obtain four groups of subbands

[0019] S2.3: Perform attention calculation on the image pair, perceive the style features of the target domain through the attention module, and To calculate the cross attention, first Flattened into Then make As the query vector q H ,make as a key and value vector k H ,v H , use q H ,k H ,v H Perform cross attention calculation to obtain HH3, and Combine them for wavelet unpooling and perform group self-attention calculation on the obtained results.

[0020] Furthermore, the TTRPA mechanism described in S5 is a two-stage teacher-student region proposal alignment, which is constructed as follows:

[0021] S5.1: Pre-train the object detection model using data and annotations from the source domain: The resulting model is used as the student model, and a model with the same network structure as the teacher model is defined. The weights of the teacher model are updated from the weights of the student network via an exponential moving average (EMA), and its supervision loss is defined as:

[0022]

[0023] S5.2: Perform domain adaptation training: target domain image I T After weak enhancement, it is fed into the teacher model to generate pseudo labels for the first stage The student model is also supervised by pseudo labels. In the second stage, the region proposal P obtained by the teacher model at the final layer of the decoding layer is Tch-Decoder And the region proposal P obtained by the student model in the first layer of the decoding layer Stu-Decoder Merge and sort by Top K according to their classification scores to get a new region proposal B stage2 :

[0024] B stage2 =TopK(P Stu-Decoder +P Tch-Decoder ) (2)

[0025] S5.3: B stage2 This is fed into the subsequent decoders of the student and teacher models to obtain their category predictions:

[0026]

[0027] Among them B Stu stage2 and B Tch stage2 are the bounding box predictions of the student model and the teacher model, C Stu stage2 and C Tch stage2 are the category predictions of the student model and the teacher model respectively;

[0028] S5.4: Calculation and KL divergence, and based on it, construct the consistency loss L S consistancy :

[0029]

[0030] Where ζ is the scaling factor;

[0031] S5.5: Build domain classifiers in the backbone, encoder, and decoder for feature alignment to achieve adversarial loss L adv Optimization goal:

[0032]

[0033] S and D represent the student model and domain classifier respectively, and the loss L for domain classification is D is the weighted sum of each module;

[0034] S5.6: The overall objective of the model is the sum of the supervised loss, unsupervised loss, and domain classification loss:

[0035]

[0036] Among them, Lconsistancy stage2 is the consistency loss of the second stage of the model, and β is the weight factor of the domain classification loss and the multi-scale alignment loss.

[0037] Furthermore, paired images are used as input in the training phase, and single images are used as input in the inference phase. The input images are processed offline by WCT2 to generate the output images I corresponding to the target domain style. S_Tstyle , using ResNet50 as the feature extraction network, cross-attention calculation is performed on the extracted feature maps to perceive the style characteristics of the target domain.

[0038] Furthermore, in the pre-training stage of the model, the model uses Deformable DETR as the detector, and the teacher model and the student model share the same network structure. In the domain adaptation training stage of the model, the prediction of the teacher model is merged with the region proposal from the first decoder of the student model, and the results are re-ranked according to the category prediction to obtain the new top K region proposals.

[0039] The above technical solution can achieve the following beneficial effects:

[0040] This method combines the object content of the source domain and the style of the target domain into image pairs. The IAG module extracts high-order domain-invariant features and perceives the target domain style information. Through a two-stage teacher-student region proposal alignment (TTRPA) strategy, instance-level feature alignment is used to guide the model to assign higher attention weights to more effective regions, and a consistency loss is constructed to achieve suboptimal supervision. This method surpasses baseline methods in three domain adaptation scenarios, achieving state-of-the-art performance. Experimental results verify its effectiveness. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] Figure 1 This is a block diagram of the style-aware object detection method for underwater domain adaptation.

[0042] Figure 2 It is a method flow chart.

[0043] Figure 3 It is the constituent instance graph of the image pair.

[0044] Figure 4 It is an image attention guidance (IAG) module structure based on wavelet transform.

[0045] Figure 5 It is the structure of the two-stage teacher-student region proposal alignment (TTRPA) method.

[0046] Figure 6 It is the t-SNE visualization result diagram of different training stages. DETAILED DESCRIPTION

[0047] The following is combined with Figure 1-6 The present invention is further described as follows:

[0048] A style-aware object detection method for underwater domain adaptation is proposed. It constructs a wavelet-based Image Pair Attention Guidance (IAG) module, combining the source domain's object content and the target domain's style into image pairs. This information is then used as a set of complementary information for high-order feature modeling. Combining a cross-attention module with a discrete wavelet transform, the method performs differentiated content perception on the four subbands after wavelet decomposition, extracting high-order domain-invariant features and guiding the student model to learn cross-domain invariant information from the image pairs. Furthermore, a two-stage Teacher-Student Region Proposal Alignment (TTRPA) strategy is proposed. This strategy leverages instance-level feature alignment to guide the model to assign higher attention weights to more effective regions, and a consistency loss is constructed to achieve suboptimal supervision, thereby improving the model's generalization ability.

[0049] The method flow proposed by the present invention is as follows Figure 2 As shown. First, let the source domain image I S is the content image, the target domain image I T For the style image, the output image I with the target domain style is generated offline through the WCT2 model S_Tstyle , and with I S Constitute image pair I P =(I S ,I S_Tstyle), the two images in an image pair share the same annotations. Using the two-stage Deformable DETR as the backbone network, model training is divided into two stages. In the pre-training stage, feature extraction is performed using a ResNet50 pair to obtain multi-scale feature maps. An image-pair-based attention guidance module (IAG) is constructed to perform differential content perception on the four subbands after wavelet decomposition, guiding the student model to learn cross-domain invariant information from the "image pair." The teacher network's weights are updated from the student network's weights using an exponential moving average (EMA). In the domain adaptation training stage, the teacher network provides pseudo-labels, which are used to calculate the consistency loss between the student network's predictions on unlabeled data in the target domain and the teacher network's predictions on the target domain. Simultaneously, the teacher model outputs a high-confidence prediction box in the last decoder of the Deformable DETR. This prediction is merged with the student model's region proposals from the first decoder and re-ranked to obtain the top K region proposals. These are fed into the subsequent decoders of the student and teacher models, respectively, to obtain their respective category predictions. A consistency loss is constructed based on the classification results of the two predictions for suboptimal supervision. The Teacher model is encouraged to transfer the knowledge of relatively correct samples to the Student model, and the less confident samples in the Student model are transferred to the Teacher model, thereby improving the quality of pseudo labels through weighted consistency loss.

[0050] The image pair proposed by the present invention is constructed as follows Figure 3 As shown. Using the source domain image I S As the content image, the target domain image I T As a style image. Each content image I S Combined with the randomly selected style image, the WCT2 model is used to generate an output image I with the target domain style offline. S_Tstyle , and with I S The image pairs are fed into the feature extraction network of the model.

[0051] The structure of the image attention guidance (IAG) module based on wavelet transform proposed in this invention is as follows: Figure 4 As shown. Apply wavelet downsampling to the high-order feature map. P =(I S ,I S_Tstyle ) After Resnet50 feature extraction, three feature maps of different scales are obtained Using Haar wavelet Downsampling is performed to obtain four groups of subbands Then the attention calculation of the image pair is performed, and the style characteristics of the target domain are perceived through the attention module. To calculate the cross attention, first Flattened into Then make As the query vector q H ,make as a key and value vector k H ,v H . Use q H ,k H ,v H Perform cross attention calculation to get HH3. Combine them for wavelet unpooling and perform group self-attention calculation on the obtained results.

[0052] The two-stage teacher-student region proposal alignment (TTRPA) method proposed in this paper is as follows Figure 5 shown.

[0053] In the first step, the object detection model is pre-trained using the data and annotations from the source domain. The resulting model is used as the student model, and a model with the same network structure as the teacher model is defined. The weights of the teacher model are updated from the weights of the student network using the exponential moving average (EMA). The supervision loss is defined as:

[0054]

[0055] Among them, L S box is the bounding box regression loss of the model on the source domain, L S giou is the GIoU loss of the model, L S cls is the classification loss of the model, x s and y s are the prediction results and true values ​​of the model in the source domain respectively;

[0056] The second step is to conduct domain adaptation training. T After weak enhancement, it is fed into the teacher model to generate pseudo labels for the first stage The student model is also supervised by pseudo labels. In the second stage, the region proposal P obtained by the teacher model at the final layer of the decoding layer is Tch-Decoder And the region proposal P obtained by the student model in the first layer of the decoding layer Stu-Decoder Merge and sort by Top K according to their classification scores to get a new region proposal B stage2 :

[0057] B stage2 =TopK(P Stu-Decoder +P Tch-Decoder ) (2)

[0058] The third step is to stage2 This is fed into the subsequent decoders of the student and teacher models to obtain their category predictions:

[0059]

[0060] Among them B Stu stage2 and B Tch stage2 are the bounding box predictions of the student model and the teacher model, C Stu stage2 and C Tch stage2 are the category predictions of the student model and the teacher model respectively;

[0061] Step 4: Calculate and The KL divergence is constructed based on the consistency loss:

[0062]

[0063] In the fifth step, domain classifiers are built in the backbone, encoder, and decoder for feature alignment to achieve the adversarial optimization goal:

[0064]

[0065] S and D represent the student model and domain classifier respectively. The loss of domain classification is the weighted sum of each module.

[0066] Under this teacher-student region proposal alignment mechanism, the decoder of the teacher model further verifies and improves the pseudo-labels, continuously updating its internal representation to maintain predictive power and suppress potential erroneous sticky cycles.

[0067] The overall objective of the model is the sum of supervised loss, unsupervised loss, and domain classification loss:

[0068]

[0069] β is the weight factor for domain classification loss and multi-scale alignment loss.

[0070] The specific training methods are as follows:

[0071] The proposed model uses ResNet50 as the backbone network and Deformable DETR as the detector. The batch size is set to 8, and training is performed on two GTX 4090 GPUs. During the pre-training phase, a total of 70 training epochs are performed. The initial learning rate for the Deformable DETR and IAG modules is set to 2e-4, and the initial learning rate for the backbone module is set to 2e-5. The model is first trained on the source domain for 40 epochs using the original images as paired inputs. Subsequently, the learning rate for the Deformable DETR and IAG modules is reduced to 2e-5, and the model is trained for an additional 30 epochs using the style-transferred images as paired inputs. The self-training phase consists of 20 epochs, with an unsupervised coefficient λunsup of 1.0, a consistency loss parameter of 0.01, and an EMA smoothing coefficient α of 0.9996. The model is trained using the PyTorch framework with an Intersection over Union (IoU) threshold of 0.5.

[0072] Dataset and experimental scenario:

[0073] This model verifies the domain adaptation performance in three different scenarios:

[0074] Domain adaptation for underwater scenes. The DUO dataset includes 5,560 training images and 1,111 test images. The S-UODAC2020 dataset includes seven domains with different tones. Both datasets share the same target classes: sea cucumber, starfish, scallop, and sea urchin. In this scenario, DUO serves as the source domain and SUODAC as the target domain.

[0075] Weather Domain Adaptation. The Cityscapes dataset contains 2,975 training images and 500 test images. The FoggyCityscapes dataset uses a fog synthesis algorithm for cityscapes, preserving the same classes and content. In the domain adaptation test for this scenario, Cityscapes is used as the source domain, and FoggyCityscapes is used as the target domain.

[0076] Real-world domain adaptation. The Sim10K dataset contains 10,000 images synthesized by a game engine and serves as the source domain training set. Cityscapes is the target domain. Domain adaptation for this scenario was tested only on the "Car" category.

[0077] Comparison of experimental results:

[0078] Domain adaptation for underwater scenes. In the experiments for this scenario, WCT2 was used to generate image pairs with the style of the target domain offline. As shown in Table 1, the proposed model achieved the best performance in the process of transferring from DUO to S-UODAC2020. Notably, the model showed good results for the salient target "sea urchin" and the background easily confused target "sea cucumber". We also tested the detection performance of the model on the DUO dataset under non-cross-domain conditions. As shown in Table 2, the model achieved the best detection results.

[0079] Table 1 DUO→S-UODAC cross-domain detection results

[0080]

[0081] Table 2 DUO→DUO non-cross-domain detection results

[0082]

[0083] Weather scene domain adaptation. In the weather scene domain adaptation experiment, WCT2 was used to offline generate image pairs with the target domain style. As shown in Table 3, the proposed model demonstrated strong generalization capabilities. It achieved optimal detection performance for the salient object "Car" and effectively distinguished between the easily confused categories "Bicycle" and "Motorcycle."

[0084] Table 3 Cityscapes→Foggy Cityscapes cross-domain detection results

[0085]

[0086] Real-world domain adaptation. In the virtual-to-real-world domain adaptation experiments, image pairs were constructed directly from the source domain images themselves. As shown in Table 4, the model maintained the best performance in the generalization test between virtual and real environments, demonstrating its effectiveness in extracting domain-invariant features.

[0087] Table 4 BDD100K→Cityscapes cross-domain detection results

[0088]

[0089] Ablation experiment:

[0090] To evaluate the effectiveness of each module in the proposed model, we conducted ablation experiments with various configurations on the DUO and SUODAC datasets. We investigated the contribution of each component by adding each part in turn and observing the change in mAP performance. The results are summarized in Table 5.

[0091] Table 5 Ablation experiment (DUO→S-UODAC)

[0092]

[0093] t-SNE visualization:

[0094] We use t-SNE to visualize the feature similarity distribution. Figure 6 As shown in Figure 3, we randomly select 300 images from the source domain and 200 images from the target domain, and project the final layer features of the model at different training stages onto a two-dimensional plane. The green and purple points represent the source and target domain data, respectively. A good mixture of green and purple points indicates that the model is able to capture more domain-invariant features. Figure 6 The results in suggest that the IAG module helps extract domain-invariant features, while TTRPA promotes feature alignment.

[0095] The above are all preferred embodiments of the present invention. For ordinary technicians in this technical field, without departing from the principle of the present invention, various equivalent modifications to the present invention are within the scope of protection of the claims attached to this application.

Claims

1. A style-aware object detection method for underwater domain adaptation, characterized by: The method is as follows: S1: Pre-training the object detection model using data and annotations from the source domain Using the source domain image I S ∈R H×W×3 As the content image, the target domain image I T ∈R H×W×3 As a style image, each content image I S Combined with the randomly selected style image, the style transfer model is used to generate an output image I with the target domain style offline. S_Tstyle ∈R H×W×3 , and with I S Constitute image pair I P =(I S ,I S_Tstyle ), the two images in an image pair share the same annotation; S2: Image pair I P =(I S ,I S_Tstyle ) Feature extraction is performed through ResNet50, and then the IAG module is used to enable the model to focus on the feature representation before and after stylization; S3: The extracted features are regressed and classified through the Encoder layer and the Decoder layer to obtain the supervision loss of the model. The supervision loss L sup Defined as: Among them, L S box is the bounding box regression loss of the model on the source domain, L S giou is the GIoU loss of the model, L S cls is the classification loss of the model, x s and y s are the prediction results and true values ​​of the model in the source domain respectively; S4: The obtained model is used as the student model, and a model with the same network structure as the teacher model is defined. The weight of the teacher model is updated from the weight of the student network through the exponential moving average EMA, and the model is trained through supervised learning using classification loss. S5: The teacher model and the student model are aligned at the instance level for the data in the source and target domains through the TTRPA mechanism; S6: Construct the total target loss function and train the network to obtain the final target detection result.

2. The style-aware object detection method for underwater domain adaptation according to claim 1, characterized in that: The IAG module in S2 is an image-to-attention guidance module based on wavelet transform, and its construction method is as follows: S2.1: Constructing hue-based style transfer image pairs I P =(I S ,I S_Tstyle ); S2.2: Apply wavelet downsampling to high-order feature maps, I P =(I S ,I S_Tstyle ) After Resnet50 feature extraction, three feature maps of different scales are obtained Using Haar wavelet Downsampling is performed to obtain four groups of subbands S2.3: Perform attention calculation on the image pair, perceive the style features of the target domain through the attention module, and To calculate the cross attention, first Flattened into Then make As the query vector q H ,make as a key and value vector k H ,v H , use q H ,k H ,v H Perform cross attention calculation to obtain HH3, and Combine them for wavelet unpooling and perform group self-attention calculation on the obtained results.

3. The style-aware object detection method for underwater domain adaptation according to claim 1, characterized in that: The TTRPA mechanism described in S5 is a two-stage teacher-student region proposal alignment, which is constructed as follows: S5.1: Pre-train the object detection model using data and annotations from the source domain: The resulting model is used as the student model, and a model with the same network structure as the teacher model is defined. The weights of the teacher model are updated from the weights of the student network via an exponential moving average (EMA), and its supervision loss is defined as: S5.2: Perform domain adaptation training: target domain image I T After weak enhancement, it is fed into the teacher model to generate pseudo labels for the first stage The student model is also supervised by pseudo labels. In the second stage, the region proposal P obtained by the teacher model at the final layer of the decoding layer is Tch-Decoder And the region proposal P obtained by the student model in the first layer of the decoding layer Stu-Decoder Merge and sort by Top K according to their classification scores to get a new region proposal B stage2 : B stage2 =TopK(P Stu-Decoder +P Tch-Decoder ) (2) S5.3: B stage2 This is fed into the subsequent decoders of the student and teacher models to obtain their category predictions: Among them B Stu stage2 and B Tch stage2 are the bounding box predictions of the student model and the teacher model, C Stu stage2 and C Tch stage2 are the category predictions of the student model and the teacher model respectively; S5.4: Calculation and KL divergence, and based on it, construct the consistency loss L S consistancy : Where ζ is the scaling factor; S5.5: Build domain classifiers in the backbone, encoder, and decoder for feature alignment to achieve adversarial loss L adv Optimization goal: S and D represent the student model and domain classifier respectively, and the loss L for domain classification is D is the weighted sum of each module; S5.6: The overall objective of the model is the sum of the supervised loss, unsupervised loss, and domain classification loss: Among them, L consistancy stage2 is the consistency loss of the second stage of the model, and β is the weight factor of the domain classification loss and the multi-scale alignment loss.

4. The style-aware object detection method for underwater domain adaptation according to claim 1, characterized in that: Paired images are used as input in the training phase, and single images are used as input in the inference phase. The input images are used to generate the output images I corresponding to the target domain style offline through WCT2. S_Tstyle , using ResNet50 as the feature extraction network, cross-attention calculation is performed on the extracted feature maps to perceive the style characteristics of the target domain.

5. The style-aware object detection method for underwater domain adaptation according to claim 1, characterized in that: In the pre-training stage of the model, the model uses Deformable DETR as the detector, and the teacher model and the student model share the same network structure. In the domain adaptation training stage of the model, the prediction of the teacher model is merged with the region proposal from the first decoder of the student model, and the new top K region proposals are re-ranked according to the results of the category prediction.