Unmanned underwater vehicle cross-domain target detection method based on domain self-adaption
By constructing a cross-domain object detection data set and introducing spatial dropout, combined with the teacher-student model framework, the problem of UUV object detection algorithm dependence on training data is solved, and high-precision detection in different waters and target types is achieved, which improves the safe navigation and accident response capabilities of UUVs.
Patent Information
- Application Number
- CN202510320785.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-18
- Publication Date
- 2025-07-22
AI Technical Summary
The existing unmanned underwater vehicle (UUV) target detection algorithms rely on a large amount of training data, and it is difficult to achieve high-precision detection in different equipment, different waters, and different target types, especially in the problem of scarce training data in new ship sinking accidents.
By constructing a cross-domain object detection data set based on domain adaptation, using labeled ship optical remote sensing images as the source domain, and unlabeled side-swept sonar (SSS) images as the target domain, combining CycleGAN and CUT migration paths to change the image style, and introducing spatial dropout and teacher-student model framework for cross-domain detection, optimizing the loss function to improve detection accuracy.
It effectively alleviates the demand for training samples for UUV cross-domain object detection, improves detection accuracy, reduces the model's dependence on source domain features, and improves detection performance on target domains.
Smart Images

Figure CN120356079A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of unmanned underwater vehicles, and particularly relates to a cross-domain object detection method for unmanned underwater vehicles based on domain adaptation. Background Technique
[0002] The object detection technology of unmanned underwater vehicles (UUVs) plays an extremely important role in improving their navigation safety and the efficiency of ocean accident emergency response. UUVs are usually equipped with devices such as side-scan sonar (SSS). The SSS scans the underwater environment by emitting and receiving sound waves, and can generate high-resolution images in real time, which can be used to identify unknown underwater obstacles or objects, such as sunken ships, waste, submarine pipelines, etc. During the mission execution of UUVs, if these underwater obstacles are not discovered in time, they may cause serious damage to UUVs and even lead to mission failure. Therefore, by accurately detecting objects in advance, potential dangerous objects can be identified to avoid collisions, which can greatly ensure the safe navigation of UUVs and the smooth execution of missions. In addition, in case of emergencies at sea, such as ship sinkings or damage to offshore facilities, UUVs equipped with SSS can quickly scan the accident scene, provide large-scale and highly accurate underwater images, and help search and rescue personnel locate key data such as the position, depth, and shape of the accident ship or sunken object, thus significantly improving the search efficiency.
[0003] In recent years, due to the many advantages of deep learning-based unmanned underwater vehicle (UUV) object detection, it has gradually become a research hotspot. Data-driven deep learning object detection algorithms have made significant progress. Compared with traditional detection algorithms, they have great advantages. However, in the practical application of data-driven deep learning UUV object detection algorithms, many difficulties are faced. Because deep learning detection algorithms rely heavily on training data, too little data volume will not be able to fully learn all the features of the data, and the model is easily trapped in an overfitting state. For the UUV object detection task, collecting a sufficient number of training images is extremely challenging and time-consuming. At this time, the model has a good recognition effect on the training set, but the detection accuracy on the validation set is very poor. If the model is required to have a high detection accuracy under different devices, different waters, and different target types, a large amount of data collection and annotation need to be carried out in these scenarios, which is very time-consuming and difficult.
[0004] In addition, different from common object detection tasks, there is also an imperceptible problem in the UUV object detection task. That is, in the UUV object detection task, there cannot be a large number of real-time updated training images containing the required objects. A typical scenario is as follows: A new type of ship with a different form from traditional ships encounters an unpredictable danger and sinks during a transportation mission. Now, it is necessary to detect its specific location and then salvage it to analyze the cause of the accident. In this scenario, the training images of the new type of ship available for deep learning network learning are extremely scarce.
[0005] Regarding the problem of insufficient training data for UUV object detection, some scholars have conducted research. For example, Ye et al. applied VGG-11 and ResNet-18 to object recognition in underwater SSS images and used transfer learning to fine-tune the fully connected layer, which partially alleviated the problem of low recognition rate due to insufficient data samples. Huo et al. used semi-synthetic data and different modality data to fine-tune the parameters of VGG-19 to further improve its generalization ability. Using semi-synthetic data requires segmenting the object in the optical image and then transplanting it into the SSS image. There are still significant differences between the synthetic images and the SSS image samples, which directly affects the object recognition effect. Ganin et al. introduced the classic gradient reversal layer and designed instance-level and image-level alignments to improve the performance of the target domain. Shen et al. proposed a stacked complementary loss method based on gradient separation, which uses Faster R-CNN as the detector. Nguyen et al. proposed a method based on Faster R-CNN, which consists of two customized modules, namely local uncertainty attention alignment and multi-level uncertainty-aware context alignment. Shi et al. implemented general scale-aware domain adaptive Faster R-CNN and multi-label learning to reduce the negative transfer effect during training.
[0006] UUV cross-domain object detection has long been dominated by the outdated two-stage detector Faster R-CNN, which contains a region proposal network in most methods to perform local adaptation. Although this has been proven to improve the adaptation effect, in terms of efficiency and accuracy in the field of object detection, even the most powerful existing architecture (Faster RCNN+ResNet101) is far inferior to the recently proposed simple but more efficient YOLOv8. The outdated detector will limit the accuracy of cross-domain object detection. In addition, the Faster R-CNN architecture mainly adopted by the current domain adaptive detector needs to complete object detection through two stages, and its processing time is much higher than that of the single-stage YOLO model. Summary of the Invention
[0007] To overcome the deficiencies of the prior art, the present invention provides a cross-domain object detection method for an unmanned underwater vehicle based on domain adaptation. By using labeled ship optical remote sensing images as the source domain and unlabeled SSS images as the target domain, a UUV cross-domain object detection dataset is constructed for training a cross-domain object detection network. By fusing the domain adaptation framework with a conventional object detector, the demand of the network for training samples is alleviated, and the accuracy of UUV cross-domain object detection is improved. In addition, the method of the present invention introduces spatial dropout into the domain adaptation object detection network to prevent the model from overlearning the source domain features, further improving the detection accuracy on the target domain.
[0008] The technical solution adopted by the present invention to solve its technical problems is as follows:
[0009] Step 1: Define the source domain D s and the target domain D t ; The source domain is a dataset containing labeled images, and the target domain is a dataset without any labeled images, only containing sample images; Use and to represent the images in the source domain and the target domain respectively;
[0010] Step 2: Represent the i target categories in the source domain as Use to represent the labeled bounding boxes of each target in the source domain image, where x i , y i is the center point coordinate of the bounding box, and w i , h i are the width and height of the bounding box;
[0011] Step 3: For image-level domain difference repair, adopt two migration paths of CycleGAN and CUT to change the style of the images in the source domain, and mix them to obtain samples close to the target domain image style; During the generation process of the cross-domain target dataset, use and to represent the source domain images and target domain images with the style of the target domain image of the class obtained through style transfer respectively;
[0012] For instance-level domain difference repair, first introduce the YOLOv8 detector into the teacher-student framework; Then introduce spatial dropout to increase the randomness and diversity of the model, and reduce the dependence of the network on the source domain features;
[0013] Step 4: Use paired images for cross-domain distillation; For each distillation iteration, use the fake target domain image with the style of the source domain image of the class as the input of the teacher network, and use the real target domain image As the input of the student network; measure the weight transfer between the teacher model and the student model by defining the distillation loss; the input of the student network includes not only the real target domain images but also the real source domain images and the fake source domain images with the style of the class target domain images
[0014] Step 5: Use the real target domain images and the fake target domain images with the style of the class source domain images as the inputs of the student network and the teacher network respectively, and then use the instance predictions of the teacher network to guide the student network;
[0015] In each step during training, set the teacher network to the evaluation mode and use non-maximum suppression (NMS) to filter the predicted bounding boxes sorted by object confidence with a CIoU threshold; then, select the bounding box with the largest class score as the pseudo-label; the final pseudo-label provides the instance-level features of the target domain for the student network; the distillation loss between the teacher network and the student network is shown as follows:
[0016]
[0017] where M C is the class with the largest score from the teacher network, M B are the coordinates of the predicted bounding box with the largest class score from the teacher network, and L det (.) represents the loss of inputting the real image into the detector;
[0018] Step 6: Add the consistency loss as shown below:
[0019]
[0020] where represents the loss of inputting the fake image with the style of the class real image into the detector;
[0021] Step 7: Train the entire network by jointly optimizing the loss of the detector and the loss of the teacher-student model, which is carried out in an end-to-end manner, and the overall loss function is shown as follows:
[0022]
[0023] A computer program that causes a computer to execute the above-mentioned cross-domain object detection method for an unmanned underwater vehicle.
[0024] An electronic device includes: a processor and a memory; the memory is used to store a computer program, and the processor is used to execute the computer program stored in the memory so that the electronic device executes the above-mentioned cross-domain target detection method for an unmanned underwater vehicle.
[0025] A computer-readable storage medium stores a computer program thereon, and when the computer program is executed by a processor, the above-mentioned cross-domain target detection method for an unmanned underwater vehicle is implemented.
[0026] A chip includes: a processor, which is used to call and run a computer program from a memory, so that a device installed with the chip executes the above-mentioned cross-domain target detection method for an unmanned underwater vehicle.
[0027] A computer program product includes a computer storage medium, the computer storage medium stores a computer program, the computer program includes instructions that can be executed by at least one processor, and when the instructions are executed by the at least one processor, the above-mentioned cross-domain target detection method for an unmanned underwater vehicle is implemented.
[0028] The beneficial effects of the present invention are as follows:
[0029] 1. The present invention uses labeled ship optical remote sensing images as the source domain and unlabeled SSS images as the target domain to construct a cross-domain SSS image target detection dataset. By selecting two migration paths such as CycleGAN and CUT to change the style of the images in the source domain and mixing them, more abundant samples with a style closer to the target domain images are obtained.
[0030] 2. The present invention fuses the domain adaptation framework with a conventional target detector, alleviates the network's demand for training samples, and improves the cross-domain target detection accuracy of UUV. In addition, the present invention introduces spatial dropout into the domain adaptation target detection network to prevent the model from overlearning the source domain features, further improving the detection accuracy on the target domain. Description of the Drawings
[0031] Figure 1 It is a comparison schematic diagram of standard dropout and spatial dropout;
[0032] Figure 2 It is a schematic diagram of the generation process of the cross-domain target detection dataset (for repairing domain shift at the image level);
[0033] Figure 3 It is a network architecture diagram;
[0034] Figure 4 It is a visualization schematic diagram of the detection results of multiple models in the test set. Detailed Embodiments
[0035] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.
[0036] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.
[0037] The present invention proposes a cross-domain object detection method for unmanned underwater vehicles based on domain adaptation to alleviate the dependence of deep learning object detection algorithms on a large amount of labeled data. Taking the side-scan sonar image (SSS) carried by a UUV as an example, in this method, the labeled ship optical remote sensing image is used as the source domain, and the unlabeled SSS image is used as the target domain to construct a UUV cross-domain object detection dataset for training the cross-domain object detection network. By fusing the domain adaptation framework with a conventional object detector, the network's demand for training samples is alleviated, and the accuracy of UUV cross-domain object detection is improved. In addition, the present invention introduces spatial dropout into the domain adaptation object detection network to prevent the model from overlearning the source domain features, further improving the detection accuracy on the target domain.
[0038] 1. Application of Dropout technology;
[0039] Dropout technology is an easy-to-implement regularization method, initially mainly used for fully connected layers. By randomly discarding some neurons, it breaks the mutual dependence between neurons, improves the model's robustness to input changes and noise, and thus reduces the risk of overfitting. Nowadays, dropout is widely used in deep learning tasks such as image classification and image segmentation, becoming an important means to solve the overfitting problem in the training process. However, in object detection networks, the application of dropout is relatively rare. This is mainly because object detection networks are usually more complex than classification or segmentation networks. In addition to class prediction, they also need to accurately predict the position (bounding box) of the object. During training, dropout randomly suppresses the activation of some neurons, resulting in the loss of some spatial information in the feature map, which may have a negative impact on the detection task that requires precise positioning. In addition, since bounding box regression requires the network to have rich spatial information and high feature resolution, the operation of randomly discarding features may lead to inaccurate bounding box prediction and increase the instability during training. At the same time, some studies have shown that although standard dropout is very effective when applied to fully connected layers, its effectiveness will decrease when used together with convolutional layers.
[0040] Different from the data distribution consistency in ordinary tasks, cross-domain object detection faces significant differences in the data distributions of the source domain and the target domain. The source domain contains abundant annotated images, while the target domain usually lacks label information. In this case, the model is prone to overfitting the source domain data, resulting in poor generalization ability for the target domain. Therefore, the cross-domain detection task requires the model to extract features that are common in both the source domain and the target domain, namely "domain-invariant features", to achieve better generalization. Dropout increases the randomness and diversity of the model by randomly discarding some neurons during training, forcing the model to learn more robust features, thereby reducing the dependence on source-domain specific features. In this way, the model does not overly rely on certain specific features but instead focuses on those feature representations that are stable across different domains, enhancing the model's adaptability and robustness to domain variations. At the same time, source domain shift also causes the model to overly rely on the feature distribution of the source domain, resulting in poor performance on the target domain. Dropout can be used as a regularization method to reduce the impact of source domain data shift on the model to a certain extent. On the other hand, the random discard mechanism of dropout can also help the model adapt to the noise and variations in the data, making its performance on the target domain more robust. In addition, combined with Bayesian learning theory, dropout can be interpreted as "Monte Carlo Dropout" or Bayesian approximation, which helps estimate the uncertainty of the model in the target domain, prompts the model to focus on more general features, effectively alleviates the domain shift problem, and further improves the domain adaptation ability.
[0041] For the above reasons, the present invention introduces spatial dropout into the model and adds it to the last layer of the feature extraction backbone network. Different from standard dropout, spatial dropout discards in the spatial dimension of the feature map, that is, randomly discards entire channels at all positions in the same feature map. This method can preserve the spatial structure of the feature map and ensure that the relationships between features are maintained. At the same time, in each iteration, spatial dropout randomly deletes some feature map channels and forces the network to summarize the remaining feature maps, effectively avoiding the problem of adjacent pixels transmitting similar information. Figure 1 The schematic diagrams of standard dropout and spatial dropout are given.
[0042] 2. UUV cross-domain object detection method based on domain adaptation;
[0043] For a domain adaptation object detection network, two domains need to be defined first, namely the source domain D s and the target domain D t . The source domain is a dataset containing a large number of labeled images. The target domain is a dataset without any labeled images and only contains a small number of samples. Use and to represent the images in the source domain and the target domain. Denote the i target categories in the source domain as Use to represent the labeled bounding boxes of each target in the source domain image, where x i , y i is the center point coordinate of the bounding box, and w i , h i are the width and height of the border. The purpose of the domain adaptation network is to transfer the knowledge learned in the source domain to the target domain, and then achieve detection in the target domain. It needs to fix the offsets between different domains at the instance level and the instance level. The domain offset at the image level is mainly reflected in aspects such as image style and image size. The domain offset at the instance level is mainly reflected in aspects such as the appearance, size, and perspective of the targets in the image.
[0044] The classic domain - adaptive object detection framework is an adversarial process, and its core is the gradient reversal layer. It can reduce the classification error during forward training while maximizing the binary classification error and learning domain - invariant features during backpropagation. The classic domain - adaptive object detection framework uses distance metrics such as maximum mean discrepancy to calculate the domain offset. The latest domain - adaptive object detection framework uses a teacher - student model based on the knowledge distillation structure, which can improve the stability against data variance. Currently, this type of framework has been proven to be effective and has gradually become a research hotspot. In the domain - adaptive object detection framework based on the teacher - student model, the design of the teacher model and the student model is an important part. Research shows that the teacher model and the student model composed of outdated detectors will limit the accuracy of cross - domain object detection. On this basis, it is necessary to focus on how to fix the offsets between different domains at the image level and the instance level. At the same time, in the cross - domain object detection task, the model often tends to learn the labeled source domain information, resulting in low accuracy for the target domain.
[0045] Therefore, for the repair of domain differences at the image level, the present invention selects two transfer paths to change the style of the images in the source domain and mixes them to obtain more abundant samples closer to the target domain image style. The process of generating the cross - domain object dataset is as Figure 2 shown. Denote the source domain image with the target domain image style of the class object and the target domain image with the source domain image style of the class as and respectively through style transfer.
[0046] For the repair of domain differences at the instance level, the present invention first introduces the YOLOv8 detector into the latest teacher - student framework. Further, spatial dropout is introduced to increase the randomness and diversity of the model and reduce the network's dependence on source domain features. The architecture of the entire network is as Figure 3 shown.
[0047] 3. Loss function;
[0048] For the cross - domain object detection task based on the teacher - student model, it is necessary to design respective loss functions for the teacher model and the student model.
[0049] In the teacher - student model, the teacher model does not share weights with the student model, but uses the EMA weights of the student model. It can aggregate information after each step instead of each epoch. In addition, since the weight average improves the output of all layers, not just the top output, the target model has better intermediate representations. To make the teacher network produce more accurate predictions to provide better guidance for the student model, the present invention uses paired images for cross - domain distillation. For each distillation iteration, a fake target - domain image in the style of the class source - domain image is used as the input to the teacher network, and the real target - domain image is used as the input to the student network. Such processing can also make the student network increase its bias towards the features of the target - domain image during learning and reduce its dependence on the source - domain features. Based on this, the distillation loss is defined to measure the weight transfer between the teacher model and the student model. On the other hand, for the student network, in addition to the real target - domain image as its input, there are also the real source - domain images and the fake source - domain images in the style of the class target - domain image For the latter two, although their data distributions are different at the scene level, they belong to the same label space. To further enable the features in the two image styles to be better learned, a reasonable constraint is that the prediction outputs of the student network for the images in these two labeled domains are consistent.
[0050] The real target - domain image and the fake target - domain image in the style of the class source - domain image are respectively used as the inputs to the teacher network and the student network, and then the instance predictions with higher probabilities of the teacher network are used to guide the student network through distillation. Specifically, at each step during training, by setting the teacher network to the evaluation mode and using non - maximum suppression (NMS) to filter the predicted bounding boxes sorted by object confidence with a CIoU threshold. Then, the bounding box with the largest class score is selected as the pseudo - label. The final pseudo - label is provided to the instance - level features of the target domain of the student network.
[0051] The distillation loss between the teacher network and the student network is shown as follows.
[0052]
[0053] Where M Cis the class with the maximum score from the teacher network. M B are the predicted bounding box coordinates with the maximum class score from the teacher network.
[0054] As the model is trained and the weight parameters of the student network are passed to the teacher network via EMA, the student network will gradually approach the true target domain. At the same time, it will also weakly supervise the pseudo-labels in the teacher network's predictions for the input Although this weak supervision is not very accurate, these pseudo-labels play an important role in promoting instance-level adaptation.
[0055] For the student network, although the input source domain real images and the fake source domain images with the style of the class target domain images have different scene-level data distributions, they belong to the same label space. To further enable better learning of the features in the two image styles, a reasonable constraint is that the student network's prediction outputs for the images in these two labeled domains are consistent. Therefore, to ensure that their outputs are as consistent as possible, a new consistency constraint is added. The consistency loss for... is shown in the following equation.
[0056]
[0057] The entire network is trained by jointly optimizing the loss regarding the detector and the loss of the teacher-student model.
[0058] This is done in an end-to-end manner, and the overall loss function is shown in the following equation.
[0059]
[0060] Example:
[0061] To comprehensively demonstrate the performance of the method of the present invention in the UUV cross-domain object detection task, experiments were conducted on the SSS image domain adaptive dataset. In addition, to further illustrate the model performance, the present invention was experimentally compared with SSDA-YOLO and VersatileTeacher. The experimental results are shown in Table 1 below. As can be seen from the following table, compared with the compared models, the model of the present invention has higher accuracy. This proves the effectiveness of the present invention.
[0062] Table 1 Experimental results on the SSS image domain adaptive dataset
[0063] Model F1 mAP50 mAP50-95 SSDA-YOLO 0.608 63.5% 37.8% VersatileTeacher 0.676 68.6% 40.9% Our 0.713 76.3% 47.6%
[0064] In the typical scenarios of the test set, the detection results of multiple models are visualized as Figure 3As shown. Scenario 1: Both real targets and false targets exist simultaneously. Scenario 2: Only real targets exist. Scenario 3: The real target is an object not common in the training dataset. Scenario 4: Only false targets exist (the false target has a similarity to the shape of a common ship).
[0065] Through Figure 4 It can be seen that for Scenario 1, the present invention has the highest confidence in real targets and does not identify false targets as real targets. For Scenario 2, the present invention has the highest confidence. For Scenario 3, only the present invention successfully detects the target. For Scenario 4, the present invention does not identify false targets as real targets. In summary, it can be seen that the object detection accuracy, reliability, and stability of the present invention are superior to other methods.
Claims
1. A cross-domain target detection method for an unmanned underwater vehicle based on domain adaptation, characterized in that, It includes the following steps: Step 1: Define the source domain D s and the target domain D t ; the source domain is a dataset containing labeled images, and the target domain is a dataset without any labeled images, only containing sample images; use and to represent the images in the source domain and the target domain respectively; Step 2: Represent the i target categories in the source domain as Use to represent the labeled bounding boxes of each target in the source domain image, where x i , y i is the center point coordinate of the bounding box, and w i , h i are the width and height of the bounding box; Step 3: For image-level domain difference repair, two migration paths of CycleGAN and CUT are adopted to change the style of images in the source domain, and they are mixed to obtain samples close to the style of target domain images; during the generation process of the cross-domain target dataset, and are used to represent the source domain images with the style of the target domain of the class and the target domain images with the style of the source domain of the class obtained through style transfer; For instance-level domain difference repair, first, introduce the YOLOv8 detector in the teacher-student framework; then introduce spatial dropout to increase the randomness and diversity of the model and reduce the network's dependence on source domain features; Step 4: Use paired images Perform cross - domain distillation; for each distillation iteration, use the fake target - domain image in the style of the source - domain images as the input to the teacher network, and use the real target - domain image as the input to the student network; define the distillation loss to measure the weight transfer between the teacher model and the student model; in addition to the real target - domain image as the input to the student network, there are also the real source - domain images and the fake source - domain images in the style of the target - domain images Step 5: Use the real target domain image and the fake target domain image with the style of the source domain image as the inputs of the student network and the teacher network respectively, and then use the instance prediction of the teacher network to guide the student network; In each step during training, by setting the teacher network to the evaluation mode and using non-maximum suppression (NMS) to filter the predicted bounding boxes sorted by object confidence with the CIoU threshold; then, select the bounding box with the largest class score as the pseudo-label; the final pseudo-label provides the instance-level features of the target domain to the student network; the distillation loss between the teacher network and the student network is shown as follows: Among them, M C is the class with the maximum score from the teacher network, and M B are the predicted bounding box coordinates with the maximum class score from the teacher network. L det (.) represents the loss of inputting the real image into the detector; Step 6: Add the consistency loss as shown in the following formula: Among them, represents the loss of inputting a fake image in the style of a real image into the detector; Step 7: Train the entire network by jointly optimizing the loss regarding the detector and the loss of the teacher-student model, which is carried out in an end-to-end manner, and the overall loss function is shown as follows:
2. A computer program, characterized in that, The computer program causes the computer to execute the method as claimed in claim 1.
3. An electronic device, characterized in that, It includes: A processor and a memory; The memory is used to store the computer program, and the processor is used to execute the computer program stored in the memory so that the electronic device executes the method as claimed in claim 1.
4. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the method as claimed in claim 1.
5. A chip, characterized in that, It includes: A processor for calling and running the computer program from the memory, so that the device installed with the chip executes the method as claimed in claim 1.
6. A computer program product, characterized in that, The computer program product includes a computer storage medium, the computer storage medium stores the computer program, and the computer program includes instructions that can be executed by at least one processor. When the instructions are executed by the at least one processor, the method as claimed in claim 1 is implemented.