Marine target unsupervised learning detection method based on auxiliary generation network
By constructing a pseudo detection box encoding to generate simulated images, using the auxiliary generation network for unsupervised learning, the problems of insufficient samples and high labeling cost in maritime target detection are solved, and efficient target detection is achieved.
Patent Information
- Application Number
- CN202510476275.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2025-07-18
AI Technical Summary
Maritime target detection faces the problems of insufficient samples and high labeling costs, especially the difficulty of obtaining enough training samples for certain special types of maritime targets, which limits the training and performance of supervised learning models.
Build a pseudo detection box encoding to generate simulated images, perform unsupervised learning through the auxiliary generation network, and use the pseudo detection box encoding and image generation module to train the object detection and discrimination module to realize unsupervised object detection.
Without manual annotation, the accuracy and adaptability of the model to detect maritime targets is improved, the labeling cost is reduced, and the detection effect is improved.
Smart Images

Figure CN120339587A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of target detection, and particularly relates to a method for detecting marine targets. Background Art
[0002] With the increasing frequency of marine activities, the demand for marine target monitoring is continuously growing. However, due to the scarcity and difficulty of collection of some targets, problems such as insufficient samples and high annotation costs are faced in marine target detection, seriously restricting the development and application in this field. Especially for some special types of marine targets, it is particularly difficult to obtain enough training samples, limiting the training and performance of supervised learning models.
[0003] Unsupervised learning object detection technology is gradually becoming a research hotspot due to its characteristic of not requiring labeled data, which can significantly reduce the annotation cost and improve the adaptability and flexibility of the model to new scenarios.
[0004] In the field of target detection, the paper CutLER (Cut and Learn for Unsupervised Object Detection and Instance Segmentation) proposed a scheme that uses a self-supervised model to automatically discover targets in images and uses them as ground truth to train detection and segmentation models, without the need for manual annotation throughout the process. This method further improves the model performance through multiple self-training. However, it mainly focuses on the detection of vertical boxes in natural images, which is significantly different from the detection of oblique boxes in SAR images.
[0005] Some research focuses on enhancing the quality of SAR images to improve the detection effect. For example, Xinyang et al. proposed an unsupervised domain adaptation method based on CycleGAN (Xinyang, D., et al. Ship detection in low-quality SAR images via an unsupervised domain adaption method. IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, 2023, 16, 1234-1242.), which maps high-quality SAR images to the low-quality domain through image conversion and uses the cycle consistency loss to maintain the invariance of target localization. The generated images combined with the original labels are used as training data, effectively improving the detection accuracy under complex interference. However, this method still relies on labeled data and has not completely got rid of the limitations of supervised learning.
[0006] On the other hand, Liu Xing et al. proposed an unsupervised domain - adaptive SAR ship detection method based on cross - domain feature interaction and data contribution balance (Unsupervised Domain - Adaptive SAR Ship Detection Based on Cross - Domain Feature Interaction and Data Contribution Balance, Remote Sensing, 2024). This method designed a CycleGAN - SCA module for unbiased feature generation from the optical domain to the SAR domain. However, this method belongs to the category of transfer learning, and its source domain still requires finely labeled data.
[0007] In addition, the Chinese invention patent with the authorization announcement number CN118397256B proposed a method and device for detecting ship targets in SAR images. It automatically generates part - level labels through Gaussian heatmaps, reducing the cost of manual annotation. Although this method can automatically mine the key areas of the hull, its detection model still needs to rely on SAR image samples with labeled bounding boxes for training. Summary of the Invention
[0008] The present invention proposes an unsupervised learning detection method for maritime targets based on an auxiliary generation network, and its purpose is to solve the problems of few specific targets and high cost of labeled data in the field of maritime target detection.
[0009] The technical solution of the present invention is as follows: An unsupervised learning detection method for maritime targets based on an auxiliary generation network, the steps include: Step 1: Collect real images containing maritime targets to obtain a target data set; Step 2: Construct a pseudo - detection - box generation module for generating random pseudo - detection - box encodings; Step 3: Construct an image generation module, which is used to obtain corresponding simulated images according to the input pseudo - detection - box encodings; Step 4: Construct a target detection discrimination module, which includes a feature learning network, a real - fake discrimination branch, and a target detection branch; the feature learning network is used to extract features from the input images, the real - fake discrimination branch is used to judge whether the input image is a real image or a simulated image according to the extracted features, and the target detection branch is used to perform target detection on the input image according to the extracted features to obtain the detection bounding boxes of maritime targets in the image; Step 5: Train the target detection discrimination module based on the real images, the pseudo - detection - box encodings output by the pseudo - detection - box generation module, and the simulated images output by the image generation module according to the pseudo - detection - box encodings; Step 6: Input the image to be detected into the target detection and discrimination module. The predicted detection box coordinate information output by the target detection branch is the target detection result.
[0010] As a further improvement of the unsupervised learning detection method for maritime targets based on the auxiliary generation network: The pseudo-detection box encoding refers to an encoding vector containing the coordinate information of more than one pseudo-detection box. The length of the encoding vector is L. Suppose the upper limit of the number of pseudo-detection box coordinate information contained in the encoding vector is M, and a pseudo-detection box coordinate information includes n position information, then L > M * n. Suppose a certain encoding vector contains m pseudo-detection box coordinate information. The m pseudo-detection box coordinate information is arranged in sequence starting from the 1st position of the encoding vector in a unified order, and the remaining L - m * n positions are filled with random values.
[0011] As a further improvement of the unsupervised learning detection method for maritime targets based on the auxiliary generation network: The distribution of the number of pseudo-detection box coordinate information in all encoding vectors is set according to the distribution of the number of targets in the target dataset.
[0012] As a further improvement of the unsupervised learning detection method for maritime targets based on the auxiliary generation network: The pseudo-detection box coordinate information includes 5 position information, expressed as: (x, y, h, w, θ); where x is the abscissa value of the upper left corner of the detection box, y is the abscissa value of the upper left corner of the detection box, h is the length of the detection box, w is the width of the detection box, and θ is the tilt angle of the detection box.
[0013] As a further improvement of the unsupervised learning detection method for maritime targets based on the auxiliary generation network, the pseudo-detection box coordinate information is randomly generated in the following way: x and y are random values greater than 0 and less than the image size; θ is a random value greater than or equal to 0 and less than 180°; h is a random value subject to a Gaussian distribution: w is calculated according to the generated h and the randomly generated aspect ratio obtained as follows: .
[0014] As a further improvement of the unsupervised learning detection method for maritime targets based on the auxiliary generation network, in the pseudo-detection box generation module: , where is a random number generated from a standard normal distribution; The aspect ratio is generated in the following way: , , , is a random number generated from a standard normal distribution.
[0015] As a further improvement to the above-mentioned unsupervised learning detection method for maritime targets based on an auxiliary generation network: The image generation module includes 1 fully connected layer, 2 deconvolution layers, 1 lightweight module, 2 deconvolution layers, 1 lightweight module, 4 deconvolution-lightweight-batch normalization modules, and 1 deconvolution layer connected in sequence; The deconvolution-lightweight-batch normalization module includes 1 deconvolution layer, 1 lightweight module, and 1 batch normalization layer connected in sequence.
[0016] As a further improvement to the above-mentioned unsupervised learning detection method for maritime targets based on an auxiliary generation network: The lightweight module is the DDA-Fire Module, which includes a dynamic compression layer, a multi-scale expansion layer, and a channel-space dual attention layer, with the input being features , and the output being features , and the working process is as follows: (1) In the dynamic compression layer, the multi-layer perceptron MLP obtains the complexity based on the features , and then inputs the features into a 1x1 convolutional layer to obtain the features . The features then pass through the first deformable convolutional layer to obtain the features ; (2) In the multi-scale expansion layer, the features pass through the second deformable convolutional layer to obtain the features with the output channel number of . At the same time, the features pass through the dilated convolutional layer to obtain the features with the output channel number of . Then, the features and the features are concatenated along the channel dimension to obtain the expanded features ; (3) In the channel-space dual attention layer, the features are respectively input into the channel attention module and the spatial attention module to obtain and ; ; ; Among them, is the sigmoid function, is a multi-layer perceptron composed of 3 fully connected layers, is average pooling, is convolution; Multiply and Feature is obtained through weighted fusion to enhance the sensitivity to the edges and textures of the oblique boxes and alleviate the problem of gradient disappearance; (4) Finally, the feature is connected with the input feature through residual connection to obtain the feature .
[0017] As a further improvement of the unsupervised learning detection method for marine targets based on the auxiliary generation network: The RoI Transformer object detection discrimination module is used as the object detection discrimination module, and the original object classification branch in the RoI Transformer is used as the true / false discrimination branch.
[0018] As a further improvement of the unsupervised learning detection method for marine targets based on the auxiliary generation network, in step 5, steps 5-1 to 5-3 are repeated several times to complete the training: Step 5-1: Use the pseudo-detection box generation module to obtain e pseudo-detection box encodings, and then input the obtained pseudo-detection box encodings into the image generation module respectively to obtain the corresponding e simulated images; Step 5-2: Mix the e real images and the e simulated images, input them into the object detection discrimination module respectively, and obtain the corresponding image true / false discrimination results from the true / false discrimination branch , and then calculate the discrimination loss function corresponding to the image according to the image true / false discrimination result and the true label of the input image ; Update the network parameters of the feature learning network, the true / false discrimination branch, and the image generation module based on the average value of the e discrimination loss functions; The calculation method of the discrimination loss function is as follows: ; In the formula, the true / false discrimination result is the probability that the image is a real image; the true label is 1 or 0. When the image is a real image and otherwise ; Step 5-3: Input the e simulated images into the object detection discrimination module respectively, obtain the predicted detection box coordinate information from the object detection branch, and then calculate the object detection loss function according to the predicted detection box coordinate information and the pseudo-detection box coordinate information in the pseudo-detection box encoding corresponding to the original simulated image, and update the network parameters of the object detection branch and the image generation module based on the average value of the e object detection loss functions.
[0019] Compared with the prior art, the present invention has the following beneficial effects: 1. The present invention constructs a pseudo-detection box code to obtain a simulated image, and then uses the simulated image and the pseudo-detection box code as training samples to train the target detection and discrimination module, achieving unsupervised learning and target detection, and solving the problems of few specific targets and high cost of labeled data in the field of maritime target detection.
[0020] 2. When generating the pseudo-detection box code, the present invention adopts a random generation method, and the generated parameters conform to the real distribution, laying a foundation for subsequent image generation, ensuring that the obtained simulated image and pseudo-detection box are similar to the real situation, and improving the adaptability of the model and the accuracy of detection.
[0021] 3. The image generation module of the present invention can combine with the subsequent true / false discrimination branch to generate a simulated image that is consistent with the real image distribution and more realistic. At the same time, the lightweight module Fire Module is used in this module, which can replace part of the 3x3 convolution with 1x1 convolution, significantly reducing the number of network parameters, and can improve the network performance and inference speed while maintaining the quality of the simulated image.
[0022] 4. The present invention mixes real images and simulated images to perform adversarial training on the target detection and discrimination module and the image generation module, improving the realism of the simulated image, thus laying a foundation for unsupervised learning and enhancing the detection ability of the model.
[0023] 5. The present invention effectively re-uses the RoI Transformer model, especially using its target classification branch for image true / false judgment, efficiently utilizing the feature learning network shared by the two branches without increasing the workload and without destroying the structure of the RoI Transformer backbone network. And during training, the target detection branch and the true / false discrimination branch form a two-way constraint through the shared feature extraction network: on the one hand, the target detection branch supervises the target position of the simulated image through the pseudo-detection box to ensure the semantic rationality of the simulated image; on the other hand, the true / false discrimination branch forces the simulated image to approach the real distribution through adversarial training, indirectly improving the adaptability of the detection branch to the real scene, and finally being able to achieve the oblique box detection of ship targets in SAR images without the need for manually labeled oblique box detection labels.
[0024] 6. The image generation module uses the improved DDA-Fire Module to replace the traditional lightweight module FireModule, and this module introduces a dynamic compression rate (complexity ), and dynamically adjust the number of channels according to the complexity of the input features, which can avoid the loss of small target information; at the same time, introduce deformable convolutional kernels to learn the offset of target deformation and accurately model the geometric features of inclined box targets; finally, construct a dual attention mechanism to enhance the weights of important channels through channel attention and suppress background noise; generate a spatial weight map through spatial attention to focus on the target edge area.
[0025] In summary, based on the framework of the auxiliary generation network, the present invention designs a method for unsupervised object detection. This method can generate simulation images consistent with the distribution of real images through an image generation module, optimize the quality of simulation images through a true / false discrimination branch, and achieve object detection of simulation images through the object detection branch and corresponding pseudo-detection box labels in the object detection discrimination module. Until the training is completed, since the generated simulation images are consistent with the distribution of real images and have a high degree of verisimilitude, and the strong correspondence between pseudo-detection boxes and simulation images is constrained during the training process, the trained object detection branch can also complete the object detection task of real images, thus realizing object detection under unsupervised learning and reducing the cost of manual annotation. Brief Description of the Drawings
[0026] Figure 1 is a schematic diagram of the principle of the method of the present invention; Figure 2 is a network structure diagram of the image generation module; Figure 3 is a network structure diagram of the DDA-Fire Module. Detailed Description of the Embodiments
[0027] The technical solution of the present invention will be described in detail below with reference to the drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments.
[0028] An unsupervised learning detection method for maritime targets based on an auxiliary generation network, the steps include: Step 1: Collect real images containing maritime targets to obtain a target data set.
[0029] The real image is a captured ship SAR image. After collection, the size is adjusted to 512×512 to construct a SAR image ship target data set, and then the SAR image ship target data set is randomly divided into training and test samples according to a 7:3 ratio to construct a training set and a test set.
[0030] Step 2: Construct a pseudo-detection box generation module for generating random pseudo-detection box encodings.
[0031] The pseudo-detection box encoding refers to an encoding vector containing the coordinate information of one or more pseudo-detection boxes.
[0032] The length of the encoding vector is L. Suppose the upper limit of the number of pseudo-detection box coordinate information contained in the encoding vector is M, and a piece of pseudo-detection box coordinate information includes n position information (occupying n positions), then L > M * n.
[0033] The distribution of the number of pseudo-detection box coordinate information in all encoding vectors is set according to the distribution of the number of targets in the target dataset. In this embodiment, a total of e = 128 encoding vectors (pseudo-detection box encodings) are generated, and on average, each image slice in the target dataset has 1.46 targets. Therefore, it is set that the first 64, the 65th to 96th, the 97th to 113th, the 114th to 121st, the 122nd to 125th, the 126th to 127th, and the 128th encoding vectors contain 1, 2, 3, 4, 5, 6, and 7 targets respectively (one target corresponds to one piece of pseudo-detection box coordinate information). Suppose a certain encoding vector contains m pieces of pseudo-detection box coordinate information, and the m pieces of pseudo-detection box coordinate information are arranged in sequence starting from the 1st position of the encoding vector in a unified order, and the remaining L - m * n positions are filled with random values (such as random values within 0 to 1).
[0034] Specifically, the pseudo-detection box coordinate information includes 5 position information (i.e., n = 5), expressed as: (x, y, h, w, θ); where x is the abscissa value of the upper left corner of the detection box, y is the abscissa value of the upper left corner of the detection box, h is the length of the detection box, w is the width of the detection box, and θ is the tilt angle of the detection box.
[0035] The pseudo-detection box coordinate information is randomly generated in the following way: x and y are random values greater than 0 and less than the image size. In this embodiment, 0 < x, y < 500; θ is a random value greater than or equal to 0 and less than 180°; h is a random value following a Gaussian distribution. In this embodiment ; where is a random number generated by a standard normal distribution; w is calculated according to the already generated h and the randomly generated aspect ratio : , and the aspect ratio is generated in the following way: , , , is a random number generated by a standard normal distribution.
[0036] Step 3, construct an image generation module, and the image generation module is used to obtain a corresponding simulation image according to the input pseudo-detection box encoding.
[0037] The image generation module contains an improved lightweight module DDA - Fire Module to reduce the number of parameters and speed up the model training and inference speed.
[0038] Specifically, as Figure 2, the image generation module includes a fully connected layer, two transposed convolutional layers, a lightweight module, two transposed convolutional layers, a lightweight module, four transposed convolutional-lightweight-batch normalization modules, and a transposed convolutional layer connected in sequence. The transposed convolutional-lightweight-batch normalization module includes a transposed convolutional layer, a lightweight module, and a batch normalization layer connected in sequence.
[0039] Furthermore, although the conventional lightweight module Fire Module achieves lightweight design through a compression-expansion structure and multi-scale convolution, there are still the following problems in complex scenarios (such as oblique box target detection in SAR images): First, the Squeeze compression layer uses a fixed compression ratio, which may cause loss of important features, especially the key information of small targets or sparse targets is over-compressed by the above fixed compression method; Second, the 3x3 convolutional kernel in the Expand expansion layer can only capture local features of a fixed shape and cannot adapt to geometric changes such as rotation and deformation of oblique box targets; Third, there is a lack of a dynamic attention mechanism, which cannot distinguish the importance of channels or spatial positions, resulting in a decrease in the generation or detection accuracy under noise interference.
[0040] To solve the above problems, the present invention uses an improved DDA-Fire Module (dynamic deformable attention-based lightweight module) to replace the original Fire Module in the image generation module.
[0041] As Figure 3 , the DDA-Fire Module includes a dynamic compression layer, a multi-scale expansion layer, and a channel-spatial dual attention layer, with the input being the feature , and the output being the feature , and the working process is as follows: (1) In the dynamic compression layer, the multi-layer perceptron MLP obtains the complexity according to the feature , and then inputs the feature into a 1x1 convolutional layer to obtain the feature , and the feature then passes through the first deformable convolutional layer (the deformable convolutional layer is Deformable Conv2D, abbreviated as DefConv) to obtain the feature . The deformable convolution dynamically adjusts the sampling position of the convolutional kernel through the learned offset, thereby enhancing the modeling ability for oblique boxes or deformed targets.
[0042] represents the floor function.
[0043] The complexity is adaptively generated by a 3-layer MLP, which can avoid information loss caused by a fixed compression rate.
[0044] (2) In the multi-scale expansion layer, the feature passes through the second deformable convolutional layer (3×3 deformable convolution) to obtain a feature with an output channel number of . At the same time, the feature passes through the dilated convolutional layer (3×3 dilated convolution with a dilation rate of 2) to obtain a feature with an output channel number of . Then, the feature and the feature are concatenated along the channel dimension to obtain the expanded feature .
[0045] (3) In the channel-spatial dual attention layer, the feature is respectively input into the channel attention module and the spatial attention module to obtain and .
[0046] Specifically: ; ; Among them, is the sigmoid function, is a multi-layer perceptron composed of 3 fully connected layers, is average pooling, is convolution.
[0047] Multiply and and fuse them with weights to obtain the feature , thereby enhancing the sensitivity to the edges and textures of the oblique boxes and alleviating the problem of gradient disappearance.
[0048] (4) Finally, perform a residual connection between the feature and the input feature to obtain the feature .
[0049] The improved DDA-Fire Module introduces a dynamic compression rate (complexity ), dynamically adjusts the number of channels according to the complexity of the input feature, and avoids the loss of small target information; at the same time, it introduces a deformable convolution kernel to learn the offset of the target deformation and accurately model the geometric features of the oblique box target; finally, it constructs a dual attention mechanism, enhances the weights of important channels through channel attention, and suppresses background noise; generates a spatial weight map through spatial attention and focuses on the target edge area.
[0050] Step 4: Construct a target detection and discrimination module, which includes a feature learning network, a true / false discrimination branch, and a target detection branch. The feature learning network is used to extract features from the input image. The true / false discrimination branch is used to determine whether the input image is a real image or a simulated image based on the extracted features. The target detection branch is used to perform target detection on the input image based on the extracted features to obtain the detection boxes of marine targets in the image.
[0051] In this embodiment, the RoI Transformer target detection and discrimination module is used as the target detection and discrimination module. The original target classification branch in RoI Transformer is used for classifying the target type. Here, this branch is used as the true / false discrimination branch to judge the authenticity of the image.
[0052] Step 5: Train the target detection and discrimination module based on real images, the pseudo-detection box codes output by the pseudo-detection box generation module, and the simulated images output by the image generation module.
[0053] As Figure 1 , during training, repeat steps 5-1 to 5-3 a certain number of times (1000 times in this embodiment): Step 5-1: Use the pseudo-detection box generation module to obtain e pseudo-detection box codes, and then input the obtained pseudo-detection box codes into the image generation module respectively to obtain the corresponding e simulated images.
[0054] In this embodiment, the batch size e = 128.
[0055] Step 5-2: Mix e real images and e simulated images, and input them into the target detection and discrimination module respectively. Obtain the corresponding image true / false discrimination results from the true / false discrimination branch , and then according to the image true / false discrimination results and the true label of the input image calculate the discriminant loss function corresponding to this image: ; In the formula, the true / false discrimination result is the probability that the image is a real image; the true label is 1 or 0. When the image is a real image , otherwise ; Based on the average value of e discriminant loss functions update the network parameters of the feature learning network, the true / false discrimination branch, and the image generation module.
[0056] The purpose of this step is to simultaneously improve the discrimination ability of the authenticity discrimination branch for real and fake images and the realism of the simulated images output by the image generation module. Eventually, in the collaborative optimization of the two, the discrimination ability of the authenticity discrimination branch gradually becomes stronger, and the forgery ability of the image generation module also gradually becomes stronger.
[0057] Step 5-3: Input the e simulated images into the target detection and discrimination module respectively, obtain the predicted detection box coordinate information from the target detection branch, then calculate the target detection loss function according to the predicted detection box coordinate information and the pseudo-detection box coordinate information in the pseudo-detection box encoding corresponding to the original simulated image, and update the network parameters of the target detection branch, the feature learning network, and the image generation module based on the average value of the e target detection loss functions.
[0058] When calculating the target detection loss function, the intersection over union can be used as the target detection loss function.
[0059] In this step, in addition to improving the detection ability of the target detection branch, the image generation module is also re-constrained based on the difference between the pseudo-detection box and the detection result of the simulated image to ensure that the ship target position information in the generated simulated image can accurately correspond to the pseudo-detection box, and improve the generation effect of the ship target in the simulated image.
[0060] Step 6: Input the image to be detected into the target detection and discrimination module, and the predicted detection box coordinate information output by the target detection branch is the target detection result.
[0061] It should be noted that for those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and without departing from the spirit or basic characteristics of the present invention, the present invention can be implemented in other specific forms. The scope of the present invention is defined by the claims rather than the above description.
Claims
1. An unsupervised learning detection method for maritime targets based on an auxiliary generation network, characterized in that the steps Including: Step 1: Collect real images containing marine targets to obtain a target dataset; Step 2: Construct a pseudo-detection box generation module for generating random pseudo-detection box encodings; Step 3: Construct an image generation module, which is used to obtain corresponding simulated images according to the input pseudo-detection box encodings; Step 4: Construct a target detection discrimination module, which includes a feature learning network, a real / fake discrimination branch, and a target detection branch; The feature learning network is used to extract features from the input images. The real / fake discrimination branch is used to determine whether the input image is a real image or a simulated image according to the extracted features. The target detection branch is used to perform target detection on the input image according to the extracted features to obtain the detection boxes of marine targets in the image; Step 5: Train the target detection discrimination module based on the real images, the pseudo-detection box encodings output by the pseudo-detection box generation module, and the simulated images output by the image generation module according to the pseudo-detection box encodings; Step 6: Input the image to be detected into the target detection discrimination module, and the predicted detection box coordinate information output by the target detection branch is the target detection result.
2. The unsupervised learning detection method for maritime targets based on an auxiliary generation network according to claim 1, characterized in that: The pseudo-detection box encoding refers to an encoding vector containing more than one pseudo-detection box coordinate information; The length of the encoding vector is L. Let the upper limit of the number of pseudo-detection box coordinate information contained in the encoding vector be M. One pseudo-detection box coordinate information includes n position information, then L > M * n; Suppose a certain encoding vector contains m pseudo-detection box coordinate information. The m pseudo-detection box coordinate information is arranged in sequence starting from the 1st position of the encoding vector in a unified order, and the remaining L - m * n positions are filled with random values.
3. The unsupervised learning detection method for maritime targets based on an auxiliary generation network according to claim 2, characterized in that: The distribution of the number of pseudo-detection box coordinate information in all encoding vectors is set according to the distribution of the number of targets in the target dataset.
4. The unsupervised learning detection method for maritime targets based on an auxiliary generation network according to claim 2, wherein: The pseudo-detection box coordinate information includes 5 position information, expressed as: (x, y, h, w, θ); where x is the abscissa value of the upper left corner of the detection box, y is the abscissa value of the upper left corner of the detection box, h is the length of the detection box, w is the width of the detection box, and θ is the tilt angle of the detection box.
5. The unsupervised learning detection method for maritime targets based on an auxiliary generation network according to claim 4, wherein The coordinates of the pseudo-detection box are randomly generated in the following way: x and y are random values greater than 0 and less than the image size; θ is a random value greater than or equal to 0 and less than 180°; h is a random value following a Gaussian distribution: w is calculated based on the generated h and the randomly generated aspect ratio Calculated as: .
6. The unsupervised learning detection method for maritime targets based on an auxiliary generation network according to claim 5, characterized in that, In the pseudo-detection box generation module: , where is a random number generated from the standard normal distribution; Aspect ratio is generated in the following way: , , , is a random number generated from the standard normal distribution.
7. The unsupervised learning detection method for maritime targets based on an auxiliary generation network according to claim 1, characterized in that: The image generation module includes 1 fully connected layer, 2 deconvolution layers, 1 lightweight module, 2 deconvolution layers, 1 lightweight module, 4 deconvolution-lightweight-batch normalization modules, and 1 deconvolution layer connected in sequence; The deconvolution-lightweight-batch normalization module includes 1 deconvolution layer, 1 lightweight module, and 1 batch normalization layer connected in sequence.
8. The unsupervised learning detection method for maritime targets based on an auxiliary generation network according to claim 1, wherein: The lightweight module is the DDA-Fire Module, which includes a dynamic compression layer, a multi-scale expansion layer, and a channel-spatial dual attention layer. The input is a feature and the output is a feature . The working process is as follows: In the dynamic compression layer, the multi-layer perceptron MLP obtains the complexity according to the features and then inputs the features into the 1x1 convolutional layer to obtain the features . The features are then passed through the first deformable convolutional layer to obtain the features ; (2)In the multi-scale expansion layer, the feature passes through the second deformable convolutional layer to obtain a feature with an output channel number of ; meanwhile, the feature passes through the dilated convolutional layer to obtain a feature with an output channel number of ; then, the feature and the feature are concatenated along the channel dimension to obtain the expanded feature ; ; ; In the channel spatial dual attention layer, the feature is respectively input into the channel attention module and the spatial attention module to obtain and ; ; ; Among them, is the sigmoid function, is a multi-layer perceptron composed of 3 fully connected layers, is average pooling, is convolution; Combine and through weighted fusion to obtain the feature , so as to enhance the sensitivity to the edges and textures of the oblique boxes and alleviate the problem of gradient disappearance; (4) Finally, connect the feature with the input feature through a residual connection to obtain the feature .
9. The unsupervised learning detection method for maritime targets based on an auxiliary generation network according to claim 1, wherein: Use the RoI Transformer target detection discrimination module as the target detection discrimination module, and use the original target classification branch in the RoI Transformer as the real / fake discrimination branch.
10. The unsupervised learning detection method for maritime targets based on an auxiliary generation network according to claim 1, characterized in that, In step 5, repeat steps 5-1 to 5-3 several times to complete the training: Step 5-1: Use the pseudo-detection box generation module to obtain e pseudo-detection box encodings, and then input the obtained pseudo-detection box encodings into the image generation module respectively to obtain the corresponding e simulated images; Step 5-2: Mix e real images and e simulated images, and input them into the target detection and discrimination module respectively, and obtain the corresponding image authenticity discrimination results from the authenticity discrimination branch , and then according to the image authenticity discrimination results and the true labels of the input images calculate the discrimination loss function corresponding to the image; update the network parameters of the feature learning network, the authenticity discrimination branch, and the image generation module based on the average value of the e discrimination loss functions; Discriminant loss function is calculated as follows: ; In the formula, the true / false discrimination result is the probability that the image is a real image; the true label is 1 or 0. When the image is a real image , otherwise ; Step 5-3: Input the e simulation images into the target detection discrimination module respectively, obtain the predicted detection box coordinate information from the target detection branch, then calculate the target detection loss function based on the predicted detection box coordinate information and the pseudo-detection box coordinate information in the pseudo-detection box encoding corresponding to the original simulation image, and update the network parameters of the target detection branch and the image generation module based on the average value of the e target detection loss functions.
Citation Information
Patent Citations
SAR image ship target detection method and device
CN118397256B