Target shape self-adaptive seabed shipwreck detection method based on side-scan sonar image
By using dynamic rotation convolution, feature decoupling head and S-A tag allocation strategies in the S3DR-Det model, the inconsistency problem in the detection of the submarine wreck target in the side-sweep sonar image is solved, and efficient and accurate detection effects are achieved.
Patent Information
- Application Number
- CN202510067675.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-16
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-01-16
AI Technical Summary
The prior art is difficult to accurately detect the high-level-width ratio subsea wreck targets placed in any direction in the side-sweep sonar image, and the rotation target detection has problems such as inconsistent targets with anchor frames, inconsistent classification features and regression features, and inconsistent rotation frame quality and label allocation strategy.
The S3DR-Det model is proposed, and high-quality rotation features are extracted through dynamic rotation convolution (DRC), and the feature decoupling head (FDM) inputs the rotation change characteristics and rotation invariant characteristics into the classification and regression task branches respectively. The S-A label allocation strategy is used in the training strategy, and the label allocation is integrated into IoU, distance between center points and angle differences.
It realizes efficient detection of multi-directional, high aspect ratio targets in the contralateral sweep sonar image, solves the inconsistency between the target and anchor frame, classification features and regression features, and rotates frame quality and label allocation strategy, and improves detection accuracy and model performance.
Smart Images

Figure CN120014424A_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the research field of seabed target detection in side-scan sonar images, and in particular relates to a seabed shipwreck detection method with adaptive target shape. Background Art
[0002] The detection and identification of seabed targets plays an extremely important role in underwater search and rescue, marine engineering construction, marine topography and geomorphology measurement, marine resource investigation and other fields. However, due to the complex marine environment, imaging conditions and measurement methods, its detection is more difficult than natural image target detection, and the detection accuracy is difficult to meet the needs. It has become a hot spot and difficulty in current research. Acoustic detection uses the acoustic image formed by the echo information from the target to detect seabed targets through manual interpretation of the image. Acoustic detection is widely used in underwater target detection due to its mature technology, intuitiveness, high efficiency and easy use. As the main equipment for seabed topography and geomorphology detection, side scan sonar has become the mainstream equipment for underwater target detection due to its high imaging resolution.
[0003] Currently, with the development of deep convolutional neural networks (DCNNs), many structures carefully designed for the characteristics of sonar images have achieved remarkable results. Most object detection methods mainly detect objects through horizontal bounding boxes (HBBs). The most important feature of horizontal bounding boxes is that their sides are parallel to the horizontal and vertical axes of the image. Among them, the two-stage detection model R-CNN based on horizontal boxes first generates a region proposal network for candidate object boxes to preliminarily screen out candidate regions that may contain objects; then these candidate boxes are classified and bounding box regression networks are used to further determine the category and precise location of the object in the candidate box. Single-stage detection models such as SSD and YOLO series simplify the object detection task into one stage, without using the intermediate candidate region generation step, and directly predict the target category and bounding box position on the input image or feature map. The above methods all perform object detection based on horizontal boxes. However, the target entities in side-scan sonar images (such as the most common shipwreck targets) are usually placed in arbitrary directions and have high aspect ratios. Using horizontal box detection cannot accurately represent targets in arbitrary directions, and will introduce a lot of background information, which brings great challenges to the detection algorithm to accurately locate oriented objects.
[0004] Rotated object detection is a very challenging task, which is more difficult and complex than traditional object detection, mainly reflected in the following three aspects:
[0005] (1) Inconsistency between target and anchor box
[0006] The convolutional features of current networks are usually axis-aligned and have a fixed receptive field. However, when faced with objects distributed in any direction in sonar images, there is a misalignment between the anchor frame and the convolutional features, making it difficult to accurately characterize the characteristics of the object. In other words, the anchor frames obtained by existing methods are of low quality and cannot cover the objects, resulting in inconsistencies between the objects and the anchor frames, and the features of the internal area of the anchor frame are difficult to represent the entire object. This phenomenon is more significant for objects with high aspect ratios. For example, the aspect ratio of underwater shipwreck targets is usually between 1 / 3 and 1 / 10. This misalignment phenomenon will aggravate the imbalance between the target and background information and hinder performance.
[0007] (2) Classification features and regression features are inconsistent
[0008] In the submarine target detection model, classification and regression tasks are performed based on features extracted from the backbone network, and these features are usually rotation invariant. However, in the sonar submarine target detection task, the target is distributed in any direction. In the classification task, we need to use fixed features to determine the category of the target, that is, rotation invariant features. Since the target in the side scan sonar image has the characteristics of rotating in multiple directions, it is difficult for us to obtain accurate target position information. Therefore, as the angle changes, we need to extract features at different angles to perceive the position change of the target, so as to accurately locate it, that is, rotation change features.
[0009] (3) Inconsistency between rotation box quality and label assignment strategy
[0010] For directional targets with high aspect ratios, IoU is very sensitive to changes in angles. Slight changes in angles can cause IoU to change dramatically. Moreover, a high IoU does not necessarily mean a good classification effect. Due to objects with high aspect ratios, it is difficult to accurately frame various features of the target when presetting anchor frames. Although some high IoU frames summarize the main position information of the target and may have better results in regression, they lack key features for classification, resulting in poor results. They are low-quality samples but are retained. Some low-IoU frames may frame key features and key position nodes to make them work well. Such high-quality samples are treated as negative samples. Therefore, the existing label assignment method only distinguishes positive and negative samples based on IoU scores, which will lead to an imbalance between positive and negative samples, thereby affecting model performance. Summary of the invention
[0011] The present invention proposes S 3The DR-Det model solves the inconsistency problem in rotation target detection from three levels. First, in the feature extraction stage, we designed a dynamic rotation convolution, which can extract high-quality rotation features based on the target's directional information. Next, due to the inconsistency of the features required for classification and regression tasks, we designed a feature decoupling head to input the rotation-changing features and rotation-invariant features into different task branches, respectively, to make classification and regression more accurate. Finally, we proposed the SA label assignment strategy in the training strategy, introduced the concept of alignment, and integrated information such as IoU, distance between center points, and angle difference to more comprehensively evaluate the sample quality for label assignment. The three modules are efficiently coupled together to ultimately achieve efficient and accurate detection.
[0012] In order to achieve the above object, the technical solution of the present invention is:
[0013] A method for detecting a sunken ship based on a target shape adaptively based on side scan sonar images, comprising the following steps:
[0014] The first step is data set preprocessing
[0015] The images in the dataset are grayed out and the dataset is divided into training set, validation set and test set.
[0016] The second step is to build and train the network model
[0017] The network model includes Backbone, Neck and Head. The network model is trained and verified using the training set and the verification set to obtain a trained network model.
[0018] (1) Backbone
[0019] Backbone extracts image features through a series of convolutional layers and activation functions, gradually reducing the spatial dimension of the image while increasing the number of channels. The standard convolution in the ResNet backbone is replaced by the DRC module. The specific processing process of the DRC module is: first input the feature map into the deep convolution, then perform layer normalization and ReLU activation, and then merge the activated features through average pooling and maximum pooling to obtain rich features. The merged feature vector is passed through a linear layer and different activation functions to obtain the predicted rotation angle α=[α1,...,α n ] and weights ω=[ω1,...,ω n ] Each convolution kernel is rotated according to the rotation angle and weight size, the rotated convolution kernel is convolved with the feature map, and the output features are added pixel by pixel to obtain the rotation feature.
[0020] (2) Neck
[0021] The neck part uses the feature pyramid network FPN (Feature Pyramid Networks), which is located between the backbone and the head. The FPN integrates the rotation features extracted by the backbone accordingly.
[0022] Furthermore, FPN constructs bottom-up and top-down feature fusion paths to fuse feature maps of different scales, thereby generating a feature pyramid with rich multi-scale information. Each layer of the feature pyramid corresponds to a specific scale range, allowing the model to process objects of different sizes at the same time.
[0023] (3)Head
[0024] Head generates the final detection result from the feature map provided by Neck, namely the category, bounding box information and confidence of the target. Head adopts the adaptive feature decoupling head structure FDM, which generates rotational variation features and rotational invariant features through the fused features and then inputs them into the regression sub-network and classification sub-network to generate the final prediction, thus improving the prediction accuracy of the model.
[0025] Furthermore, the adaptive feature decoupling head structure includes an anchor frame optimization module and a dynamic refinement module.
[0026] The anchor box optimization module includes anchor box regression and rotational convolution feature alignment operations; anchor box regression is to optimize the horizontal anchor box to an anchor box that is close to the target shape and has a certain rotation angle; the rotational convolution feature alignment operation is to dynamically and adaptively align the target features according to the shape, size and direction of the corresponding anchor box.
[0027] The dynamic refinement module adds a dynamic rotary encoder DRE before the classification and regression subnetwork and the classification subnetwork. The dynamic rotary encoder encodes the directional information to generate a feature map with multi-directional channels. The dynamic rotary encoder is a k×k×N filter that can actively rotate N-1 times during the convolution process to generate a feature map with N directional channels. For a feature map A and a DRE, the output S of the i-th direction is expressed as:
[0028]
[0029] In the formula, α i represents the angle of filter rotation, n represents the nth direction channel, represents the nth dynamic rotary encoder, and A represents the nth feature map.
[0030] Furthermore, during the network model training process, an alignment dynamic label assignment strategy based on spatial matching prior information is adopted. The alignment dynamic label assignment strategy uses alignment as an indicator to measure the quality of the anchor box, which is defined as follows:
[0031]
[0032] Among them, IoU pre is the rotation IoU value before regression, IoU post is the regressed rotation IoU value, ad is the alignment, d is the distance between the center point of the rotation prediction box and the center point of the rotation real box, θ is the angle difference between the rotation prediction box and the rotation real box, max d with max θ They represent the maximum possible distance and the maximum angle difference respectively. α, β, and γ in the formula are weighted hyperparameters used to measure the influence of different items.
[0033] In the training phase, we first calculate the ad of the real box GT and the predicted box, and then select the anchor boxes whose ad values are greater than or equal to a certain threshold as positive samples, and those below the threshold as negative samples. For the GT box that does not match any anchor box, the anchor box with the highest ad score is used as the positive candidate box for compensation, thereby realizing the dynamic selection of label assignment.
[0034] Step 3: Model Evaluation
[0035] By putting the test set into the trained model for detection, the target information in the side-scan sonar image is obtained, and the model performance is evaluated by the precision P (precision), recall R (recall), and mean average precision (mean average precision) indicators.
[0036] The beneficial effects of the present invention are as follows: the present invention proposes an S-type detection method for sunken ship targets in side-scan sonar images. 3 DR-Det is used to detect multi-directional, high-aspect-ratio targets in side-scan sonar images. Through the DRC, FDM, and SA label assignment strategies we proposed, the inconsistency between the target and the anchor box, the inconsistency between the classification feature and the regression feature, and the inconsistency between the quality of the rotation box and the existing label assignment strategy are solved in the feature extraction stage, detection stage, and training stage of the model. Each module in the model in this paper is highly coupled in function and structure, solving the problems caused by high-aspect-ratio targets in any direction at different stages of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] Figure 1 This is the overall network structure diagram;
[0038] Figure 2This is a schematic diagram of the principle of the dynamic rotation convolution DRC module;
[0039] Figure 3 This is a schematic diagram of the feature decoupling detection head structure;
[0040] Figure 4 Some of the test results are shown below. DETAILED DESCRIPTION
[0041] In order to make the method problems solved by the present invention, the method solutions adopted and the method effects achieved clearer, the present invention is further described in detail below in conjunction with the accompanying drawings and experiments. It is understood that the specific experiments described herein are only used to explain the present invention, rather than to limit the present invention. It should also be noted that, for the convenience of description, only the parts related to the present invention are shown in the accompanying drawings, rather than all the contents.
[0042] The overall structure of the model of the present invention is as follows Figure 1 As shown in the figure, it includes three parts: Backbone, Neck and Head. In the Backbone part, we designed a DRC convolution module, in which the dynamic rotation convolution kernel can be dynamically rotated according to different input feature maps to extract the rotation features of targets in any direction, thereby improving the model's ability to represent targets in different directions. We replace the standard convolution in the ResNet backbone with DRC as the basis, and integrate the dynamic rotation convolution DRC module for construction. The specific principle of the DRC module is as follows: Figure 2 As can be seen, we first input the feature map into the deep convolution, then perform layer normalization and ReLU activation, and then merge the activated features through average pooling and maximum pooling to obtain rich features. The merged feature vector is passed through the linear layer and different activation functions to obtain the predicted rotation angle α=[α1,...,α n ] and weights ω=[ω1,...,ω n ]. Each convolution kernel is rotated according to the rotation angle and weight size, the rotated convolution kernel is convolved with the feature map, and the output features are added pixel by pixel.
[0043] In order to solve the problem of inconsistent features required by the classification and regression task branches, we designed an adaptive feature decoupling head structure FDM. The specific structure is as follows: Figure 3As shown in the figure, the adaptive feature decoupling head structure includes an anchor box optimization module and a dynamic refinement module. The dynamic refinement module is to add a dynamic rotary encoder before the classification and regression task branches. First, the anchor box optimization module generates high-quality anchor boxes, and adaptively aligns the features according to the corresponding anchor boxes and dynamic rotation convolutions, and inputs the features corresponding to the classification and regression tasks into different task branches. The anchor box optimization module (AFO) contains an anchor box regression and rotation convolution feature alignment operation. The anchor box regression optimizes the horizontal anchor box into a high-quality anchor box with a certain rotation angle that is close to the target shape. The rotation convolution feature alignment operation dynamically and adaptively aligns the target features according to the shape, size and direction of the corresponding anchor box. Then, the dynamic rotary encoder (DRE) in the dynamic refinement module is used to encode the direction information to generate a feature map with multi-directional channels. DRE is a k×k×N filter that can actively rotate N-1 times during the convolution process to generate a feature map with N (default is 8) directional channels. For a feature map A and a DRE (represented by E()), the output S of the i-th direction can be expressed as
[0044]
[0045] In the formula, α i Represents the angle of filter rotation, and n represents the nth directional channel. By using DRE on the convolutional layer, we can obtain features sensitive to rotation changes based on directional information encoding. Then, we can select the directional channel with the strongest response as the input feature of the classification task.
[0046] During the overall model training process, we proposed an alignment dynamic label assignment strategy SA based on spatial matching prior information (SMPI). It mainly refers to the detection model distinguishing positive and negative samples during the training phase, and matching appropriate supervision targets to different positions of the feature map for calculating losses and completing gradient updates. We comprehensively consider three factors: IoU, center point distance difference, and angle difference, and design the SA label assignment strategy based on these three points as the basis for distinguishing positive and negative samples, more comprehensively evaluating the quality of anchor boxes, and effectively improving the performance of the model. We introduce alignment as an indicator to measure the quality of anchor boxes, which is defined as follows:
[0047]
[0048] Among them, IoU pre is the rotation IoU value before regression, IoU post is the regressed rotation IoU value, ad is the alignment, d is the distance between the center point of the rotation prediction box and the center point of the rotation real box, θ is the angle difference between the rotation prediction box and the rotation real box, max d with max θThey represent the maximum possible distance and the maximum angle difference respectively. α, β, and γ in the formula are weighted hyperparameters used to measure the influence of different items.
[0049] During the regression process, effective interference suppression can collect higher quality anchor frames and make training more stable. Our penalty term for regression uncertainty consists of three parts: IoU, distance between center points, and angle difference. We can effectively select high-quality and highly aligned rotation anchor frames from three different aspects.
[0050] In the training phase, we first calculate the ad of the ground truth box GT and the predicted box, and then select those anchor boxes whose ad values are greater than or equal to a certain threshold as positive samples, and those below the threshold as negative samples. For the GT box that does not match any anchor box, we use the anchor box with the highest ad score as a positive candidate box for compensation, thereby achieving dynamic selection of label assignment.
[0051] The first step is data set preprocessing
[0052] We selected a shipwreck target dataset with high aspect ratio and multi-directional rotation for experimental verification. The experimental datasets were collected by domestic research institutes and manufacturers using mainstream side-scan sonar equipment in different sea areas and collected on the Internet. The dataset contains a total of 1,691 shipwreck sample images, and the targets in the dataset are placed in any direction and have the characteristics of high aspect ratio, which is suitable for verifying the effectiveness of the dynamic rotating target detection model. We divided the entire dataset into training set, validation set and test set in a ratio of 5:2:3.
[0053] Since the training samples come from a wide range of sources, in order to ensure the generalization of the model, we grayscale all samples and then input them into the network for training. During the training process, only random horizontal flipping is used to avoid overfitting, and a small learning rate is used to avoid drastic changes in the rotation angle. The optimizer used for training is SGD, the initial learning rate is set to 0.0025, the momentum is set to 0.9, and the weight decay is set to 0.0001 to avoid overfitting or underfitting. The training includes 500 warm-up iterations before starting training, and no pre-trained weights are used during the training process, and training starts from scratch. This model is implemented based on Python, using the PyTorch deep learning framework, the operating system is Windows 11, and the hardware used for the experiment includes Intel Core i7-14650HX CPU, NVIDIA GeForce RTX 4060Laptop GPU and 64GB memory.
[0054] Step 2: Model training
[0055] The model training process is carried out according to the technical solution.
[0056] Step 3: Model Evaluation
[0057] In order to comprehensively and objectively evaluate the prediction effects of different models, the present invention evaluates and optimizes model performance through the following coefficients: average precision AP (average precision) and mean AP (mean AP).
[0058] Precision and recall: In the classification task of predicting whether an image contains a bag, the four elements of precision and recall can be explained as follows: TP (true positive): the positive sample is correctly marked as a positive sample in the prediction result; TN (true negative): the negative sample is correctly marked as a negative sample in the prediction result; FP (false positive): the positive sample is incorrectly marked as a negative sample in the prediction result; FN (false negative): the negative sample is incorrectly marked as a positive sample in the prediction result. The calculation relationship is:
[0059]
[0060]
[0061] Average Precision: The geometric meaning of average precision AP is the area corresponding to the PR curve as shown in formula (5), where the interpolation summation method can be used to approximate the integral.
[0062]
[0063] We conducted comparative experiments with existing rotated object detection methods, including a two-stage detection model and a single-stage detection model, and the results are shown in Table 1. From the results in the table, our model achieved an AP result of 89.68%, which is better than all the two-stage and single-stage detection models in the table, and obtained higher Recall values and AP values, which shows the superiority of the algorithm in the detection of rotated objects.
[0064] Table 1 Comparison of experimental results with different models
[0065]
[0066]
[0067] In order to verify the effects of DRC, FDM and SA designed in the present invention, we conducted ablation experiments and used different module combinations to verify their effects on the model. The results are shown in Table 2.
[0068] Table 2 Ablation experiments of modules in the model
[0069]
[0070] We will configure the proposed DRC, FDM and SA in the baseline model to conduct ablation experiments to verify the effect of each part on the model, and conduct experiments on the dataset. From the results, we can find that each module has a significant improvement when configured separately in the baseline model. Compared with the baseline model using static convolution, the addition of the DRC module enables the rotation convolution kernel to dynamically align the angle, indicating the adaptability and effectiveness of the DRC module in capturing rotated targets. The FDM module can effectively decouple the extracted rotation features and input the corresponding features into the feature branch network to achieve the effect of fine positioning and classification. Therefore, adding the FDM module will effectively improve the model detection accuracy and achieve a higher AP value. Finally, SA introduces an alignment that is more in line with the rotation box to measure the quality of the anchor box. The input IoU, center point distance difference and angle difference are introduced in the alignment to more comprehensively and realistically select high-quality anchor boxes to improve model accuracy.
[0071] Some representative test results are as follows: Figure 4 As shown in the figure, it can be seen that although the targets are distributed in any direction in the image, the detection model can still accurately identify the targets and accurately select them according to the target direction, which effectively improves the positioning and recognition accuracy of the model. However, we can also find that for some targets, the detection frame is slightly offset. This is because some shipwreck targets have sunk to the bottom of the sea for a long time and some parts are buried, making it difficult to distinguish the edge contour of the shipwreck from the background of the seabed, resulting in positioning deviation. At the same time, due to observation conditions, instruments and other factors, the target imaging quality is poor, and affected by suspended matter in the water, the effective echo of the target is covered, making it difficult to accurately identify the target.
[0072] Finally, it should be noted that the above experiments are only used to illustrate the method scheme of the present invention, rather than to limit it. Although the present invention has been described in detail, ordinary method personnel in the field should understand that modifying the aforementioned method scheme, or equivalently replacing part or all of the method features therein, does not cause the essence of the corresponding method scheme to deviate from the scope of the method scheme of the present invention.
Claims
1. A method for detecting submarine sunken ships based on target shape adaptively based on side scan sonar images, characterized in that: The specific steps include: The first step is data set preprocessing Grayscale the images in the dataset and divide the dataset into training set, validation set and test set; The second step is to build and train the network model The network model includes Backbone, Neck and Head. The network model is trained and verified using the training set and the verification set to obtain a trained network model. (1) Backbone Backbone extracts image features through a series of convolutional layers and activation functions, gradually reducing the spatial dimension of the image while increasing the number of channels; the standard convolution in the ResNet backbone is replaced by the DRC module; The specific processing process of the DRC module is as follows: first, the feature map is input into the deep convolution, then layer normalization and ReLU activation are performed, and then the activated features are merged through average pooling and maximum pooling to obtain rich features. The merged feature vector is passed through the linear layer and different activation functions to obtain the predicted rotation angle α=[α1,...,α n ] and weights ω=[ω1,...,ω n ]; Rotate each convolution kernel according to the rotation angle and weight, convolve the rotated convolution kernel with the feature map, and add the output features pixel by pixel to obtain the rotation feature; (2) Neck The neck part uses the feature pyramid network FPN, which is located between the backbone and the head. The FPN integrates the rotation features extracted by the backbone accordingly. (3)Head Head generates the final detection result from the feature map provided by Neck, namely the category, bounding box information and confidence of the target; Head adopts the adaptive feature decoupling head structure FDM, The adaptive feature decoupling head structure FDM uses the fused features to generate rotational variation features and rotational invariant features, which are then input into the regression sub-network and the classification sub-network to generate the final prediction, thereby improving the prediction accuracy of the model. The adaptive feature decoupling head structure includes an anchor frame optimization module and a dynamic refinement module; The anchor box optimization module includes anchor box regression and rotation convolution feature alignment operations; anchor box regression is to optimize the horizontal anchor box to an anchor box that is close to the target shape and has a certain rotation angle; The rotation convolution feature alignment operation dynamically and adaptively aligns the target features according to the shape, size and orientation of the corresponding anchor box; The dynamic refinement module adds a dynamic rotary encoder DRE before the classification and regression subnetwork and the classification subnetwork. The dynamic rotary encoder encodes the direction information to generate a feature map with multi-directional channels. Step 3: Model Evaluation By putting the test set into the trained model for detection, the target information in the side-scan sonar image is obtained, and the model performance is evaluated by the precision rate P, recall rate R, and average precision mean indicators.
2. The method for detecting a sunken ship based on a target shape adaptively based on side scan sonar images according to claim 1, characterized in that: FPN fuses feature maps of different scales by constructing bottom-up and top-down feature fusion paths, thereby generating a feature pyramid with rich multi-scale information; each layer of the feature pyramid corresponds to a specific scale range, allowing the model to process objects of different sizes simultaneously.
3. According to the method for adaptively detecting submarine shipwrecks based on side scan sonar images of target shape according to claim 1, the dynamic rotary encoder is a k×k×N filter that can actively rotate N-1 times during the convolution process to generate a feature map with N directional channels. For a feature map A and a DRE, the output S in the i-th direction is expressed as: In the formula, α i represents the angle of filter rotation, n represents the nth direction channel, represents the nth dynamic rotary encoder, and A represents the nth feature map.
4. According to the target shape adaptive submarine shipwreck detection method based on side scan sonar images in claim 1, in the process of network model training, an alignment degree dynamic label allocation strategy based on spatial matching prior information is adopted. The alignment degree dynamic label allocation strategy uses alignment degree as an indicator to measure the quality of the anchor frame, which is defined as follows: in, IoU pre is the rotation IoU value before regression, IoU post is the regressed rotation IoU value, ad is the alignment, d is the distance between the center point of the rotation prediction box and the center point of the rotation real box, θ is the angle difference between the rotation prediction box and the rotation real box, max d with max θ Represent the maximum possible distance and the maximum angle difference respectively. α, β, and γ in the formula are weighted hyperparameters used to measure the influence between different items. During the training phase, we first calculate the ad of the real box GT and the predicted box, and then select the anchor boxes whose ad values are greater than or equal to a certain threshold as positive samples, and those below the threshold as negative samples; for the GT box that does not match any anchor box, the anchor box with the highest ad score is used as the positive candidate box for compensation, thereby realizing dynamic selection of label assignment.
Citation Information
Patent Citations
Ocean ship detection method and system based on Gaussian prior label distribution and feature decoupling
CN116823838A
Underwater sonar image target detection method and system based on neural network
CN118628898A
Infrared ship detection method based on improved RT-DETR algorithm
CN119169453A
Training method and training apparatus for rotating-ship target detection model, and storage medium
WO2023116631A1
Cited By
Image classification method and device, computer program product and image classification system
CN121236455A