Ship fine-grained classification method and system based on double-branch collaborative decoupling network

By using a dual-branch collaborative decoupling network for feature extraction and an asynchronous feature decoupling module, the problems of inter-class feature separation and intra-class feature consistency in fine-grained ship classification are solved, thereby improving the classification accuracy of remote sensing ship monitoring.

CN120997561APending Publication Date: 2025-11-21WUHAN UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510977909.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-16
Publication Date
2025-11-21

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively preserve the overall outline and microstructure of ships in fine-grained classification. Furthermore, traditional methods tend to obscure intra-class semantic relationships when separating features between classes, making it difficult to balance the dynamic requirements of feature decoupling and aggregation.

Method used

A method based on a dual-branch collaborative decoupling network is adopted. Through a feature extractor, a multi-semantic collaborative attention module, and an asynchronous feature decoupling module, inter-class feature separation and intra-class feature consistency are achieved. The feature response is enhanced by a multi-semantic space grouping module and a channel attention sparse module. The asynchronous feature decoupling module improves classification accuracy through biaxial pooling and feature aggregation.

Benefits of technology

It improves the accuracy of fine-grained ship classification, making it particularly suitable for remote sensing ship monitoring in complex scenarios, and enhances the ability to identify inter-class similarities and intra-class differences.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120997561A_ABST
    Figure CN120997561A_ABST
Patent Text Reader

Abstract

The invention discloses a ship fine-grained classification method based on a double-branch collaborative decoupling network. The method comprises the following steps: inputting a ship image into a ship fine-grained classification network to obtain a ship classification result; the ship fine-grained classification network comprises a feature extractor, a multi-semantic collaborative attention module and an asynchronous feature decoupling module, and the processing process of the ship fine-grained classification network comprises the following steps: inputting a ship image after data enhancement into the feature extractor for feature extraction to obtain a preliminary double-view feature map; inputting the initial double-view feature map into a multi-semantic collaborative attention module for spatial coding and channel self-attention to obtain an enhanced feature map; and inputting the enhanced feature map into an asynchronous feature decoupling module for feature separation and feature aggregation, and obtaining a ship fine-grained classification result through a full connection layer. According to the method, the challenges of complex background and easy confusion of appearance in a ship fine-grained classification task are well handled, and the precision of ship fine-grained classification is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the fields of computer vision and remote sensing image processing technology, and in particular to a fine-grained classification method, system, storage medium and electronic device for ships based on a dual-branch cooperative decoupling network. Background Technology

[0002] With the acceleration of globalization and the rapid development of the marine economy, ship monitoring is becoming increasingly important in areas such as marine resource development, maritime traffic management, environmental protection, and national defense. Remote sensing technology, with its wide coverage, real-time capabilities, and multi-source data fusion capabilities, has become a core tool for ship monitoring. Among these, fine-grained ship classification requires precise identification of ship types, functions, and subcategories, such as distinguishing between cargo ships, fishing vessels, and warships. This is crucial for improving marine management efficiency, combating illegal fishing, and ensuring maritime security. However, achieving high-precision fine-grained classification remains a significant challenge due to the strong coupling between ship targets and complex backgrounds (such as waves, clouds, and port facilities) in remote sensing imagery, as well as the similarities in appearance between ship subcategories and the differences in appearance within subcategories.

[0003] To address complex background problems, existing deep learning-based fine-grained ship classification methods primarily employ two strategies: one is a multi-stage local feature fusion scheme, which progressively focuses on key hull components through cascaded attention modules; the other is a channel attention weighting mechanism, which uses channel-dimensional feature response values ​​to filter discriminative regions. However, both of these mainstream methods suffer from significant multi-scale representation deficiencies—the former, using serially stacked convolutional modules for feature extraction, suffers from the loss of fine-grained details as the receptive field expands due to layer-by-layer downsampling; the latter relies on single-scale global average pooling for feature compression, making it difficult to simultaneously preserve the overall ship outline and microstructure. Furthermore, regarding the unique problem of "high inter-class similarity and significant intra-class differences" in fine-grained ship classification, traditional methods use Softmax cross-entropy loss for end-to-end training, essentially only constructing linear decision boundaries between categories and failing to fully exploit the discriminative feature relationships between samples. In recent years, contrastive learning techniques have shown potential in mitigating appearance confusion by uncovering discriminative relationships between samples. However, existing methods often employ synchronous optimization strategies, which can easily blur semantic relationships within classes while forcibly separating features between classes, making it difficult to balance the dynamic requirements of feature decoupling and aggregation. Summary of the Invention

[0004] This application provides a ship fine-grained classification method, system, storage medium, and electronic device based on a dual-branch collaborative decoupling network, which can ensure sufficient separation of features between classes and semantic consistency of features within classes, further improving the performance of ship fine-grained classification.

[0005] This application provides a fine-grained ship classification method based on a dual-branch cooperative decoupling network, including: Acquire images of the ship; Ship images are input into a fine-grained ship classification network to obtain ship classification results; The ship fine-grained classification network includes a feature extractor, a multi-semantic collaborative attention module, and an asynchronous feature decoupling module. The processing steps of the ship fine-grained classification network include: The augmented ship image is input into a feature extractor to extract features and obtain a preliminary dual-view feature map. The initial dual-view feature map is input into the multi-semantic collaborative attention module for spatial encoding and channel self-attention to obtain the enhanced feature map; The enhanced feature map is input into the asynchronous feature decoupling module for feature separation and feature aggregation, and then the fine-grained classification result of the ship is obtained through a fully connected layer.

[0006] Furthermore, according to the above-mentioned ship fine-grained classification method based on a dual-branch cooperative decoupling network, the multi-semantic cooperative attention module includes a multi-semantic spatial grouping module and a channel attention sparse module. The step of inputting the preliminary dual-view feature map into the multi-semantic cooperative attention module for spatial encoding and channel self-attention to obtain an enhanced feature map includes: The initial dual-view feature map is input into the multi-semantic space grouping module for global pooling, multi-scale feature extraction and stitching operations to obtain the first feature map. The first feature map is input into the channel attention sparse module, and the enhanced feature map is obtained based on the adaptive pooling compression and channel self-attention joint computation framework.

[0007] Furthermore, according to the above-mentioned ship fine-grained classification method based on a dual-branch cooperative decoupling network, the step of inputting the preliminary dual-view feature map into the multi-semantic space grouping module for global pooling, multi-scale feature extraction, and concatenation operations to obtain the first feature map includes: The initial dual-view feature map is subjected to biaxial average pooling along the X and Y directions to obtain two one-dimensional sequence structures. The channel dimension of the one-dimensional sequence structure is divided into four independent sub-features. Apply depthwise one-dimensional convolutions with different kernel sizes to each group of sub-features to obtain the modified sub-features; The modified sub-features in a certain direction are concatenated, and spatial attention weights are generated using group normalization and the Sigmoid function. The first feature map is generated based on spatial attention weights and preliminary dual-view feature maps.

[0008] Furthermore, according to the above-mentioned ship fine-grained classification method based on a dual-branch collaborative decoupling network, the step of inputting the first feature map into the channel attention sparse module and obtaining the enhanced feature map based on the adaptive pooling compression and channel self-attention joint computation framework includes: The first feature map is compressed using adaptive average pooling to obtain a compressed feature map; Generate a query matrix based on the compressed feature map. Key matrix Sum matrix Then according to Calculate the attention distribution matrix; The top 50% of channels are selected, and the channels are sparsified to obtain enhanced feature maps.

[0009] Furthermore, according to the above-mentioned ship fine-grained classification method based on a dual-branch cooperative decoupling network, the asynchronous feature decoupling module includes a feature separation module, a feature aggregation module, and a fully connected layer. The step of inputting the enhanced feature map into the asynchronous feature decoupling module for feature separation and aggregation, and then obtaining the ship fine-grained classification result through the fully connected layer, includes: The enhanced features are pooled, an orthogonal projection matrix is ​​constructed to perform a linear transformation on the feature space, sample pairs of the same image are aggregated, and sample pairs of different images are separated. Reconstruct feature associations by bringing images of ships of a certain category closer to the agents belonging to that category to obtain the aggregation result; The aggregation results are input into a fully connected layer to obtain fine-grained classification results for ships.

[0010] Furthermore, according to the above-mentioned ship fine-grained classification method based on a dual-branch cooperative decoupling network, the feature separation module includes a dual-channel pooling layer and a separation block, performs pooling operations on the enhanced features, constructs an orthogonal projection matrix to perform linear transformation on the feature space, aggregates sample pairs of the same image, and separates sample pairs of different images, including: Global average pooling and local max pooling operations are performed on the enhanced features to obtain hull morphology and structural detail feature maps. The hull shape and structural detail feature maps are mapped to the decoupled space to generate feature vectors; The feature vectors are mirrored and flipped along the batch dimension to construct the feature matrix; By constraining the feature maps of different views of the same image to be close to each other, and separating the views of different images from each other, we obtain separated feature maps.

[0011] Furthermore, according to the above-mentioned ship fine-grained classification method based on a dual-branch cooperative decoupling network, the feature aggregation module includes an aggregation block and a multi-agent module. The reconstruction of feature associations, which brings ship images of a certain category closer to the agent to which that category belongs, to obtain the aggregation result, includes: The separated feature maps are subjected to convolution, normalization, and activation operations to obtain the aggregated block output feature map; The feature map output from the aggregated block is input into the multi-agent module, which constrains ship images of a certain category to be close to the agent to which that category belongs, thus obtaining a separation feature map.

[0012] Furthermore, according to the above-mentioned ship fine-grained classification method based on a dual-branch cooperative decoupling network, during the training process, the feature maps of different views of the same image are constrained to be close to each other, while the views of different images are separated from each other, by using a feature separation loss function. The feature separation loss function is:

[0013] in, For feature separation loss function, For batches, For mirror separation loss function, The aggregation loss function; The feature separation loss function is:

[0014]

[0015]

[0016] in, For feature separation loss function, For the feature map after separation, For another feature map that is far away, To block gradient backpropagation for the operator, These represent the degree of separation and the degree of aggregation between two feature maps a and b, respectively. The aggregation loss function is: .

[0017] Furthermore, according to the above-mentioned fine-grained ship classification method based on a dual-branch cooperative decoupling network, during the training process, a feature aggregation loss function is used to constrain ship images of a certain category to be close to the agent to which that category belongs. The feature aggregation loss function is:

[0018] in, For feature aggregation loss function, The loss is calculated based on the feature map's proximity to the subclass proxy. To avoid the losses caused by non-subclass proxies, The losses incurred due to the need to maintain differentiation among the various agents; The loss for feature maps that are closer to the subclass proxy is:

[0019] The loss for feature maps that are closer to the subclass proxy is:

[0020] The losses incurred by maintaining differentiation among agents are:

[0021] in, The feature map of the aggregated block shows the category labels as follows: , Represents a set ; The total loss function used to train the fine-grained ship classification network is:

[0022]

[0023] in, This represents the final model inference category result. These are the actual labels of the input image. It is the cross-entropy loss function.

[0024] This application also provides a fine-grained ship classification system based on a dual-branch cooperative decoupling network, including: The acquisition module is used to acquire images of the ship. The ship classification module is used to input ship images into a fine-grained ship classification network to obtain ship classification results; The ship fine-grained classification network includes a feature extractor, a multi-semantic collaborative attention module, and an asynchronous feature decoupling module. The processing steps of the ship fine-grained classification network include: The augmented ship image is input into a feature extractor to extract features and obtain a preliminary dual-view feature map. The initial dual-view feature map is input into the multi-semantic collaborative attention module for spatial encoding and channel self-attention to obtain the enhanced feature map; The enhanced feature map is input into the asynchronous feature decoupling module for feature separation and feature aggregation, and then the fine-grained classification result of the ship is obtained through a fully connected layer.

[0025] This application also provides a computer-readable storage medium storing multiple instructions adapted for loading by a processor to execute any of the above-described ship fine-grained classification methods based on a dual-branch cooperative decoupling network.

[0026] This application also provides an electronic device, including a processor and a memory, wherein the processor is electrically connected to the memory, the memory is used to store instructions and data, and the processor is used in the steps of the ship fine-grained classification method based on a dual-branch cooperative decoupling network described in any of the above claims.

[0027] This application provides a ship fine-grained classification method, system, storage medium, and electronic device based on a dual-branch collaborative decoupling network. The ship fine-grained classification network includes: a feature extractor for initial feature extraction from input ship images; a multi-semantic collaborative attention module, which first extracts multi-semantic spatial information through a multi-semantic spatial grouping module and provides spatial priors for channel attention calculation, then refines the semantic understanding of local sub-features through a channel sparse attention module, mitigating semantic differences in the multi-semantic spatial grouping module; the two work collaboratively to enhance the feature response of key ship parts; and an asynchronous feature decoupling module, which achieves full decoupling of ship fine-grained features through a dual-axis pooling strategy, separation blocks, aggregation blocks, multiple proxy modules, and a loss function constraint that integrates feature decoupling and classification accuracy. This method, by designing a spatial-channel attention collaborative optimization and feature decoupling architecture, effectively addresses the challenges of complex backgrounds and easily confused appearances in ship fine-grained classification tasks, improving the accuracy of ship fine-grained classification, and is particularly suitable for remote sensing ship fine-grained monitoring and classification tasks in complex scenarios. Attached Figure Description

[0028] The technical solution and other beneficial effects of this application will become apparent from the following detailed description of specific embodiments in conjunction with the accompanying drawings.

[0029] Figure 1 A flowchart of a ship fine-grained classification method based on a dual-branch cooperative decoupling network provided in this application embodiment.

[0030] Figure 2 This is a schematic diagram of the structure of the multi-semantic collaborative attention module provided in an embodiment of this application.

[0031] Figure 3 This is a structural diagram of the asynchronous feature decoupling module provided in an embodiment of this application.

[0032] Figure 4 A visualization of the confusion matrix for the fine-grained classification results provided in this application.

[0033] Figure 5This is a schematic diagram of the structure of a ship fine-grained classification system based on a dual-branch cooperative decoupling network provided in an embodiment of this application.

[0034] Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0035] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0036] This application provides a ship fine-grained classification method, system, storage medium, and electronic device based on a two-branch cooperative decoupling network. The ship fine-grained classification system based on a two-branch cooperative decoupling network provided in this application can be integrated into an electronic device, such as a terminal or server. The terminal can include a tablet computer, laptop computer, personal computer (PC), microprocessor box, or other devices.

[0037] Please see Figure 1 , Figure 1 The flowchart of the ship fine-grained classification method based on a dual-branch cooperative decoupling network provided in this application embodiment is applied in electronic devices. The ship fine-grained classification method based on a dual-branch cooperative decoupling network includes the following steps: S1: Acquire ship images and perform data augmentation on the ship images.

[0038] S2, the augmented ship image is input into the ship fine-grained classification network to obtain the ship classification result.

[0039] The ship fine-grained classification network includes a feature extractor, a multi-semantic collaborative attention module, and an asynchronous feature decoupling module. The processing steps of the ship fine-grained classification network include: S21, The data-enhanced ship image is input into the feature extractor for feature extraction to obtain a preliminary dual-view feature map.

[0040] The feature extractor includes a ResNet50 backbone network for preliminary feature extraction from the input remotely sensed ship imagery.

[0041] Specifically, the feature extractor's processing steps include: S211, the input ship image undergoes a dual-view enhancement strategy, which involves two random image enhancement processes: random color jitter (adjusting brightness, contrast, saturation, and hue with a probability of 0.5), introducing a Gaussian blur kernel with a probability of 0.5, randomly performing horizontal or vertical flipping, and applying continuous rotations from 0° to 180°, resulting in two sets of enhanced views. .

[0042] S212, the two sets of enhanced views obtained in step S211 The input is fed into the ResNet50 backbone network to obtain the preliminary extracted feature maps. , ,in , This is the batch size of the input images. It is the number of channels in the feature map obtained after passing through ResNet50. and These are the height and width of the feature map, respectively.

[0043] S22, the initial dual-view feature map is input into the multi-semantic collaborative attention module for spatial encoding and channel self-attention to obtain the enhanced feature map.

[0044] Figure 2 This is a schematic diagram of the structure of the multi-semantic collaborative attention module provided in the embodiments of this application, such as... Figure 2 As shown, the multi-semantic collaborative attention module includes a multi-semantic space grouping module and a channel attention sparse module. Step S22 includes: S221, the preliminary dual-view feature map is input into the multi-semantic space grouping module for global pooling, multi-scale feature extraction and splicing operations to obtain the first feature map.

[0045] The multi-semantic space grouping module employs key components such as biaxial global pooling, channel grouping, and multi-scale deep one-dimensional convolutional clusters. First, the input feature map is processed by biaxial global pooling in the X and Y directions. Then, the pooling results in different directions are grouped along the channel dimension. Multi-scale feature extraction is then performed using multi-scale deep one-dimensional convolutional clusters, and the extracted results are concatenated along the X and Y directions. Finally, group normalization and sigmoid activation are applied, multiplying the result with the original input feature map to model the local texture and global morphological features of the long, narrow target of the ship, suppressing background noise such as ocean waves.

[0046] In one embodiment, step S221 includes: S2211, perform biaxial average pooling on the preliminary dual-view feature map along the X and Y directions to obtain two one-dimensional sequence structures.

[0047] The obtained preliminary dual-view feature map is subjected to biaxial average pooling along the X and Y directions to obtain a single-view feature map. For example, the operation steps for the other view are similar. The pooling result is two unidirectional one-dimensional sequence structures. and .

[0048] S2212 divides the channel dimension of the one-dimensional sequence structure into four independent sub-features.

[0049] Will and The channel dimension is divided into 4 independent sub-features, and the number of channels in each sub-feature is... The representation of each group's sub-features is shown in the formula:

[0050]

[0051] In the formula, Represents the i-th sub-feature. .

[0052] S2213, apply a depthwise one-dimensional convolution with different kernel sizes to each group of sub-features to obtain the modified sub-features.

[0053] For each group of sub-features, a depthwise one-dimensional convolution with a different kernel size is applied. Small kernels focus on local textures, while large kernels model the overall structure. This process can be expressed by the following formula:

[0054]

[0055] In the formula, This represents the i-th sub-feature after a one-dimensional depthwise convolution. This represents the size of the convolution kernel used for the i-th sub-feature. , This represents a one-dimensional convolution operation.

[0056] S2214 concatenates the modified sub-features in a certain direction and uses group normalization and the Sigmoid function to generate spatial attention weights.

[0057] Sub-features in a certain direction are concatenated, and spatial attention weights are generated using group normalization and the sigmoid function, expressed as the formula:

[0058]

[0059] In the formula, Represents the Sigmoid activation function. and These represent the group normalization of the four sub-features along the height and width dimensions, respectively. This represents the concatenation function.

[0060] S2215, generates the first feature map based on spatial attention weights and preliminary dual-view feature maps.

[0061] Combine the initial dual-view feature map with , Multiplying them yields the output features, as shown in the formula:

[0062] S222, the first feature map is input into the channel attention sparse module, and the enhanced feature map is obtained based on the adaptive pooling compression and channel self-attention joint calculation framework.

[0063] The channel sparse attention module adopts a joint computation framework of adaptive pooling compression and channel self-attention, including a 7×7 pooling dimensionality reduction layer, a Q, K, V matrix calculation layer generated by depthwise separable convolution, and a channel similarity matrix sparse filtering layer. It achieves cross-channel semantic association enhancement by retaining the top 50% of the weighted channels, which is used to correct spatial positioning bias and suppress intra-class redundant features.

[0064] In one embodiment, step S222 includes the following steps: S2221, The first feature map is compressed using adaptive average pooling to obtain a compressed feature map.

[0065] The first feature map is compressed using 7×7 adaptive average pooling, reducing the resolution of the feature map to [value missing]. The formula is expressed as:

[0066] In the formula, where This means using a convolution with a kernel size of 7×7 to reduce the size of the feature map from... Become The pooling function.

[0067] S2222, Generate a query matrix based on the compressed feature map. Key matrix Sum matrix Then according to Calculate the attention distribution matrix.

[0068] Semantic similarity between channels is modeled using channel self-attention. First, a query matrix is ​​generated based on the compressed feature map. Key matrix Sum matrix Then according to Calculate the attention distribution matrix The process can be expressed by the following formula:

[0069]

[0070]

[0071]

[0072] In the formula, This is a depthwise convolution operation with a kernel size of 1×1, unlike the generation of kernels for each spatial location in the Transformer. In CSA, each channel is generated independently. , .

[0073] S2223, select the top 50% of channels with weights, perform sparsification filtering on the channels, and obtain the enhanced feature map.

[0074] By selecting the top 50% of channels with the highest weights and performing channel sparsification, the computational cost is reduced while maintaining focus on important features, resulting in an enhanced feature map. The formula is expressed as:

[0075] In the formula, This is a channel filtering function that selects the top 50% of channels for activation. Indicates the kernel size as Convolution reduces the size of the feature map from Become Pooling functions, This represents the activation function.

[0076] The above steps operate on one view; the same logic applies to the other view, ultimately yielding the output feature maps of both views after passing through the multi-semantic collaborative attention module. , .

[0077] S23, the enhanced feature map is input into the asynchronous feature decoupling module for feature separation and feature aggregation, and then the fine-grained classification result of the ship is obtained through the fully connected layer.

[0078] The asynchronous feature decoupling module includes a dual-path pooling layer, a separation block, an aggregation block, a multi-proxy module, and a fully connected layer. It is used to decouple individual ship features and solve the problem of appearance confusion between ship classes with similar appearance features and within ship classes with different appearance features.

[0079] Figure 3This is a structural diagram of the asynchronous feature decoupling module provided in the embodiments of this application, as shown below. Figure 3 As shown, the asynchronous feature decoupling module includes a feature separation module, a feature aggregation module, and a fully connected layer. Step S23 includes the following steps: S231 performs pooling operations on the enhanced features, constructs an orthogonal projection matrix to perform linear transformation on the feature space, aggregates sample pairs of the same image, and separates sample pairs of different images.

[0080] The feature aggregation module includes a dual-pooling layer and a splitting block. The dual-pooling layer extracts different features from the feature map, and the splitting block constructs an orthogonal projection matrix to perform a linear transformation on the feature space, thus aggregating sample pairs from the same image and separating sample pairs from different images.

[0081] In one embodiment, the feature separation module includes a dual-channel pooling layer and a separation block, and step S231 includes: S2311 performs global average pooling and local max pooling operations on the enhanced features to obtain hull morphology and structural detail feature maps.

[0082] The dual-pooling layer captures hull morphology and focuses on structural details by performing global average pooling and local max pooling operations on the enhanced features, respectively. The formula is expressed as follows:

[0083]

[0084] S2312 maps the hull shape and structural detail feature maps to the decoupled space to generate feature vectors.

[0085] The separation block includes two sets of separation layers. Each separation layer consists of one linear layer, one normalization layer, and one ReLU activation layer, used to map the output feature map of step S2311 to the decoupling space and generate feature vectors. .

[0086] S2313 performs a mirror flip operation on the feature vectors along the batch dimension to construct the feature matrix.

[0087] The feature matrix is ​​constructed by mirroring and flipping along the batch dimension. To ensure any position eigenvectors and Originating from different samples, only batch size needs to be guaranteed. If it is even, then and The j-th vector and It must be a feature obtained by transforming two different images.

[0088] S2314, constrain the feature maps of different views of the same image to be close to each other, and separate the views of different images from each other, to obtain separated feature maps.

[0089] S232, Reconstruct feature associations by bringing images of ships of a certain category closer to the agents to which that category belongs, and obtain the aggregation result.

[0090] The feature aggregation module includes an aggregation block and a multi-agent module. The feature association is reconstructed through the aggregation block, and then the multi-agent module constrains the ship images of a certain category to be close to the agent to which that category belongs.

[0091] In one embodiment, step S232 includes: S2321 performs convolution, normalization, and activation operations on the separated feature maps to obtain the aggregated block output feature map.

[0092] The aggregation block consists of two aggregation layers. Each aggregation layer includes a 1×1 convolution, a normalization layer, and a ReLU activation layer to process the latent relationships of the separated features and accelerate the aggregation process. The formula is expressed as:

[0093] In the formula, This serves as the output feature map for the first-stage feature separation and the input feature map for the second-stage feature aggregation. This is the output feature map of the aggregated block. and These are 1×1 and 3×3 convolution operations, respectively. Based on the bibranch expansion, the formula is expressed as:

[0094] In the formula, This represents the value set for the experiment's batch size. Represents the number of categories. represent The category to which it belongs.

[0095] S2322, input the feature map output by the aggregation block into the multi-agent module, constrain the ship images of a certain category to be close to the agent to which the category belongs, and obtain the separation feature map.

[0096] The feature aggregation process is completed using a multi-agent module, employing multiple agents to address the issue of varying appearances even among ships of the same category. A multi-vector agent set is introduced. Where N represents the number of agents, K represents the number of ship types, and C represents the number of channels. The agent vector is updated through a dynamic clustering mechanism. ,in For the first Class 1 The mean of the features corresponding to each agent.

[0097] S233: Input the aggregation results into the fully connected layer to obtain the fine-grained classification results of ships.

[0098] The training process of the ship fine-grained classification network is described below: (1) Input the ship image training set into the ship fine-grained classification network to obtain the ship fine-grained classification results; (This process is similar to the application process, please refer to the above steps, and will not be repeated here) (2) Construct a total loss function based on the ship fine-grained classification results, and iteratively train the ship fine-grained classification network based on the total loss function.

[0099] The total loss function used to train the fine-grained ship classification network is:

[0100]

[0101] in, This represents the final model inference category result. These are the actual labels of the input image. It is the cross-entropy loss function.

[0102] During training, the feature separation loss function constrains the feature maps of different views of the same image to be close to each other, and separates the views of different images from each other. The feature separation loss function is:

[0103] in, For feature separation loss function, For batches, For mirror separation loss function, The aggregation loss function; The feature separation loss function is:

[0104]

[0105]

[0106] in, For feature separation loss function, For the feature map after separation, For another feature map that is far away, To prevent model collapse during training, the operator blocks gradient backpropagation. These represent the degree of separation and the degree of aggregation between two feature maps a and b, respectively. The aggregation loss function is: .

[0107] During training, a feature aggregation loss function is used to constrain ship images of a certain category to be closer to the agent belonging to that category. The feature aggregation loss function is:

[0108] in, For feature aggregation loss function, The loss is calculated based on the feature map's proximity to the subclass proxy. To avoid the losses caused by non-subclass proxies, The losses incurred due to the need to maintain differentiation among the various agents; The loss for feature maps that are closer to the subclass proxy is:

[0109] The loss for feature maps that are closer to the subclass proxy is:

[0110] The losses incurred by maintaining differentiation among agents are:

[0111] in, The feature map of the aggregated block shows the category labels as follows: , Represents a set .

[0112] The following is a specific example: The experiment in this embodiment was conducted in an NVIDIA RTX A5000 hardware environment and a Python software environment.

[0113] The constructed ship fine-grained classification system based on a dual-branch collaborative decoupling network was optimized using the SGD optimizer. The initial learning rate was set to 0.01 and adjusted using a cosine annealing strategy. A total of 100 iterations were trained, including a 10-iteration warm-up phase to stabilize the training process.

[0114] The dataset used in this embodiment is the FGSCM-52 dataset released by the National University of Defense Technology, which contains 9562 images. The training set, validation set, and test set are divided in a 3:1:1 ratio, that is, the number of images in the training set, validation set, and test set are 5738, 1912, and 1912, respectively.

[0115] This invention was compared with five existing methods published in authoritative journals for fine-grained ship classification, namely: [Zhuang P, Wang Y, Qiao Y. Learning attentive pairwise interaction for fine-grained classification[C] / / Proceedings of the AAAI conference on artificial intelligence. 2020, 34(07): 13130-13137.] (Comparison Method 1), [Nauta M, VanBree R, Seifert C. Neural prototype trees for interpretable fine-grained image recognition[C] / / Proceedings of the IEEE / CVF conference on computer vision and pattern recognition. 2021: 14933-14943.] (Comparison Method 2), [Sun H, He X, Peng Y. Sim-trans: Structure information modeling transformer for fine-grained visual categorization[C] / / Proceedings of the 30th ACM internationalconference on multimedia. 2022: 5853-5861.] (Compare method 3), [Xu Q, Wang J, JiangB, et al. Fine-grained visual classification via internal ensemble learning transformer[J]. IEEE Transactions on Multimedia, 2023, 25: 9015-9028.] (Compare method 4), [Xiong W, Xiong Z, Yao L, et al. Cog-net: A cognitive network for fine-grained ship classification and retrieval in remote sensing images[J].[IEEE Transactions on Geoscience and Remote Sensing, 2024, 62: 1-17.] (Comparison with Method 5).

[0116] Figure 4 The confusion matrix visualization of the fine-grained classification results provided in this application is as follows: Figure 4 As shown. Figure 4 From top to bottom and left to right, each row represents: the visualization results of the confusion matrix of comparison method 1, comparison method 2, comparison method 3, comparison method 4, comparison method 5, and the visualization results of the confusion matrix of the present invention. The comparison results show that the classification performance of each method varies. The color distribution intuitively reflects the advantages and limitations of the models in various fine-grained ship classification tasks. Among them, the confusion matrix of the present invention has the brightest and most continuous diagonal color, indicating that it has a high confidence level in classifying most categories, especially in categories with high inter-class similarity, where the off-diagonal area has fewer bright yellow spots than other comparison methods. Taking the Asahi-class destroyer (number 50) as an example, the average accuracy of other comparison methods (excluding the present invention) in classifying this category is less than 50%, often confusing it with the Murasame-class destroyer (number 26) or the Garibaldi aircraft carrier (number 33). The present invention, however, achieves a 100% accuracy rate in classifying this category, further verifying that the present invention has better feature discrimination ability for easily confused categories of ships. Overall, the confusion matrix of this invention exhibits a cleaner diagonal distribution and sparser misclassified regions, confirming the effectiveness of the multi-semantic collaborative attention and asynchronous contrast decoupling mechanism. Especially in challenging samples with large intra-class differences (e.g., bulk carriers of the same type with different paint schemes) and high inter-class similarities (e.g., destroyers and attack ships), the method of this invention achieves more refined semantic distinctions through decoupling loss constraints and local feature enhancement.

[0117] Based on the method described in the above embodiments, this embodiment will further describe the ship fine-grained classification system from the perspective of a dual-branch cooperative decoupling network. The ship fine-grained classification system based on the dual-branch cooperative decoupling network can be implemented as an independent entity or integrated into an electronic device. The electronic device can be a terminal, server, or other device. The terminal can include a tablet computer, a laptop computer, a personal computer (PC), a microprocessor box, or other devices.

[0118] Please see Figure 5 , Figure 5This application provides a detailed description of a fine-grained ship classification system based on a dual-branch cooperative decoupling network, applicable to electronic devices. This fine-grained ship classification system may include: The acquisition module is used to acquire images of the ship. The ship classification module is used to input ship images into a fine-grained ship classification network to obtain ship classification results; The ship fine-grained classification network includes a feature extractor, a multi-semantic collaborative attention module, and an asynchronous feature decoupling module. The processing steps of the ship fine-grained classification network include: The augmented ship image is input into a feature extractor to extract features and obtain a preliminary dual-view feature map. The initial dual-view feature map is input into the multi-semantic collaborative attention module for spatial encoding and channel self-attention to obtain the enhanced feature map; The enhanced feature map is input into the asynchronous feature decoupling module for feature separation and feature aggregation, and then the fine-grained classification result of the ship is obtained through a fully connected layer.

[0119] In specific implementation, the above modules and / or units can be implemented as independent entities, or they can be arbitrarily combined and implemented as the same or several entities. For the specific implementation of the above modules and / or units, please refer to the previous method embodiments. For the specific beneficial effects that can be achieved, please also refer to the beneficial effects in the previous method embodiments, which will not be repeated here.

[0120] In addition, this application also provides an electronic device, which may be a computer, tablet computer, or other similar device. This electronic device can implement the steps of any embodiment of the ship fine-grained classification method based on a dual-branch cooperative decoupling network provided in this application. Therefore, it can achieve the beneficial effects that any ship fine-grained classification method based on a dual-branch cooperative decoupling network provided in this invention can achieve, as detailed in the preceding embodiments, and will not be repeated here.

[0121] Figure 6 The diagram illustrates a specific structural block diagram of an electronic device provided in an embodiment of the present invention. This electronic device can be used to implement the ship fine-grained classification method based on a dual-branch cooperative decoupling network provided in the above embodiments. The electronic device 500 can be a terminal, server, or other device. The terminal can include a tablet computer, laptop computer, personal computer (PC), microprocessor box, or other devices.

[0122] RF circuit 510 is used to receive and transmit electromagnetic waves, converting electromagnetic waves into electrical signals and vice versa, thereby enabling communication with communication networks or other devices. RF circuit 510 may include various existing circuit elements used to perform these functions, such as antennas, radio frequency transceivers, digital signal processors, encryption / decryption chips, subscriber identity modules (SIM cards), memory, etc. RF circuit 510 can communicate with various networks such as the Internet, corporate intranets, and wireless networks, or communicate with other devices via wireless networks. The aforementioned wireless networks may include cellular telephone networks, wireless local area networks (WLANs), or metropolitan area networks (MANs). The aforementioned wireless networks may use various communication standards, protocols, and technologies, including but not limited to Global System for Mobile Communication (GSM), Enhanced Data GSM Environment (EDGE), Wideband Code Division Multiple Access (WCDMA), Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Wireless Fidelity (Wi-Fi) (such as IEEE 802.11a, IEEE 802.11b, IEEE 802.11g, and / or IEEE 802.11n), Voice over Internet Protocol (VoIP), Worldwide Interoperability for Microwave Access (Wi-Max), other protocols for email, instant messaging, and short messages, and any other suitable communication protocols, including those that have not yet been developed.

[0123] The memory 520 can be used to store software programs and modules, such as the program instructions / modules corresponding to those in the above embodiments. The processor 580 executes various functional applications and data processing by running the software programs and modules stored in the memory 520, such as taking pictures with the front-facing camera, processing the captured images, and switching the display colors of the content displayed on the screen. The memory 520 may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 520 may further include memory remotely located relative to the processor 580, and these remote memories can be connected to the electronic device 500 via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.

[0124] The input unit 530 can be used to receive input numeric or character information, and to generate a keyboard and mouse related to user settings and function control. Display unit 540 can be used to display information input by the user or information provided to the user, as well as various graphical user interfaces, which can be composed of graphics, text, icons, video, and any combination thereof. Display unit 540 may include display panel 541, which may optionally be configured in the form of LCD (Liquid Crystal Display), OLED (Organic Light-Emitting Diode), or other similar forms.

[0125] Audio circuitry 560, speaker 561, and microphone 562 provide an audio interface between the user and electronic device 500. Audio circuitry 560 converts received audio data into electrical signals and transmits them to speaker 561, where speaker 561 converts them into sound signals for output. Conversely, microphone 562 converts collected sound signals into electrical signals, which are then received by audio circuitry 560, converted back into audio data, and processed by processor 580. The audio data is then transmitted via RF circuitry 510 to, for example, another terminal, or output to memory 520 for further processing. Audio circuitry 560 may also include an earphone jack to facilitate communication between external headphones and electronic device 500.

[0126] Electronic device 500, through transmission module 570 (e.g., Wi-Fi module), can help users receive requests, send information, etc., providing users with wireless broadband internet access. Although transmission module 570 is shown in the figure, it is understood that it is not an essential component of electronic device 500 and can be omitted as needed without changing the essence of the invention.

[0127] The processor 580 is the control center of the electronic device 500. It connects to various parts of the phone via various interfaces and lines, and performs various functions and processes data of the electronic device 500 by running or executing software programs and / or modules stored in the memory 520, and by calling data stored in the memory 520, thereby providing overall monitoring of the electronic device. Optionally, the processor 580 may include one or more processing cores; in some embodiments, the processor 580 may integrate an application processor and a modem processor, wherein the application processor mainly handles the operating system, user interface, and applications, and the modem processor mainly handles wireless communication. It is understood that the modem processor may also not be integrated into the processor 580.

[0128] Electronic device 500 also includes a power supply 590 (such as a battery) that supplies power to various components. In some embodiments, the power supply may be logically connected to processor 580 through a power management system, thereby enabling functions such as charging, discharging, and power consumption management through the power management system. The power supply 590 may also include one or more DC or AC power supplies, recharging systems, power fault detection circuits, power converters or inverters, power status indicators, and other arbitrary components.

[0129] Although not shown, the electronic device 500 also includes cameras (such as front-facing cameras and rear-facing cameras), Bluetooth modules, etc., which will not be described in detail here. Specifically, in this embodiment, the display unit of the electronic device is a touch screen display, and the mobile terminal also includes a memory and one or more programs, wherein one or more programs are stored in the memory and configured to be executed by one or more processors. One or more programs contain instructions for performing the following operations: Acquire images of the ship; Ship images are input into a fine-grained ship classification network to obtain ship classification results; The ship fine-grained classification network includes a feature extractor, a multi-semantic collaborative attention module, and an asynchronous feature decoupling module. The processing steps of the ship fine-grained classification network include: The augmented ship image is input into a feature extractor to extract features and obtain a preliminary dual-view feature map. The initial dual-view feature map is input into the multi-semantic collaborative attention module for spatial encoding and channel self-attention to obtain the enhanced feature map; The enhanced feature map is input into the asynchronous feature decoupling module for feature separation and feature aggregation, and then the fine-grained classification result of the ship is obtained through a fully connected layer.

[0130] In practice, the above modules can be implemented as independent entities or combined in any way to be implemented as the same or several entities. For the specific implementation of the above modules, please refer to the previous method implementation examples, which will not be repeated here.

[0131] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by instructions, or by instructions controlling related hardware. These instructions can be stored in a computer-readable storage medium and loaded and executed by a processor. Therefore, embodiments of the present invention provide a storage medium storing multiple instructions that can be loaded by a processor to execute the steps of any embodiment of the ship fine-grained classification method based on a dual-branch cooperative decoupling network provided by the present invention.

[0132] The computer-readable storage medium may include: read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.

[0133] Since the instructions stored in the storage medium can execute the steps in any embodiment of the ship fine-grained classification method based on a dual-branch cooperative decoupling network provided in the embodiments of the present invention, the beneficial effects that any ship fine-grained classification method based on a dual-branch cooperative decoupling network provided in the embodiments of the present invention can achieve can be realized. For details, please refer to the previous embodiments, which will not be repeated here.

[0134] The foregoing has provided a detailed description of a ship fine-grained classification method, system, storage medium, and electronic device based on a dual-branch cooperative decoupling network provided in the embodiments of this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the embodiments above are only for the purpose of helping to understand the method and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. A fine-grained classification method for ships based on a dual-branch cooperative decoupling network, characterized in that, The method includes: Acquire images of the ship; Ship images are input into a fine-grained ship classification network to obtain ship classification results; The ship fine-grained classification network includes a feature extractor, a multi-semantic collaborative attention module, and an asynchronous feature decoupling module. The processing steps of the ship fine-grained classification network include: The augmented ship image is input into a feature extractor to extract features and obtain a preliminary dual-view feature map. The initial dual-view feature map is input into the multi-semantic collaborative attention module for spatial encoding and channel self-attention to obtain the enhanced feature map; The enhanced feature map is input into the asynchronous feature decoupling module for feature separation and feature aggregation, and then the fine-grained classification result of the ship is obtained through a fully connected layer.

2. The ship fine-grained classification method based on a dual-branch cooperative decoupling network according to claim 1, characterized in that, The multi-semantic collaborative attention module includes a multi-semantic spatial grouping module and a channel attention sparse module. The step of inputting the initial dual-view feature map into the multi-semantic collaborative attention module for spatial encoding and channel self-attention to obtain an enhanced feature map includes: The initial dual-view feature map is input into the multi-semantic space grouping module for global pooling, multi-scale feature extraction and stitching operations to obtain the first feature map. The first feature map is input into the channel attention sparse module, and the enhanced feature map is obtained based on the adaptive pooling compression and channel self-attention joint computation framework.

3. The ship fine-grained classification method based on a dual-branch cooperative decoupling network according to claim 2, characterized in that, The process of inputting the preliminary dual-view feature map into the multi-semantic space grouping module for global pooling, multi-scale feature extraction, and concatenation operations to obtain the first feature map includes: The initial dual-view feature map is subjected to biaxial average pooling along the X and Y directions to obtain two one-dimensional sequence structures. The channel dimension of the one-dimensional sequence structure is divided into four independent sub-features. Apply depthwise one-dimensional convolutions with different kernel sizes to each group of sub-features to obtain the modified sub-features; The modified sub-features in a certain direction are concatenated, and spatial attention weights are generated using group normalization and the Sigmoid function. The first feature map is generated based on spatial attention weights and preliminary dual-view feature maps.

4. The ship fine-grained classification method based on a dual-branch cooperative decoupling network according to claim 3, characterized in that, The step of inputting the first feature map into the channel attention sparse module and obtaining the enhanced feature map based on the adaptive pooling compression and channel self-attention joint computation framework includes: The first feature map is compressed using adaptive average pooling to obtain a compressed feature map; Generate a query matrix based on the compressed feature map. Key matrix Sum matrix Then according to Calculate the attention distribution matrix; The top 50% of channels are selected, and the channels are sparsified to obtain enhanced feature maps.

5. The ship fine-grained classification method based on a dual-branch cooperative decoupling network according to claim 1, characterized in that, The asynchronous feature decoupling module includes a feature separation module, a feature aggregation module, and a fully connected layer. The enhanced feature map is input into the asynchronous feature decoupling module for feature separation and aggregation, and then the fine-grained classification result of the ship is obtained through the fully connected layer, including: The enhanced features are pooled, an orthogonal projection matrix is ​​constructed to perform a linear transformation on the feature space, sample pairs of the same image are aggregated, and sample pairs of different images are separated. Reconstruct feature associations by bringing images of ships of a certain category closer to the agents belonging to that category to obtain the aggregation result; The aggregation results are input into a fully connected layer to obtain fine-grained classification results for ships.

6. The ship fine-grained classification method based on a dual-branch cooperative decoupling network according to claim 5, characterized in that, The feature separation module includes a dual-pooling layer and a separation block. It performs pooling operations on the enhanced features, constructs an orthogonal projection matrix to perform a linear transformation on the feature space, aggregates sample pairs from the same image, and separates sample pairs from different images, including: Global average pooling and local max pooling operations are performed on the enhanced features to obtain hull morphology and structural detail feature maps. The hull shape and structural detail feature maps are mapped to the decoupled space to generate feature vectors; The feature vectors are mirrored and flipped along the batch dimension to construct the feature matrix; By constraining the feature maps of different views of the same image to be close to each other, and separating the views of different images from each other, we obtain separated feature maps.

7. The ship fine-grained classification method based on a dual-branch cooperative decoupling network according to claim 5, characterized in that, The feature aggregation module includes an aggregation block and a multi-agent module. The reconstructed feature association involves bringing ship images of a certain category closer to the agents belonging to that category to obtain the aggregation result, including: The separated feature maps are subjected to convolution, normalization, and activation operations to obtain the aggregated block output feature map; The feature map output from the aggregated block is input into the multi-agent module, which constrains ship images of a certain category to be close to the agent to which that category belongs, thus obtaining a separation feature map.

8. The ship fine-grained classification method based on a dual-branch cooperative decoupling network according to claim 6 or 7, characterized in that, During training, the feature maps of different views of the same image are constrained to be close to each other, while the views of different images are separated from each other, by using the feature separation loss function. The feature separation loss function is: in, For feature separation loss function, For batches, For mirror separation loss function, The aggregation loss function; The feature separation loss function is: in, For feature separation loss function, For the feature map after separation, For another feature map that is far away, To block gradient backpropagation for the operator, These represent the degree of separation and the degree of aggregation between two feature maps a and b, respectively. The aggregation loss function is: 。 9. The ship fine-grained classification method based on a dual-branch cooperative decoupling network according to claim 6 or 7, characterized in that, During training, a feature aggregation loss function is used to constrain ship images of a certain category to be closer to the agent belonging to that category. This feature aggregation loss function is: in, For feature aggregation loss function, The loss is calculated based on the feature map's proximity to the subclass proxy. To avoid the losses caused by non-subclass proxies, The losses incurred due to the need to maintain differentiation among the various agents; The loss for feature maps that are closer to the subclass proxy is: The loss for feature maps that are closer to the subclass proxy is: The losses incurred by maintaining differentiation among agents are: in, The feature map of the aggregated block shows the category labels as follows: , Represents a set ; The total loss function used to train the fine-grained ship classification network is: in, This represents the final model inference category result. These are the actual labels of the input image. It is the cross-entropy loss function.

10. A fine-grained ship classification system based on a dual-branch cooperative decoupling network, characterized in that, include: The acquisition module is used to acquire images of the ship. The ship classification module is used to input ship images into a fine-grained ship classification network to obtain ship classification results; The ship fine-grained classification network includes a feature extractor, a multi-semantic collaborative attention module, and an asynchronous feature decoupling module. The processing steps of the ship fine-grained classification network include: The augmented ship image is input into a feature extractor to extract features and obtain a preliminary dual-view feature map. The initial dual-view feature map is input into the multi-semantic collaborative attention module for spatial encoding and channel self-attention to obtain the enhanced feature map; The enhanced feature map is input into the asynchronous feature decoupling module for feature separation and feature aggregation, and then the fine-grained classification result of the ship is obtained through a fully connected layer.

Citation Information

Cited By

  • LIBS multi-component quantitative analysis method based on spectral mechanism constraint and feature decoupling

    CN122508078A