ConvNeXt model SAR (Synthetic Aperture Radar) ship classification method based on multilevel feature collaborative interaction

By combining Canny edge detection and multi-level branched ConvNeXt model in SAR ship classification, multi-scale features are extracted and fused, and the problem of existing models ignoring traditional features and insufficient fusion of scales is achieved, and high-precision ship classification is achieved.

CN120388301AActive Publication Date: 2025-07-29HARBIN ENG UNIV
View PDF 10 Cites 0 Cited by

Patent Information

Application Number
CN202510401718.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-01
Publication Date
2025-07-29
Estimated Expiration
2045-04-01

AI Technical Summary

Technical Problem

The existing SAR ship classification model based on deep learning ignores traditional manual features and lacks mining and fusion of features at different scales, resulting in difficulty in improving classification accuracy.

Method used

The Canny-based edge detection method is used to extract traditional manual features, and combine the ConvNeXt model with multi-level branch structure for feature extraction and cross-fusion, retaining feature information at different scales, and optimizing the model through multi-task regression and cross-entropy loss function.

Benefits of technology

It improves the accuracy and accuracy of SAR ship classification, realizes efficient feature characterization capabilities, and adapts to target recognition in complex marine environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120388301A_ABST
    Figure CN120388301A_ABST
Patent Text Reader

Abstract

The invention provides a ConvNeXt model SAR (Synthetic Aperture Radar) ship classification method based on multilevel feature collaborative interaction. In order to effectively utilize traditional manual features, the method uses an edge detection method based on Canny, captures edge information of an object in an image through multistage filtering and edge gradient analysis, and inputs a result into a ConvNeXt model for feature extraction. Besides, a multi-level branch structure is expanded in the ConvNeXt model and is used for deep extraction of feature information under different scales, so that various kinds of information of the image are captured at the same time on low-level features and high-level features. And finally, cross fusion with traditional manual features is carried out in the multi-scale feature extraction process, so that the classification precision of the model is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of target classification of synthetic aperture radar (SAR), and particularly to a SAR ship classification method based on a ConvNeXt model with multi-level feature collaborative interaction. Background Art

[0002] Synthetic Aperture Radar (SAR) is an advanced remote sensing technology that realizes precise imaging of the ground or ocean surface by transmitting high-frequency electromagnetic waves and receiving their reflected signals. Compared with traditional optical imaging technology, SAR has the ability of all-weather and all-time high-resolution imaging, and plays an important role in marine environmental monitoring and resource management. It is not restricted by weather and lighting conditions, can continuously track and monitor marine targets, and provides key technical support for the construction of a smart marine management system. At the same time, the advantages of SAR in ship intelligent information recognition and classification contribute to improving the efficiency of maritime traffic management and promoting the sustainable development of the blue economy.

[0003] In recent years, with the rapid development of deep learning technology, its application in the SAR ship target classification task has become a research hotspot, but at the same time, it also faces many challenges. First, the existing SAR ship classification models based on deep learning mainly rely on the high-level features automatically extracted by neural networks, and often ignore the traditional manual features containing rich expert experience, resulting in the training effect of the model being difficult to be fully optimized. Second, the mainstream convolutional neural network (CNN) usually extracts features from shallow to deep step by step through multiple layers of convolution, but the mining and fusion of features at different scales are still insufficient, which limits the further improvement of classification accuracy. Therefore, how to effectively combine traditional manual features and make full use of multi-scale feature information has become an important research topic for improving the performance of SAR ship classification. Summary of the Invention

[0004] The purpose of the present invention is to solve the problems in the prior art, and a SAR ship classification method based on a ConvNeXt model with multi-level feature collaborative interaction is proposed.

[0005] The present invention is realized through the following technical solutions. The present invention proposes a SAR ship classification method based on a ConvNeXt model with multi-level feature collaborative interaction, and the method includes the following steps:

[0006] Step 1: Process the currently input SAR ship sample through the Canny method to extract traditional manual features and generate a corresponding picture with edge information;

[0007] Step 2: Extract features from the generated picture with edge information, and select the ConvNeXt model as the network for feature extraction; during the process of layer-by-layer convolution, features of different scales are retained for subsequent cross-fusion of features;

[0008] Step 3: Extract features from the currently input SAR ship sample, and select the ConvNeXt model as the network for feature extraction; after passing through a multi-level branch structure and a feature cross-fusion structure, the input image is transformed into a feature vector, and it combines multi-dimensional and multi-scale feature information;

[0009] Step 4: Use a classification head to perform multi-task regression on the features, and assign weight coefficients to adapt to the scenario of SAR ship classification, and finally obtain the classification result.

[0010] Furthermore, the dataset used in Step 1 is the FUSAR-Ship ship dataset. The dataset is processed by image enhancement and the samples in the dataset are augmented.

[0011] Furthermore, edge point judgment is performed in Step 1. The specific judgment method is as follows: If the amplitude of a certain pixel point is greater than the high threshold, the pixel point is retained as an edge point; if the amplitude of a certain pixel point is less than the low threshold, the pixel point is excluded; if the amplitude of a certain pixel point is between the two thresholds, the pixel point is retained as a possible edge point; finally, all the final edge points are connected.

[0012] Furthermore, Step 2 extracts features from the edge information picture through the ConvNeXt model, which consists of 4 stages. Each stage is connected by a downsampling layer to retain 4 features of different scales for the subsequent feature cross-fusion structure.

[0013] Furthermore, the ConvNeXt model is composed of ConvNeXt Blocks, and its structure is depthwise separable convolution. Depthwise separable convolution is divided into two processes, namely depthwise convolution and pointwise convolution;

[0014] Depthwise convolution is to perform convolution operations independently on each input channel; assuming the input feature map is and the convolution kernel is then the operation of depthwise convolution can be expressed as:

[0015]

[0016] where Y d is the output after depthwise convolution, i and j are the spatial positions of the output feature map, and c is the channel index; this process is performed independently for each channel, so different convolution kernels are used for convolution on each channel;

[0017] Pointwise convolution is a 1×1 convolution operation that fuses information in the depth direction; assuming that the output feature map has a size of and the convolution kernel is then the operation of pointwise convolution can be expressed as:

[0018]

[0019] where Y p is the output after pointwise convolution, c′ is the output channel index, and C out is the number of output channels; the role of pointwise convolution is to merge all channel information at each position into new channel information.

[0020] Furthermore, in step 3, the ConvNeXt model is used to extract features from the original image, which consists of 4 stages with a multi-level branch structure. During the process of extracting feature depth in each branch, feature cross-fusion is performed.

[0021] Furthermore, a multi-level branch structure is added to the ConvNeXt model in step 3 to process features at different scales. High-scale feature maps and low-scale feature maps can effectively interact at each layer; feature maps of different scales exchange information through a fusion module to ensure that high-scale feature maps maintain clarity throughout the network and low-scale feature maps obtain sufficient context information in the deep layers of the network; assuming that at layer l, the input feature map The processing at different resolutions can be expressed as:

[0022]

[0023] where Fuse(·) represents the feature fusion operation;

[0024] Meanwhile, during the extraction process of feature maps at each scale, cross-fusion with the feature extraction of the edge information picture is performed; at a certain scale, the input feature maps are respectively and then the cross-fusion operation is:

[0025]

[0026] where is the fused feature, and

[0027] Furthermore, in step 4, the cross-entropy loss function and label smoothing are adopted to prevent overfitting problems;

[0028] The formula of the cross-entropy loss function is as follows:

[0029]

[0030] Among them, x i is the result of passing the model output through softmax, and y i represents whether it is the corresponding class label. The label smoothing method is used for the class label, which changes the probability distribution. The specific formula is:

[0031]

[0032] Among them, ε is a constant.

[0033] The present invention also proposes an electronic device, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the method for classifying SAR ships based on a ConvNeXt model with multi-level feature collaborative interaction are implemented.

[0034] The present invention also proposes a computer-readable storage medium for storing computer instructions. When the computer instructions are executed by a processor, the steps of the method for classifying SAR ships based on a ConvNeXt model with multi-level feature collaborative interaction are implemented.

[0035] Compared with the prior art, the beneficial effects of the present invention are:

[0036] The present invention provides a ConvNeXt model based on multi-level feature collaborative interaction for the method of classifying ships in SAR images. This method obtains a manually traditional feature image through an edge detection method, and inputs it together with the original image into the ConvNeXt model with multiple levels of branches for feature extraction, and performs cross-fusion between features, thereby improving the representation ability of the model. Efficient and accurate classification of ships in SAR images is achieved. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained according to the provided drawings.

[0038] Figure 1 It is a flowchart of the method for classifying SAR ships based on a ConvNeXt model with multi-level feature collaborative interaction according to the present invention.

[0039] Figure 2 It is a structural framework diagram of the ConvNeXt model.

[0040] Figure 3Schematic diagram of the image generated for edge feature extraction in the embodiment. Among them, (a) is the input picture, and (b) is the extraction result.

[0041] Figure 4 Schematic diagram of the classification result in the embodiment, where (a) and (b) are the classification results formed by different input pictures. Detailed implementation manners

[0042] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0043] Aiming at the problems in the prior art, the present invention proposes a ConvNeXt model based on multi-level feature collaborative interaction for SAR image ship classification. To effectively utilize traditional handcrafted features, a Canny-based edge detection method is used to capture the edge information of objects in the image through multi-level filtering and edge gradient analysis, and the result is input into the ConvNeXt model for feature extraction. In addition, a multi-level branch structure is extended in the ConvNeXt model for in-depth extraction of feature information at different scales, so as to capture multiple information of the image on both low-level features and high-level features. Finally, cross-fusion with traditional handcrafted features is performed during the multi-scale feature extraction process to improve the classification accuracy of the model.

[0044] Specifically, in combination with Figures 1 - 4 , the present invention proposes a SAR ship classification method based on a ConvNeXt model with multi-level feature collaborative interaction, and the method includes the following steps:

[0045] Step 1: Process the currently input SAR ship sample through the Canny method to perform traditional handcrafted feature extraction, and generate a corresponding picture with edge information.

[0046] The dataset used in Step 1 is the FUSAR-Ship ship dataset. Considering the problem of unbalanced samples in each category, the dataset is processed by image enhancement to expand the samples of the dataset. Next, the dataset is divided into a training set, a validation set, and a test set according to the ratio of 7:2:1. Finally, the training parameters are set.

[0047] In step 1, edge point judgment is performed. The specific judgment method is as follows: If the amplitude of a certain pixel point is greater than the high threshold, this pixel point is retained as an edge point; if the amplitude of a certain pixel point is less than the low threshold, this pixel point is excluded; if the amplitude of a certain pixel point is between the two thresholds, this pixel point is retained as a possible edge point; finally, all the final edge points are connected.

[0048] Step 2: Extract features from the generated picture with edge information, and select the ConvNeXt model as the network for feature extraction; during the process of layer-by-layer convolution, features of different scales are retained for subsequent cross-fusion of features;

[0049] In step 2, the ConvNeXt model is used to extract features from the edge information picture, which consists of 4 stages. Each stage is connected by a downsampling layer to retain 4 features of different scales for the subsequent feature cross-fusion structure.

[0050] The ConvNeXt model is composed of ConvNeXt Blocks, and its structure is depthwise separable convolution. Depthwise separable convolution is divided into two processes, namely depthwise convolution and pointwise convolution;

[0051] Depthwise convolution is to perform convolution operations independently on each input channel; assume the input feature map is and the convolution kernel is Then the operation of depthwise convolution can be expressed as:

[0052]

[0053] where Y d is the output after depthwise convolution, i and j are the spatial positions of the output feature map, and c is the channel index; this process is performed independently for each channel, so different convolution kernels are used for convolution on each channel;

[0054] Pointwise convolution is a 1×1 convolution operation that fuses the information in the depth direction; assume that after pointwise convolution, the size of the output feature map is and the convolution kernel is Then the operation of pointwise convolution can be expressed as:

[0055]

[0056] where Y p is the output after pointwise convolution, c′ is the output channel index, and C out is the number of output channels; the role of pointwise convolution is to merge all channel information at each position into new channel information.

[0057] Step 3: Extract features from the currently input SAR ship samples, and select the ConvNeXt model as the network for feature extraction; after passing through a multi-level branch structure and a feature cross-fusion structure, the input image is transformed into a feature vector, and it combines multi-dimensional and multi-scale feature information;

[0058] In Step 3, the ConvNeXt model is used to extract features from the original image, which consists of 4 stages with a multi-level branch structure. During the process of extracting feature depth in each branch, feature cross-fusion is performed.

[0059] In the ConvNeXt model of Step 3, a multi-level branch structure is added to process features at different scales. High-scale feature maps and low-scale feature maps can effectively interact at each layer; feature maps of different scales exchange information through a fusion module to ensure that the high-scale feature maps maintain clarity throughout the network and the low-scale feature maps obtain sufficient context information in the deep layers of the network; assuming at layer l, the input feature map The processing at different resolutions can be expressed as:

[0060]

[0061] where, Fuse(·) represents the feature fusion operation;

[0062] Meanwhile, during the extraction process of feature maps at each scale, cross-fusion with the feature extraction of the edge information picture is performed; at a certain scale, the input feature maps are respectively and Then the cross-fusion operation is:

[0063]

[0064] where, is the fused feature, and

[0065] Step 4: Use the classification head to perform multi-task regression on the features and assign weight coefficients to adapt to the SAR ship classification scenario, and finally obtain the classification result.

[0066] In Step 4, the cross-entropy loss function and label smoothing are adopted to prevent overfitting problems;

[0067] The formula of the cross-entropy loss function is as follows:

[0068]

[0069] where, x i is the result of the model output after softmax, and y iIndicates whether it is the corresponding category label. The label smoothing method is used for the category label, which changes the probability distribution. The specific formula is:

[0070]

[0071] where ε is a constant.

[0072] Embodiment

[0073] The object of the present invention is to solve the problem of SAR ship classification. The deep learning network is used to efficiently and accurately identify SAR ship targets and output the corresponding category information. To achieve this goal, an embodiment of the present invention provides a SAR ship classification method based on a ConvNeXt model with multi-level feature collaborative interaction. The basic process is as Figure 1 shown, and the method includes:

[0074] Step 1: Process the currently input SAR ship sample through the Canny method to extract traditional handcrafted features and generate a corresponding picture with edge information.

[0075] The dataset used in Step 1 is the FUSAR-Ship ship dataset, and six types of ships are selected: Cargo ship, Bulk ship, Container ship, Fishing boat, Tugboat, Tanker. After performing image enhancement processing on the existing samples, each type of ship is expanded to about 1600. Next, the dataset is divided. The training set accounts for 70% of the total number of images, the validation set accounts for 20% of the total number of images (the training set and the test set are randomly generated), and the test set accounts for 10% of the total number of images. During training, the input image is fixed at 512×512. The training batch size is 8, and the training iteration times are 500.

[0076] Before performing edge detection, it is necessary to grayscale and filter the image first. Grayscaling is to simplify the subsequent processing of the image and reduce the complexity and information processing volume of the image. It is obtained by multiplying the three channels B, G, and R of the original image by certain weights and then adding them together. The grayscale image can represent most of the features of the image with less data information.

[0077] To determine whether a certain pixel point is locally maximum in the gradient direction, non-maximum suppression is completed. It is necessary to use the Sobel filter to calculate the gradient amplitude and direction of this point. A pair of convolutional arrays (in the x and y directions) are used to calculate the derivative.

[0078]

[0079] At the same time, the following formulas are used to calculate the gradient amplitude and direction.

[0080]

[0081] Among them, for the convenience of judgment and to improve efficiency, the gradient direction θ is generally selected from one of 0°, 45°, 90°, and 135°. In the current task, the present invention selects 0°.

[0082] After that, non-maximum suppression processing is performed, which means comparing this point with two adjacent points in the gradient direction. If this point is locally maximum, it is marked as a possible edge point. In this way, non-edge pixels are excluded, and only some candidate edges are retained. However, this still includes many false edges caused by noise and other reasons. Therefore, a double-threshold hysteresis threshold is also required to reduce the number of false edges. The specific method for judging edge points is as follows: If the amplitude of a certain pixel point is greater than the high threshold, this pixel point is retained as an edge point. If the amplitude of a certain pixel point is less than the low threshold, this pixel point is excluded. If the amplitude of a certain pixel point is between the two thresholds, this pixel point is retained as a possible edge point. Finally, all the final edge points are connected.

[0083] Step 2: Extract features from the generated picture with edge information, and select the ConvNeXt model as the network for feature extraction. During the process of layer-by-layer convolution, features of different scales are retained for subsequent cross-fusion of features.

[0084] Step 2 uses the ConvNeXt model, which is a neural network architecture inspired by Transformer but still uses convolution. This model is composed of ConvNeXt Blocks, and the most important structure among them is the depthwise separable convolution. The depthwise separable convolution is mainly divided into two processes, namely depthwise convolution and pointwise convolution.

[0085] Depthwise convolution performs convolution operations independently on each input channel. Assume the input feature map is The convolution kernel is Then the operation of depthwise convolution can be expressed as:

[0086]

[0087] Among them, Y d is the output after depthwise convolution, i and j are the spatial positions of the output feature map, and c is the channel index. This process is performed independently for each channel, so different convolution kernels are used for convolution on each channel.

[0088] Pointwise convolution is a 1×1 convolution operation that fuses the information in the depth direction. Assume that after the output feature map passes through pointwise convolution, its size is The convolution kernel is Then the operation of pointwise convolution can be expressed as:

[0089]

[0090] Among them, Y p is the output after pointwise convolution, c′ is the output channel index, and C out is the number of output channels. The role of pointwise convolution is to combine all channel information at each position into new channel information.

[0091] Step 3: Extract features from the current input SAR ship samples, and select the ConvNeXt model as the network for feature extraction. After passing through the multi-level branch structure and the feature cross-fusion structure, the input image is transformed into a feature vector, and it combines multi-dimensional and multi-scale feature information.

[0092] Step 3 improves the traditional ConvNeXt Satge, adds a multi-level branch structure for processing features at different scales, and enables effective interaction between high-scale and low-scale feature maps at each layer. Feature maps of different scales exchange information through the fusion module to ensure that the high-scale feature maps maintain clarity throughout the network and the low-scale feature maps obtain sufficient context information in the deep layers of the network. Assume that at layer l, the input feature map The processing at different resolutions can be expressed as:

[0093]

[0094] Among them, Fuse(·) represents a more efficient feature fusion operation, such as convolution, upsampling, weighted fusion, etc.

[0095] At the same time, during the extraction process of feature maps at each scale, cross-fusion with the feature extraction of the edge information image is performed to further improve the expression ability of multi-resolution features. At a certain scale, the input feature maps are respectively and Then the cross-fusion operation is:

[0096]

[0097] Among them, is the fused feature, and

[0098] Step 4: Use the classification head to perform multi-task regression on the features and assign weight coefficients to adapt to the SAR ship classification scenario, and finally obtain the classification result.

[0099] Step 4 adopts the cross-entropy loss function and label smoothing to prevent overfitting problems.

[0100] The cross-entropy loss function is a common loss function in classification recognition tasks, and the formula is as follows:

[0101]

[0102] Among them, x i is the result of the model output after softmax, and y i represents whether it is the corresponding class label, which can be expressed by the following formula:

[0103]

[0104] This kind of processing will cause the relationship between the true label and other labels to be ignored. When dealing with the classification and recognition problem of the SAR ship dataset with high sample similarity and large data noise, the model is easily affected. Therefore, the label smoothing method is used later, which changes the probability distribution,

[0105]

[0106] where ε is a small constant, which makes the probability optimization target in the softmax loss no longer 1 and 0. To a certain extent, this avoids overfitting and alleviates the impact brought by wrong labels.

[0107] The present invention proposes a ConvNeXt model based on multi-level feature collaborative interaction for the SAR image ship classification method. This method obtains a manual traditional feature image through an edge detection method, and inputs it together with the original image into the ConvNeXt model with multiple levels of branches for feature extraction, and conducts cross-fusion between features, thereby improving the representation ability of the model. Efficient and accurate SAR image ship classification is achieved.

[0108] The present invention also proposes an electronic device, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the steps of the SAR ship classification method of the ConvNeXt model based on multi-level feature collaborative interaction.

[0109] The present invention also proposes a computer-readable storage medium for storing computer instructions. When the computer instructions are executed by a processor, it implements the steps of the SAR ship classification method of the ConvNeXt model based on multi-level feature collaborative interaction.

[0110] The memory in the embodiments of the present application may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read only memory (ROM), a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchlink DRAM (SLDRAM), and direct rambus RAM (DRRAM). It should be noted that the memory of the method described in the present invention is intended to include but not limited to these and any other suitable types of memories.

[0111] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired manner (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server, data center, etc. that contains one or more integrated available media. The available media can be magnetic media (such as floppy disks, hard disks, magnetic tapes), optical media (such as high-density digital video discs (DVDs)), or semiconductor media (such as solid state discs (SSDs)), etc.

[0112] In the implementation process, each step of the above method can be completed by the integrated logic circuit of the hardware in the processor or the instructions in the form of software. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being executed by the hardware processor or executed by the combination of the hardware and software modules in the processor. The software module can be located in a mature storage medium in the art such as random access memory, flash memory, read-only memory, programmable read-only memory, or electrically erasable programmable memory, register, etc. This storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps of the above method. To avoid repetition, it will not be described in detail here.

[0113] It should be noted that the processor in the embodiments of the present application may be an integrated circuit chip with signal processing capabilities. In the implementation process, the steps of the above method embodiments can be completed by the integrated logic circuit in the hardware of the processor or instructions in software form. The above-mentioned processor may be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being executed and completed by the hardware decoding processor, or executed and completed by a combination of the hardware and software modules in the decoding processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps of the above method.

[0114] The above has introduced in detail a ConvNeXt model SAR ship classification method based on multi-level feature collaborative interaction proposed by the present invention. Specific examples are used in this article to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.

Claims

1. A SAR ship classification method for ConvNeXt model based on multi-level feature collaborative interaction, characterized in that The method comprises the following steps: Step 1: The input SAR ship sample is processed by the Canny method and traditional manual feature extraction is performed to generate the corresponding image with edge information; Step 2: Extract features from the generated image with edge information and select the ConvNeXt model as the feature extraction network; retain features of different scales during the layer-by-layer convolution process for subsequent cross-fusion of features; Step 3: Extract features from the currently input SAR ship sample and select the ConvNeXt model as the feature extraction network; after a multi-level branching structure and feature cross-fusion structure, the input image is converted into a feature vector, which combines multi-dimensional and multi-scale feature information; Step 4: Use the classification head to perform multi-task regression on the features and assign weight coefficients to adapt to the SAR ship classification scenario to finally obtain the classification results.

2. The method according to claim 1, wherein The dataset used in step 1 is the FUSAR-Ship ship dataset. The dataset is processed with image enhancement and the samples of the dataset are expanded.

3. The method according to claim 1, characterized in that In step 1, edge point judgment is performed. The specific judgment method is as follows: if the amplitude of a pixel point is greater than the high threshold, the pixel point is retained as an edge point; if the amplitude of a pixel point is less than the low threshold, the pixel point is excluded; if the amplitude of a pixel point is between the two thresholds, the pixel point is retained as a possible edge point; finally, all the final edge points are connected.

4. The method according to claim 1, characterized in that: Step 2 uses the ConvNeXt model to extract features from the edge information image. It consists of four stages, and each stage is connected using a downsampling layer to retain features of four different scales for subsequent feature cross-fusion structure.

5. The method according to claim 4, wherein The ConvNeXt model consists of ConvNeXt Block, whose structure is depth-wise separable convolution. Depth-wise separable convolution is divided into two processes: channel-wise convolution and point-wise convolution. Channel-by-channel convolution is to perform convolution operations independently on each input channel; assuming the input feature map is The convolution kernel is The channel-by-channel convolution operation can be expressed as: Among them, Y d is the output after per-channel convolution, where i and j are the spatial positions of the output feature map, and c is the channel index; this process is performed independently for each channel, so different convolutional kernels are used for convolution in each channel; Point-by-point convolution is a 1×1 convolution operation that fuses information in the depth direction. Assuming that the output feature map has a size of The convolution kernel is The point-by-point convolution operation can be expressed as: Among them, Y p is the output after point-by-point convolution, c′ is the output channel index, C out is the number of output channels; the role of point-by-point convolution is to merge all channel information at each position into new channel information.

6. The method according to claim 1, characterized in that, Step 3 uses the ConvNeXt model to extract features from the original image. It contains 4 stages with a multi-level branch structure. In the process of feature depth extraction in each branch, feature cross-fusion is performed.

7. The method according to claim 6, characterized in that In step 3, a multi-level branch structure is added to the ConvNeXt model to process features at different scales. High-scale feature maps and low-scale feature maps can effectively interact at each layer. Feature maps of different scales perform information exchange through a fusion module to ensure that high-scale feature maps maintain clarity throughout the network and low-scale feature maps obtain sufficient context information in the deeper layers of the network. Assume that at layer l, the input feature map The processing at different resolutions can be expressed as: Among them, Fuse(·) represents the feature fusion operation; At the same time, in the process of extracting each scale feature map, cross-fusion with edge information image extraction features is performed; at a certain scale, the input feature maps are and The cross-fusion operation is: Among them, is the fused feature, and 8. The method according to claim 1, characterized in that, In step 4, the cross entropy loss function and label smoothing are used to prevent overfitting problems; The cross entropy loss function formula is as follows: where x i is the result of passing the model output through softmax, and y i represents whether it is the corresponding class label. The label smoothing method is used for the class label, which changes the probability distribution. The specific formula is: Here, ε is a constant.

9. An electronic device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, the steps of the method according to any one of claims 1 to 8 are implemented.

10. A computer-readable storage medium for storing computer instructions, characterized in that: When the computer instructions are executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.

Citation Information

Patent Citations

  • Ship target real-time detection method and terminal based on improved SSD model

    CN113205151A

  • Method for improving SAR ship classification precision by combining HOG features

    CN113344045A

  • Method for improving SAR image ship classification precision

    CN113344046A

  • Three-dimensional model classification fusing view features and multi-branch networks

    CN116433965A

  • Short-time heavy rainfall typing method based on feature cross fusion

    CN116628626A