Ship target identification instrument based on dense connection mechanism

By adopting a feature extraction method of dense connection mechanism and attention mechanism in the ship target recognition instrument, the problems of low ship target recognition accuracy and weak anti-interference ability in the prior art are solved, and high-precision and high-efficiency ship target recognition are achieved.

CN120070954APending Publication Date: 2025-05-30ZHEJIANG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510058358.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-14
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

When processing SAR images, existing ship target recognition instruments face problems such as low signal-to-noise ratio, complex background and low resolution, resulting in low recognition accuracy and weak anti-interference ability.

Method used

The ship target recognition instrument based on the dense connection mechanism is adopted, and multi-threshold segmentation is performed through the data preprocessing module. The feature extraction module uses the initial convolution layer, dense connection convolution layer and the end convolution layer to extract features, and combines the spatial and channel attention mechanism, and the classification identification module uses the full connection layer for classification.

Benefits of technology

It improves the accuracy and robustness of ship target recognition, enhances anti-interference ability, and achieves high-precision and high-efficiency ship target recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070954A_ABST
    Figure CN120070954A_ABST
Patent Text Reader

Abstract

The invention discloses a ship target recognition instrument based on a dense connection mechanism, which is used for carrying out ship target recognition and classification on SAR (Synthetic Aperture Radar) images and comprises a data preprocessing module, a feature extraction module and a classification and recognition module. According to the method, the swarm intelligence optimization algorithm is introduced to perform multi-threshold segmentation, and the problem of target identification caused by high signal-to-noise ratio and low resolution of the SAR image is solved. A channel attention mechanism and a space attention mechanism are introduced into a dense connection network, and the defects that a traditional ship target recognition instrument is not high in recognition precision and poor in anti-interference capacity are overcome. The finally realized ship target identification instrument has the characteristics of high classification precision, full utilization of an attention mechanism, strong anti-interference capability and fast classification speed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of ship target recognition instruments, and specifically relates to a ship target recognition instrument based on a dense connection mechanism. Background Art

[0002] Synthetic Aperture Radar (SAR) is a high-resolution imaging radar instrument. Due to its ability to operate all-weather under any weather conditions, it plays an important role especially in the recognition of ship targets at sea. As an active microwave imaging sensor, the imaging process of SAR is not affected by environmental factors such as weather, light, and clouds like traditional passive imaging sensors (such as infrared and optical sensors), enabling it to detect hidden targets in complex environments. With the development of technology, SAR has been widely used in civilian and military fields. Ship recognition at sea also involves many fields such as marine vessels, icebergs, reefs, maritime search and rescue, and monitoring of illegal fishing and smuggling. It is the "clairvoyant" for maritime operation safety. How to analyze and recognize ship targets in SAR images has become an important research topic in the future.

[0003] Due to the inherent resolution limitation and noise problems of SAR images, especially when facing dense sea areas and small ship targets, the low signal-to-noise ratio and highly complex background will greatly increase the difficulty of ship target recognition, resulting in limitations in accuracy and reliability. Traditional SAR image processing methods rely on three stages: preprocessing, feature extraction, and classification and recognition. However, these traditional methods often rely on manually designed features, which not only involve a large amount of work but also have low efficiency. Especially when dealing with high-noise and low-resolution images, their performance is severely affected. In addition, it is difficult for these methods to achieve high accuracy and high robustness when dealing with small ship targets and complex backgrounds in SAR images.

[0004] With the rapid development of deep learning technology, especially in the field of computer vision, methods based on convolutional neural networks (CNNs) have shown significant advantages in image recognition tasks. In order to further optimize the performance, the application of attention mechanism has been proven to be an effective strategy. Mimicking the attention mechanism of human vision, this method optimizes the feature extraction and utilization process by dynamically focusing on key areas in the image and ignoring irrelevant information. Especially when processing small ship targets in SAR images, the attention mechanism can enhance the model's sensitivity to target features and robustness to complex backgrounds, significantly improving detection accuracy. Therefore, integrating the attention mechanism into the ship recognition task of SAR images is a very promising research direction. At the same time, in the ship target recognition task of SAR images, convolutional neural networks still face difficulties caused by image quality, such as low resolution and noise problems, which will seriously weaken the effect of feature learning, and then affect the recognition performance, making the ship target recognition instrument have low recognition accuracy and weak anti-interference ability. Summary of the invention

[0005] In order to overcome the shortcomings of the existing ship target recognition instruments, such as low recognition accuracy, no effective use of attention mechanism, and weak anti-interference ability, the present invention aims to provide a ship target recognition instrument based on a dense connection mechanism.

[0006] The technical solution adopted by the present invention to solve its technical problem is:

[0007] A ship target recognition device based on dense connection mechanism is used for ship target recognition and classification of SAR images, including a data preprocessing module, a feature extraction module and a classification and recognition module;

[0008] The data preprocessing module is used to segment the image foreground and background by using multiple thresholds;

[0009] The feature extraction module is used to read the image after segmentation processing in the data preprocessing module, and output a set of feature values ​​of the image after passing through the initial convolution layer, the densely connected convolution layer with the attention mechanism, and the final convolution layer;

[0010] The classification recognition module is used to convert the feature map into classification output.

[0011] Furthermore, the data preprocessing module uses multiple thresholds to segment the image foreground and background, wherein the values ​​of the multiple thresholds are determined using a swarm intelligence optimization method. The module reads the SAR image to be identified and calculates its histogram as follows:

[0012]

[0013] Where k represents the gray level, x and y represent the image positions, H(k) represents the frequency of gray level k, and Igray (x, y) represents the gray value at position (x, y), W and H represent the width and height of the image respectively, and 1(·) represents the indicator function, which takes the value of 1 when the internal condition is true and 0 otherwise. The value range of k is (0, 255), and its average gray value is:

[0014]

[0015] Among them, represents the average gray value of the image;

[0016] Set the population size to N, and the initial population can be represented as a set of chromosome collections P 0 ={c 1 , c 2 ,..., c N}, where each chromosome c i is a candidate solution, and each chromosome c i is composed of multiple genes, which are used to represent the threshold of image segmentation. If m thresholds are considered, the chromosome can be expressed as a binary string c i =b i1 b i2 …b im×8 , where every 8-bit binary number corresponds to a threshold, so the total length is m×8 bits. For each gene b i in each chromosome c ij , its value (0 or 1) is randomly generated to form a random binary string, which reflects the random exploration starting point of the genetic algorithm in the search space. The binary string in each chromosome c i needs to be converted into a specific threshold set for subsequent image segmentation. Let T i be the threshold set mapped from chromosome c i , then the conversion process can be expressed as:

[0017] T i ={t i1 , t i2 ,…, t im} (3);

[0018] Among them, T i represents the threshold set mapped from chromosome c i , and t ik =int(b i(k-1)×8+1 b i(k-1)×8+2 …b ik×8 , 2), which means converting every 8-bit binary number into a decimal number;

[0019] The fitness function uses the between-class variance to evaluate the effect of image segmentation. For a given chromosome c i , the corresponding threshold set T i = {t i1 , t i2 , …, t im} divides the gray-scale range of the image into m + 1 intervals. Each threshold t ik divides the gray-scale range (0, 255) into multiple consecutive intervals, which are used to calculate the pixel weights, average gray-scale values, and contributions to the total between-class variance of each interval. The fitness function F(c i ) is the between-class variance of the i-th individual in the population and is calculated as follows:

[0020]

[0021] where w j represents the weight of the pixels in the j-th interval, which is calculated as the number of pixels in the interval divided by the total number of pixels. μ j represents the average gray-scale value of the j-th interval, and μ T represents the global average gray-scale value of the image. The between-class variance is the weighted squared difference between the interval average gray-scale value and the global average gray-scale value, which reflects the quality of the segmentation effect. The larger the between-class variance, the greater the difference between different intervals, and the better the segmentation effect;

[0022] The initial population evolves generation by generation through a natural genetic mechanism that simulates selection, crossover, and mutation, in the hope of finding the optimal solution to the problem. Calculate its fitness ratio. For each individual c i in the population, its selection probability P(c i ) is proportional to its fitness F(c i ):

[0023]

[0024] where N represents the size of the population, and c k represents traversing all individuals in the population;

[0025] Calculate the cumulative probability C i for each individual, which is used in the actual selection process:

[0026]

[0027] where c j represents all individuals from 1 to i in the population;

[0028] The selection process is to generate a random number r in the range (0, 1) and select the one that satisfies r ≤ C iThe smallest i, select the corresponding individual c i into the next-generation population. Randomly select a position l as the crossover point, where l is within (1, m×8 - 1). The crossover process is as follows: for the two selected individuals c i and c j , exchange their gene sequences after the crossover point l to generate two new individuals. The mutation process is as follows: for each gene position in each individual c i , with a very small probability p m perform an inversion, that is, change 0 to 1 and 1 to 0. Through the above genetic operations, the population will continuously evolve. Each generation of the population is generated based on the previous generation through natural selection, genetic crossover, and gene mutation until the preset maximum number of iterations L is reached, and the optimal threshold combination is obtained and the thresholds have been sorted according to their magnitudes, that is

[0029] Use the obtained optimal threshold combination to perform pixel segmentation to obtain the preprocessed image:

[0030]

[0031] where, I seg (x, y) represents the pixel value at position (x, y) in the segmented image, r j represents the target gray value of the pixel points in the jth interval, and represent the start and end thresholds of each pixel interval. Define and Using the above method to optimize the minimum gray value can adaptively optimize the threshold selection in the image segmentation process, overcoming the limitations of traditional fixed thresholds or single segmentation preprocessing methods. The method proposed in this paper can effectively handle the dynamic changes in complex scenes, improve the accuracy and robustness of image segmentation, and especially show stronger adaptability and accuracy in the case of more noise or blurred edges of ship targets

[0032] Furthermore, the feature extraction module reads the segmented image in the data preprocessing module, and after passing through the initial convolutional layer, the dense connection convolutional layer with an attention mechanism, and the ending convolutional layer respectively, outputs a set of feature values of the image, which is completed by the following process:

[0033] (1) The initial convolutional layer is used to process the input image and perform preliminary feature extraction. The initial convolutional layer consists of a two-dimensional convolutional layer, a batch normalization layer, an activation function layer, and a max pooling layer;

[0034] (1.1) The input channel number of the two-dimensional convolutional layer is 1, the output channel number is 64, the convolutional kernel size is 7x7, the stride is 2, and the padding is 3;

[0035] (1.2) The batch normalization layer performs batch normalization on 64 output channels, making the mean of the output close to 0 and the standard deviation close to 1. This helps to stabilize and accelerate the training of the neural network, while reducing the sensitivity of the model to the initial weights. For a given feature a, the output b of batch normalization can be expressed as:

[0036]

[0037] where μ B and are the mean and variance of the data in the mini-batch B respectively, γ and β are learnable parameters used to restore the scale and offset of the normalized data, and ò is a number infinitely close to 0 to prevent the denominator from being zero;

[0038] (1.3) The activation function layer uses ReLU as the activation function, which can accelerate convergence and reduce the problem of gradient vanishing;

[0039] (1.4) The max pooling layer reduces the feature dimension by taking the maximum value in the local area of the input feature map;

[0040] (2) The densely connected convolutional layer is composed of 4 densely connected blocks and 3 smooth connected blocks arranged alternately, and an attention mechanism is added to the densely connected blocks to improve the utilization rate of important features;

[0041] (2.1) Each densely connected block consists of 4 consecutive convolutional blocks, and each convolutional block includes: a batch normalization layer; an activation function layer with ReLU as the activation function; a two-dimensional convolutional layer with a convolution kernel of 3x3, 32 output channels, and a padding of 1, and an attention module;

[0042] The attention module accepts a feature map with dimensions (B, C, H, W), where B is the batch size, C is the number of channels, and H and W are the height and width of the feature map;

[0043] The feature map is input into the max pooling layer and the average pooling layer. Through the pooling operation, each channel is compressed into a single value, thereby highlighting the global statistical features of each channel. Here, the pooling window size is 1, which actually does not change the size, but only calculates the maximum value and the average value respectively;

[0044] The pooled feature map enters a fully connected layer composed of two convolutional layers. The first convolutional layer reduces the number of channels to 1 / 16 of the original to play a role in dimensionality reduction, followed by the ReLU activation function. The second convolutional layer then restores the number of channels. This process can capture the complex dependencies between channels while reducing the number of parameters and the computational cost;

[0045] The output of the fully connected layer enters the Sigmoid activation function layer and is converted into weight coefficients between 0 and 1 through the Sigmoid activation function. These weight coefficients are used to modulate the importance of each channel;

[0046] The output of the Sigmoid activation function layer enters the maximum projection layer and the average projection layer, which perform maximum and average projections on the input feature map along the channel dimension respectively, obtaining two feature maps of dimension (B, 1, H, W), representing the maximum and average information at each position respectively. This helps to extract important spatial positions;

[0047] After the two projected feature maps are concatenated, they enter a convolutional layer. This convolutional layer uses a 7x7 convolutional kernel to extract spatial features while keeping the spatial size of the feature map unchanged. This convolutional operation can extract and fuse features of local regions, enhancing the model's attention to spatial positions;

[0048] The output of the convolutional layer enters the Sigmoid activation function layer to obtain the final feature map output by the attention module;

[0049] The convolutional blocks in the dense connection block are stacked in a way of feature reuse. The specific stacking method is as follows:

[0050] Suppose the input feature map of the dense connection block has c channels. The input of the first convolutional block is the feature map of these c channels. After being processed by the first convolutional block, the number of output feature maps is the growth rate k. The output feature maps of these k channels will be concatenated with the original input feature map of c channels in the channel dimension. Therefore, the number of input feature maps of the second convolutional block becomes c + k. For each subsequent convolutional block, its input will include the output feature maps of all previous convolutional blocks and the original input feature map of the dense connection block. This means that if the dense connection block has n convolutional blocks, the number of input feature maps of the nth convolutional block will be c + (n - 1) × k; finally, the output of the dense connection block is the concatenation result of the output feature map of the last convolutional block, the output feature maps of all previous convolutional blocks, and the original input feature map of the dense connection block in the channel dimension. This ensures that each layer in the deep part of the network can directly access the original input features and the output features of all previous layers;

[0051] (2.2) The smooth connection block includes: a batch normalization layer; an activation function layer with ReLU as the activation function; a two-dimensional convolutional layer with a 1x1 convolutional kernel and the number of output channels being half of the number of input channels; an average pooling layer using a 2x2 convolutional kernel with a stride of 2;

[0052] The smooth connection block is responsible for connecting two adjacent dense connection blocks and controlling the complexity of the model by adjusting the size and number of channels of the feature map;

[0053] (3) The final convolutional layer receives the output of the last dense connection block as input, which will sequentially enter a batch normalization layer, an activation function layer with ReLU as the activation function, an adaptive average pooling layer with an output size of 1x1, and a flattening layer that converts the feature map into a one-dimensional vector. By adopting a feature extraction module based on the dense connection mechanism and combining the spatial attention and channel attention mechanisms, the present invention can automatically extract key features with higher abstraction in the image. This method effectively overcomes the problem of low classification accuracy caused by the lack of saliency in the feature extraction process in traditional image recognition technologies.

[0054] Furthermore, the classification and recognition module is a fully connected layer with an input size the same as the output size of the feature extraction module and an output size the same as the number of target classification categories, and is used to convert the feature map into a classification output.

[0055] The beneficial effects of the present invention are mainly manifested in:

[0056] 1. Innovatively adopting a swarm intelligence optimization method for multi-threshold segmentation of ship images, effectively dealing with dynamic changes in complex scenes, and improving the accuracy and robustness of the target recognition instrument;

[0057] 2. Innovatively introducing spatial and channel attention mechanisms, enabling the target recognition instrument to automatically extract and make full use of important features of the image

[0058] 3. The target recognition instrument can classify and recognize ship targets in real time according to image information, with high classification recognition accuracy, high speed, and strong anti-interference ability. BRIEF DESCRIPTION OF THE DRAWINGS

[0059] Figure 1 is a schematic structural diagram of the instrument proposed by the present invention;

[0060] Figure 2 is a schematic structural diagram of the feature extraction module of the present invention;

[0061] Figure 3 is a schematic structural diagram of the dense connection convolutional layer of the present invention;

[0062] Figure 4 is a schematic diagram of the dense connection block of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0063] The following further describes the present invention with reference to the drawings.

[0064] Refer to Figure 1, A ship target recognition instrument based on a dense connection mechanism, which is used to recognize and classify ship targets in SAR images, including three modules: a data preprocessing module 11, a feature extraction module 12, and a classification and recognition module 13; among them:

[0065] (1) The data preprocessing module 11 segments the foreground and background of the image with multiple thresholds, and the group intelligence optimization method is used to confirm the values of the multiple thresholds.

[0066] The module reads the SAR image to be recognized and calculates its histogram as:

[0067]

[0068] where k represents the gray level, x and y represent the image positions, H(k) represents the frequency of the gray level k, I gray (x, y) represents the gray value at the position (x, y), W and H respectively represent the width and height of the image, and 1(·) represents the indicator function, which takes the value of 1 when the internal condition is true and 0 otherwise. The value range of k is (0, 255). Its average gray value is:

[0069]

[0070] where, represents the average gray value of the image.

[0071] Set the population size to N, and the initial population can be represented as a set of chromosome collections P 0 ={c 1 , c 2 , …, c N}, where each chromosome c i is a candidate solution. Each chromosome c i is composed of multiple genes, which are used to represent the thresholds of image segmentation. If m thresholds are considered, the chromosome can be expressed as a binary string c i =b i1 b i2 …b im×8 , where every 8-bit binary number corresponds to a threshold, so the total length is m×8 bits. For each gene b i in each chromosome c ij , its value (0 or 1) is randomly generated to form a random binary string. This reflects the random exploration starting point of the genetic algorithm in the search space. The binary string in each chromosome c i needs to be converted into a specific threshold set for subsequent image segmentation. Let T i be the threshold set mapped from the chromosome c i , then the conversion process can be expressed as:

[0072] T i = {t i1 , t i2 , …, t im} (3);

[0073] Among them, T i represents the threshold set mapped from chromosome c i , and t ik = int(b i(k-1)×8+1 b i(k-1)×8+2 … b ik×8 , 2), which means converting every 8-bit binary number into a decimal number.

[0074] The fitness function uses the between-class variance to evaluate the effect of image segmentation. For the given chromosome c i , its corresponding threshold set T i = {t i1 , t i2 , …, t im} divides the gray level range of the image into m + 1 intervals. Each threshold t ik divides the gray level range (0, 255) into multiple continuous intervals. For example, if the threshold set is {t i1 , t i2}, then the intervals are divided into ([0, t i1 ), (t i1 , t i2 ), (t i2 , 255]). These intervals are used to calculate the pixel weights, average gray values, and contributions to the total between-class variance of each respective interval. The fitness function F(c i ) is the between-class variance of the i-th individual in the population, and is calculated as follows:

[0075]

[0076] Among them, w j represents the weight of the pixels in the j-th interval, and is calculated by dividing the number of pixels in the interval by the total number of pixels. μ j represents the average gray value of the j-th interval. μ T represents the global average gray value of the image. The between-class variance is the weighted square difference between the average gray value of the interval and the global average gray value, reflecting the quality of the segmentation effect. The larger the between-class variance, the greater the difference between different intervals, and the better the segmentation effect. The calculation method of the between-class variance provides a quantitative evaluation index for the quality of the segmentation effect.

[0077] The initial population evolves generation by generation through a natural genetic mechanism that simulates selection, crossover, and mutation, with the expectation of finding the optimal solution to the problem. Calculate its fitness proportion. For each individual c in the population i , its selection probability P(c i ) is proportional to its fitness F(c i ):

[0078]

[0079] where N is the size of the population, and c k represents traversing all individuals in the population;

[0080] Calculate the cumulative probability C i for each individual, which is used in the actual selection process:

[0081]

[0082] where c j represents all individuals from 1 to i in the population;

[0083] The selection process is to generate a random number r in the range (0, 1) and select the smallest i that satisfies r ≤ C i , and select the corresponding individual c i into the next-generation population. Randomly select a position l as the crossover point, where l is within (1, m×8 - 1). The crossover process is for the two selected individuals c i and c j , swap their gene sequences after the crossover point l to generate two new individuals. The mutation process is for each gene position in each individual c i , with a very small probability p m to reverse it, that is, 0 becomes 1 and 1 becomes 0. Through the above genetic operations, the population will continuously evolve. Each generation of the population is generated based on the previous generation through natural selection, genetic crossover, and gene mutation until the preset maximum number of iterations L is reached, obtaining the optimal threshold combination and the thresholds have been sorted according to size, that is In the selection process, always select the smallest individual that meets the conditions to ensure that each generation of the population contains individuals with higher fitness, thereby promoting the evolution and optimization of the population, and finally making the between-class variance of the demand reach the optimal value.

[0084] Use the obtained optimal threshold combination to perform pixel segmentation to obtain the preprocessed image:

[0085]

[0086] where I seg(x, y) represents the pixel value at position (x, y) in the segmented image, and r j represents the target gray value of the pixel points within the j-th interval, and represent the start and end thresholds of each pixel interval. Define and Using the above method to optimize the minimum gray value can adaptively optimize the threshold selection in the image segmentation process, overcoming the limitations of traditional fixed thresholds or single segmentation preprocessing methods. The method proposed in this paper can effectively handle the dynamic changes in complex scenes, improve the accuracy and robustness of image segmentation, and especially show stronger adaptability and accuracy in the case of more noise or blurred edges of ship targets.

[0087] (2) Referring to Figure 2 , the feature extraction module 12 reads the segmented image in the data preprocessing module. After passing through the initial convolutional layer 21, the dense connected convolutional layer 22 with an attention mechanism, and the ending convolutional layer 23 respectively, a set of feature values of the image is output. The following structures all have forward propagation functions, which enables the model to dynamically adjust and optimize its parameters based on the training results during the training process.

[0088] (2.1) The initial convolutional layer 21 is used to process the input image and perform preliminary feature extraction. The initial convolutional layer 21 consists of a two-dimensional convolutional layer, a batch normalization layer, an activation function layer, and a max pooling layer.

[0089] The input channel number of the two-dimensional convolutional layer is 1, the output channel number is 64, the convolutional kernel size is 7x7, the stride is 2, and the padding is 3.

[0090] The batch normalization layer performs batch normalization on 64 output channels, making the output mean close to 0 and the standard deviation close to 1. This helps to stabilize and accelerate the training of the neural network, while reducing the sensitivity of the model to the initial weights. For a given feature a, the output b of batch normalization can be expressed as:

[0091]

[0092] where μ B and are the mean and variance of the data in the mini-batch B respectively, γ and β are learnable parameters used to restore the scale and offset of the normalized data, and ò is a very small number to prevent the denominator from being zero. Batch normalization helps to accelerate the training process of the network and can improve the stability and generalization ability of the network.

[0093] The activation function layer uses ReLU as the activation function, which can accelerate convergence and reduce the problem of gradient disappearance.

[0094] The max pooling layer reduces the feature dimension by taking the maximum value within the local region of the input feature map.

[0095] (2.2) Refer to Figure 3 , the densely connected convolutional layer 22 is composed of 4 densely connected blocks and 3 smooth connection blocks arranged alternately, and an attention mechanism is added to the densely connected blocks to improve the utilization rate of important features.

[0096] Each densely connected block consists of 4 consecutive convolutional blocks, where each convolutional block includes a batch normalization layer; an activation function layer with ReLU as the activation function; a two-dimensional convolutional layer with a convolutional kernel of 3x3, an output channel number of 32, and a padding of 1, and an attention module.

[0097] The attention module accepts a feature map with dimensions (B, C, H, W), where B is the batch size, C is the number of channels, and H and W are the height and width of the feature map.

[0098] The feature map is input into the max pooling layer and the average pooling layer. Through the pooling operation, each channel is compressed into a single value, thereby highlighting the global statistical features of each channel. Here, the pooling window size is 1, which actually does not change the size, but only calculates the maximum value and the average value respectively. The purpose of pooling is to reduce the amount of data, lower the computational complexity, and help extract the key features of the input data.

[0099] The pooled feature map enters the fully connected layer composed of two convolutional layers. The first convolutional layer reduces the number of channels to 1 / 16 of the original to play a role in dimensionality reduction, followed by the ReLU activation function. The second convolutional layer then restores the number of channels. This process can capture the complex dependencies between channels while reducing the number of parameters and the computational cost.

[0100] The output of the fully connected layer enters the Sigmoid activation function layer and is converted into weight coefficients between 0 and 1 through the Sigmoid activation function. These weight coefficients are used to modulate the importance of each channel.

[0101] The output of the Sigmoid activation function layer enters the max projection layer and the average projection layer, and performs max and average projections on the input feature map along the channel dimension to obtain two feature maps with dimensions (B, 1, H, W), which represent the max and average values at each position respectively. This helps to extract important spatial positions.

[0102] After the two projected feature maps are concatenated, they enter a convolutional layer. This convolutional layer uses a 7x7 convolutional kernel to extract spatial features while keeping the spatial size of the feature map unchanged. This convolutional operation can extract and fuse the features of the local region, enhancing the model's attention to spatial positions.

[0103] The output of the convolutional layer enters the Sigmoid activation function layer to obtain the final feature map output by the attention module. Through multi-level processing in the attention module, the model can more effectively learn and utilize the feature information in the image, thereby improving the performance and generalization ability of the model.

[0104] Refer to Figure 4 , the convolutional blocks in the dense connection block are stacked in a way of feature reuse, and the specific stacking method is as follows:

[0105] Assume that the input feature map of the dense connection block has c channels. The input of the first convolutional block is the feature map of these c channels. After being processed by the first convolutional block, the number of output feature maps is the growth rate k. The output feature maps of these k channels will be concatenated with the original input feature map of c channels in the channel dimension. Therefore, the number of input feature maps of the second convolutional block becomes c + k. For each subsequent convolutional block, its input will include the output feature maps of all previous convolutional blocks and the original input feature map of the dense connection block. This means that if the dense connection block has n convolutional blocks, the number of input feature maps of the nth convolutional block will be c + (n - 1) × k. Finally, the output of the dense connection block is the concatenation result of the output feature map of the last convolutional block, the output feature maps of all previous convolutional blocks, and the original input feature map of the dense connection block in the channel dimension. This ensures that each layer in the deep part of the network can directly access the original input features and the output features of all previous layers. Dense connections help to promote the smooth flow of model information, the smooth propagation of gradients, feature reuse, and network compression to build a more efficient, stable, and expressive neural network model.

[0106] The smooth connection block consists of a batch normalization layer; an activation function layer with ReLU as the activation function; a two-dimensional convolutional layer with a convolution kernel of 1x1 and the number of output channels being half of the number of input channels; an average pooling layer using a 2x2 convolution kernel with a stride of 2. The smooth connection block is responsible for connecting two adjacent dense connection blocks and controlling the complexity of the model by adjusting the size and number of channels of the feature map.

[0107] (2.3) The ending convolutional layer 23 receives the output of the last dense connection block as input, and this input will sequentially enter a batch normalization layer; an activation function layer with ReLU as the activation function; an adaptive average pooling layer with an output size of 1x1; a flattening processing layer that converts the feature map into a one-dimensional vector.

[0108] (3) The classification and recognition module 13 is a fully connected layer with the same input size as the output size of the feature extraction module and the same output size as the number of target classification categories. It is used to convert the feature map into a classification output.

[0109] (4) When the SAR images in the training set are input into the instrument, they generate output results according to the current parameter configuration of the instrument. Subsequently, the instrument optimizes its parameters through the forward propagation process to improve performance and accuracy. In contrast, after the SAR images in the test set are input into the instrument, they sequentially pass through each module, but are only used to generate recognition results without involving the parameter optimization process.

[0110] The above embodiments are used to explain the present invention rather than limit the present invention. Any modifications and changes made to the present invention within the spirit and scope of the claims of the present invention fall within the protection scope of the present invention.

Claims

1. A ship target recognition device based on dense connection mechanism, characterized in that: Used to identify and classify ship targets in SAR images, including data preprocessing module, feature extraction module and classification and identification module; The data preprocessing module is used to segment the image foreground and background by using multiple thresholds; The feature extraction module is used to read the image after segmentation processing in the data preprocessing module, and output a set of feature values ​​of the image after passing through the initial convolution layer, the densely connected convolution layer with the attention mechanism, and the final convolution layer; The classification recognition module is used to convert the feature map into classification output.

2. The ship target identification device based on dense connection mechanism according to claim 1 is characterized in that: The data preprocessing module uses multiple thresholds to segment the image foreground and background, wherein the swarm intelligence optimization method is used to determine the value of the multiple thresholds. The module reads the SAR image to be identified and calculates its histogram as follows: Where k represents the gray level, x and y represent the image positions, H(k) represents the frequency of gray level k, and I gray (x, y) represents the grayscale value at the position (x, y), W and H represent the width and height of the image respectively, and 1(·) represents the indicator function, which takes the value 1 when the internal condition is true, otherwise it takes the value 0. The value range of k is (0, 255), and its average grayscale value is: in, Represents the average gray value of the image; Assuming the population size is N, the initial population can be represented as a set of chromosomes P0 = {c1, c2, ..., c N }, where each chromosome c i is a candidate solution, each chromosome c i It is composed of multiple genes and is used to represent the threshold of image segmentation. If m thresholds are considered, the chromosome can be expressed as a binary string c i =b i1 b i2 ...b im×8 , where each 8-bit binary number corresponds to a threshold, so the total length is m×8 bits. For each chromosome c i Each gene in b ij , randomly generates its value, which is 0 or 1 to form a random binary string, which reflects the random exploration starting point of the genetic algorithm in the search space. Each chromosome c i The binary string in needs to be converted into a specific threshold set for subsequent image segmentation. Let T i From chromosome c i The threshold set obtained by mapping, the conversion process can be expressed as: T i ={t i1 ,t i2 ,…,t im } (3); Among them, T i From chromosome c i The threshold set obtained by mapping, t ik = int(b i(k-1)×8+1 b i(k-1)×8+2 ...b ik×8 ,2) means converting every 8-bit binary number into a decimal number; The fitness function uses the inter-class variance to evaluate the effect of image segmentation. For a given chromosome c i , and its corresponding threshold set T i ={t i1 ,t i2 ,...,t im } Divide the grayscale range of the image into m+1 intervals, each threshold t ik The grayscale range (0, 255) is divided into multiple continuous intervals, which are used to calculate the respective pixel weights, average grayscale values, and contributions to the total inter-class variance. The fitness function F(c i ) is the inter-class variance of the i-th individual in the population, calculated as follows: Among them, w j Represents the weight of the pixel in the jth interval, which is calculated by dividing the number of pixels in the interval by the total number of pixels, μ j represents the average gray value of the jth interval, μ T Represents the global average gray value of the image. The inter-class variance is the weighted square difference between the interval average gray value and the global average gray value, which reflects the quality of the segmentation effect. The larger the inter-class variance, the greater the difference between different intervals, and the better the segmentation effect. The initial population evolves from generation to generation by simulating the natural genetic mechanism of selection, crossover and mutation, hoping to find the optimal solution to the problem and calculate its fitness ratio. For each individual c in the population i , its selection probability P(c i ) and its fitness F(c i ) is proportional to: Where N represents the size of the population, c k Indicates traversing all individuals in the population; Calculate the cumulative probability C for each individual i , used for the actual selection process: Among them, c j represents all individuals from 1 to i in the population; The selection process is to generate a random number r in the range of (0,1) and select a random number r≤C. i The smallest i will correspond to the individual c i Select into the next generation population, randomly select a position l as the crossover point, l is within (1,m×8-1), and the crossover process is for the two selected individuals c i and c j , exchange their gene sequences after the crossover point l to generate two new individuals. The mutation process is for each individual c i Each gene position in m Reverse, that is, 0 becomes 1, 1 becomes 0. Through the above genetic operations, the population will continue to evolve. Each generation of population is generated based on the previous generation through natural selection, genetic crossover and gene mutation, until the preset maximum number of iterations L is reached, and the optimal threshold combination is obtained. And the thresholds have been sorted according to size, that is, Use the obtained optimal threshold combination to perform pixel segmentation to obtain the preprocessed image: Among them, I seg (x, y) represents the pixel value at position (x, y) in the segmented image, r j represents the target grayscale value of the pixel in the jth interval, and Represents the start and end thresholds of each pixel interval, and defines and Use the above method to find the minimum gray value.

3. The ship target identification device based on dense connection mechanism according to claim 1 is characterized in that: The feature extraction module reads the image after segmentation processing in the data preprocessing module, and outputs a set of feature values ​​of the image after passing through the initial convolution layer, the densely connected convolution layer with the attention mechanism, and the final convolution layer, respectively. The following process is used to complete it: (1) The initial convolution layer is used to process the input image and perform preliminary feature extraction. The initial convolution layer consists of a two-dimensional convolution layer, a batch normalization layer, an activation function layer, and a maximum pooling layer. (1.1) The number of input channels of the two-dimensional convolutional layer is 1, the number of output channels is 64, the convolution kernel size is 7x7, the stride is 2, and the padding is 3; (1.2) The batch normalization layer batch normalizes the 64 output channels so that the mean of the output is close to 0 and the standard deviation is close to 1. This helps stabilize and accelerate the training of the neural network while reducing the model's sensitivity to the initial weights. For a given feature a, the batch normalized output b can be expressed as: Among them, μ B and are the mean and variance of the data in the mini-batch B, γ and β are learnable parameters used to restore the scale and offset of the normalized data, ò is a number infinitely close to 0 to prevent the denominator from being zero; (1.3) The activation function layer uses ReLU as the activation function, which can accelerate convergence and reduce the problem of gradient disappearance; (1.4) The maximum pooling layer reduces the feature dimension by taking the maximum value in the local area of ​​the input feature map; (2) The densely connected convolutional layer consists of four densely connected blocks and three smoothly connected blocks arranged alternately. An attention mechanism is added to the densely connected blocks to improve the utilization of important features. (2.1) Each densely connected block consists of 4 consecutive convolutional blocks, each of which includes: a batch normalization layer; an activation function layer with ReLU as the activation function; a 2D convolutional layer with a 3x3 kernel, 32 output channels, and 1 padding, and an attention module; The attention module accepts a feature map of dimension (B,C,H,W), where B is the batch size, C is the number of channels, and H and W are the height and width of the feature map; The feature map is input into the maximum pooling layer and the average pooling layer. Through the pooling operation, each channel is compressed into a single value, thereby highlighting the global statistical characteristics of each channel. The pooling window size here is 1, which does not actually change the size, but only calculates the maximum value and average value respectively; The pooled feature map enters a fully connected layer consisting of two convolutional layers. The first convolutional layer reduces the number of channels to 1 / 16 of the original to reduce the dimension, followed by a ReLU activation function. The second convolutional layer restores the number of channels. This process can capture the complex dependencies between channels while reducing the number of parameters and computational cost. The output of the fully connected layer enters the Sigmoid activation function layer and is converted into a weight coefficient between 0 and 1 through the Sigmoid activation function. These weight coefficients are used to modulate the importance of each channel. The output of the Sigmoid activation function layer enters the maximum projection layer and the average projection layer, and the input feature map is projected along the channel dimension to obtain two (B, 1, H, W) dimensional feature maps, which represent the maximum and average information at each position, respectively, which helps to extract important spatial positions; The two projected feature maps are concatenated and then enter a convolution layer, which uses a 7x7 convolution kernel to extract spatial features while keeping the spatial size of the feature map unchanged. This convolution operation can extract and fuse the features of the local area, enhancing the model's attention to the spatial position. The output of the convolutional layer enters the Sigmoid activation function layer to obtain the final feature map output by the attention module; The convolution blocks in the densely connected blocks are stacked in a feature reuse manner, and the specific stacking method is as follows: Assume that the input feature map of the dense connection block has c channels, and the input of the first convolution block is the feature map of these c channels. After being processed by the first convolution block, the number of output feature maps is the growth rate k. The output feature maps of these k channels will be concatenated with the original input feature maps of the c channels in the channel dimension. Therefore, the number of input feature maps of the second convolution block becomes c+k. For each subsequent convolution block, its input will include the output feature maps of all previous convolution blocks and the original input feature maps of the dense connection block. This means that if the dense connection block has n convolution blocks, the number of input feature maps of the nth convolution block will be c+(n-1)×k; finally, the output of the dense connection block is the concatenation of the output feature map of the last convolution block, the output feature maps of all previous convolution blocks, and the original input feature map of the dense connection block in the channel dimension, which ensures that each layer in the deep layer of the network can directly access the original input features and the output features of all previous layers; (2.2) The smooth connection block includes: a batch normalization layer; an activation function layer with ReLU as the activation function; a two-dimensional convolution layer with a convolution kernel of 1x1 and the number of output channels being half the number of input channels; an average pooling layer with a 2x2 convolution kernel and a stride of 2 for average pooling; The smooth connection block is responsible for connecting two adjacent dense connection blocks and controlling the complexity of the model by adjusting the size and number of channels of the feature map; (3) The final convolutional layer receives the output of the last densely connected block as input, which in turn enters a batch normalization layer; an activation function layer with ReLU as the activation function; an adaptive average pooling layer with an output size of 1x1; and a flattening layer that converts the feature map into a one-dimensional vector.

4. The ship target identification device based on dense connection mechanism according to claim 1 is characterized in that: The classification recognition module is a fully connected layer whose input size is the same as the output size of the feature extraction module and whose output size is the same as the target classification type, and is used to convert the feature map into a classification output.

Citation Information

Patent Citations

  • Cable insulation defect state assessment method based on high frequency local discharge signal graph

    CN107067041A

  • SAR image ship target identification method based on attention mechanism

    CN117218612A