Marine vessel detection method based on feature extraction and feature weighting selection fusion

By dynamically learning the sampling point offset and modulation factor to adjust the convolution kernel, and combining it with a hierarchical scale feature pyramid network, the problem of morphological differences of ship targets in remote sensing images is solved, and the detection accuracy and performance are improved.

CN122265870APending Publication Date: 2026-06-23SANYA SCI & EDUCATION INNOVATION PARK WUHAN UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SANYA SCI & EDUCATION INNOVATION PARK WUHAN UNIV OF TECH
Filing Date
2026-05-25
Publication Date
2026-06-23

AI Technical Summary

Technical Problem

The morphological differences and multi-scale features of ship targets in remote sensing images are difficult to process, which makes it easy for the model to miss or misdetect during feature fusion, thus affecting the detection accuracy.

Method used

We employ dynamic learning of sampling point offsets and modulation factors to adjust convolutional kernels, combined with a hierarchical feature pyramid network to enhance feature representation capabilities, and perform ship detection through feature extraction and weighted selection fusion methods.

Benefits of technology

It improved the positioning and regression accuracy of ship detection, enhanced the model's feature representation ability, and improved detection performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122265870A_ABST
    Figure CN122265870A_ABST
Patent Text Reader

Abstract

The application discloses a marine ship detection method based on feature extraction and feature weighting selection fusion. First, remote sensing ship images are collected to construct a data set. Then, a ship detection model is constructed, the remote sensing ship images are input into the model to obtain fusion features, and then the fusion features are input into a detection head of the model for detection to obtain a detection result. Then, the model is trained, the remote sensing ship images are input into the model, and after multiple rounds of training, a final model is obtained. Finally, the remote sensing ship images are input into the trained model to output ship types and positioning information. The application utilizes dynamic learning sampling point offset and modulation factors to enable convolution kernels to change according to ship shapes and enhance feature information. In the feature fusion stage, feature pyramid networks based on hierarchical scales are developed to fully utilize feature maps from different scales and enhance the feature expression capability of the model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of smart ocean and remote sensing image processing technology, and specifically to a method for detecting ships at sea based on feature extraction and feature weighted selection fusion. Background Technology

[0002] In recent years, deep learning technologies have been widely applied in the field of target detection. Ship target detection, as an important branch of target detection, aims to locate and classify maritime targets. As an important carrier of the marine economy, the trajectory positioning and monitoring of ships are of great significance in areas such as maritime traffic control, prevention of illegal activities, and resource management.

[0003] Due to differences in imaging position and angle, the shape and size of ships appear significantly different in remote sensing images, greatly increasing the difficulty of ship detection. Even ships of the same type and size may appear completely different in remote sensing images due to different imaging backgrounds, resulting in poor feature learning ability of the network during training. Consequently, some features may be overlooked during subsequent feature fusion, leading to missed or false detections. Addressing the differences in ship morphology in remote sensing images and effectively processing multi-scale features to improve the accuracy of localization regression in ship detection is an effective way to improve ship target detection performance and a major challenge currently facing ship target detection. Summary of the Invention

[0004] To address the aforementioned problems, the purpose of this invention is to provide a maritime vessel detection method based on the fusion of feature extraction and feature weighted selection. By dynamically learning the sampling point offset and modulation factor, the convolution kernel can adaptively change according to the shape of the vessel, thereby enhancing the feature information. In the feature fusion stage, feature maps from different scales were fully utilized to develop a hierarchical scale-based feature pyramid network, which enhanced the model's feature representation capabilities.

[0005] The technical solution adopted in this invention: A method for detecting ships at sea based on the fusion of feature extraction and feature weighted selection includes the following steps: S1. Collect remote sensing images of ships and construct a ship detection dataset; S2. Based on adaptive feature extraction and multi-scale feature weighted selection fusion, a ship detection model is constructed. The ship detection model includes a feature extraction network, a feature selection module, a selected feature weighted fusion module, and a detection head. The remote-sensed ship image is input into the feature extraction network to obtain the original multi-scale feature map. Then the original multi-scale feature map The input is fed into the feature selection module for localization enhancement, resulting in a multi-scale enhanced feature map. Then, the multi-scale enhanced feature map is processed. The selected feature weighted fusion module performs feature fusion to obtain a fused feature map, which is then input into the detection head to obtain the detection result. ; S3. Training the model: Input the remote-sensed ship image into the ship detection model, calculate the total loss function value, perform backpropagation, optimize the connection weights through the selected optimizer and corresponding optimizer parameters, and obtain the final ship detection model after training for multiple rounds. S4. Input the remote sensing ship image into the final ship detection model and output the ship type and location information;

[0006] In step S2, the multi-scale enhanced feature map is... The input is fed into the selected feature weighted fusion module to obtain a fused feature map. The specific steps are as follows: The selected feature weighted fusion module includes a feature filtering submodule and a feature fusion submodule; Sb1, Multi-scale Enhanced Feature Map The input feature selection submodule performs global max pooling and global average pooling along the channel dimension to obtain weights for the two channel dimensions, expressed by the formula: ; ;

[0007] In the formula, It is the feature vector obtained after global max pooling. It is the feature vector obtained after global average pooling; Indicated as to Perform a global max pooling operation; Indicated as to Perform global average pooling; Sb2, then used The activation function determines the weight value of each channel, ultimately yielding the weight of each channel, expressed by the formula: ;

[0008] In the formula, The final calculated weights for each channel; Represents a non-linear activation function; Sb3, Combine the obtained channel weights with the multi-scale enhanced feature map. Element-wise multiplication yields the filtered and enhanced feature maps; the formula is expressed as: ;

[0009] Sb4, then used The convolution operation transforms the filtered and enhanced feature maps into a single, more powerful array. The number of channels was uniformly adjusted to 256 dimensions to obtain high-level features. The formula is expressed as: ; Sb5, followed by high-level features The input feature fusion submodule utilizes a transposed convolution with a stride of 2 and a kernel size of 3. High-level characteristics Extending this process yields the transpose feature. ; Sb6, then bilinear interpolation was used to process the transposed features. Perform upsampling or downsampling to match lower-level features With the same dimensions, features are obtained. The low-level features To compare with high-level features Lower-level features; the formula is expressed as: ;

[0010] In the formula, This is represented as a transposed convolution. Represented as bilinear interpolation; Sb7, then use channel attention mechanism to extract features This is converted into corresponding attention weights, which are applied to low-level features while ensuring dimensionality consistency. The low-level features are then filtered, and finally... Characteristics of high-level personnel The fusion process is performed to obtain the aforementioned fusion features.

[0011] Preferably, in step S2, the original multi-scale feature map is used. The multi-scale enhanced feature map is obtained. The specific steps are as follows: Sa1. Use bilinear interpolation in the feature extraction network to process the original multi-scale feature map. Augmentation is performed to obtain the augmented feature map. The formula is expressed as: ;

[0012] In the formula, To introduce the actual sampling position after incorporating a learnable offset, , This is the center position of the current convolution window; The nth offset in the sampling grid of the regular convolution kernel; The learnable two-dimensional spatial offset of the nth sampling point; This is represented as an enumerated feature map on the original multi-scale feature map. The positions of all integral spaces in the space, It is a bilinear interpolation kernel. Represented as middle The value of the point, This is represented as an enumerated feature map on the original multi-scale feature map. The positions of all integration spaces in the interval correspond to the positions of all integration spaces in the interval. The values ​​above; Sa2, subsequently, a weighted mechanism was introduced to augment the feature map. By performing a weighted operation, the multi-scale enhanced feature map is obtained. The formula is expressed as: ;

[0013] In the formula, The weight of each point, For the total number of points, For the first A learnable modulation scalar at each position, ranging from 0 to 1.

[0014] Preferably, in step S3, the total loss function value includes classification loss, bounding box regression loss, and confidence loss; the classification loss is a binary cross-entropy loss for multiple targets, and the total loss function value is calculated using the following formula:

[0015] In the formula, This is the total loss function value. For classifying losses, For bounding box regression loss, For confidence loss, and The weights for different losses are given; the formula for calculating the classification loss is:

[0016] In the formula, For the Sigmoid function, It is the natural logarithm. This represents the probability that a sample is predicted as a positive example, where the sample is an object in the image to be detected; It is the sample integral value; The The calculation formula is:

[0017] In the formula, The predicted value output by the network. These are the nearest neighbor predictions output by the network. For the actual value of the sample, For sample integral values, The integral value of the nearest sample; The The calculation formula is:

[0018] In the formula, It is the true value of the target. It is the predicted value output by the model.

[0019] Preferably, in calculation After the loss, the Probiou method was also used to reduce the loss. The loss is the transformation from a horizontal bounding box (HBB) and a rotated bounding box (OBB) to a Gaussian bounding box (GBB) using the Probiou method. The transformation from HBB to GBB follows the assumption that the target region is a two-dimensional binary region with a uniform probability distribution. The formula for the mean-covariance matrix ∑ of this two-dimensional binary region distribution is:

[0020] In the formula, For the region area; Let be the mean vector of the samples. For sample values, for transpose; The binary region of the horizontal bounding box (HBB) is based on Centered on, high as , width is The mean and covariance matrix of the rectangular region are calculated using the following formulas:

[0021] Then the Parse distance The calculation is performed using the following formula:

[0022] In the formula, This represents the positional difference between the sample bounding box and the bounding box detected by the algorithm. The shape difference between the sample bounding box and the bounding box detected by the algorithm; and The calculation formula is:

[0023] In the formula, The x-coordinate of the center of the predicted bounding box, The vertical coordinate of the center of the predicted bounding box; The x-coordinate of the actual center of the annotation box. The vertical coordinate of the actual center of the annotation box; The variance of the center x-coordinate of the prediction box, The variance of the center ordinate of the prediction box, The covariance of the horizontal and vertical coordinates of the center of the predicted bounding box; Let x be the variance of the x-coordinate of the center of the true bounding box. The variance of the center ordinate of the true bounding box, The covariance of the x and y coordinates of the center of the true bounding box; However, the distance of the Barthel Since it is not a real distance, the triangle inequality cannot be satisfied; therefore, we use... Distance, the formula is:

[0024] In the formula, For Hellinger distance, The predicted bounding box and the target bounding box have the same probability distribution if and only if they are identical. The probability between the two Gaussian boxes (GBB) The calculation formula is as follows: .

[0025] Preferably, in step S2, the original multi-scale feature map Including the original multi-scale feature map Original multi-scale feature map Original multi-scale feature map and the original multi-scale feature map The multi-scale enhanced feature map Including multi-scale enhanced feature maps Multi-scale enhanced feature maps Multi-scale enhanced feature maps and multi-scale enhanced feature maps ; The feature extraction network uses the CSPDarkNet-S network. The feature extraction network constructs an adaptive feature modeling framework by introducing an adaptive feature extraction convolution module. By dynamically learning the sampling point offset and modulation factor, the convolution kernel can adaptively adjust the sampling position and weight distribution according to the geometric characteristics and feature context of the ship target.

[0026] The maritime vessel identification computer based on feature extraction and feature weighted selection fusion includes a memory, a processor, and program instructions stored in the memory for the processor to run. The processor executes the program instructions to implement the steps in the maritime vessel detection method based on feature extraction and feature weighted selection fusion described above.

[0027] A computer-readable storage medium storing a computer program that implements the above-described method for detecting marine vessels based on feature extraction and feature weighted selection fusion.

[0028] Compared with the prior art, the beneficial effects of this invention are as follows: This invention utilizes dynamic learning of sampling point offsets and modulation factors to enable convolutional kernels to adaptively change according to the shape of the ship, thereby enhancing feature information. In the feature fusion stage, it makes full use of feature maps from different scales and develops a feature pyramid network based on hierarchical scale, which enhances the feature representation capability of the model. Attached Figure Description

[0029] Figure 1 This is a flowchart of the marine vessel detection method based on feature extraction and feature weighted selection fusion according to the present invention; Figure 2 This is a structural block diagram of the feature fusion module of the present invention; Figure 3 This is a framework diagram of the marine vessel detection method based on the fusion of feature extraction and feature weighted selection according to the present invention. Detailed Implementation

[0030] The technical solutions of the embodiments of this application will be further described clearly and completely below with reference to the accompanying drawings. It should be noted that the described embodiments are only some embodiments of this application, and not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.

[0031] To make the inventive objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be further described in detail below with reference to the accompanying drawings: In order to better understand the above-mentioned objectives, features, and advantages of this invention, the advantages of this invention will be further illustrated below by comparing the embodiments with the accompanying drawings and specific implementation methods.

[0032] Example 1 This embodiment provides a method for marine vessel detection based on the fusion of feature extraction and feature weighted selection, such as... Figure 1 As shown, it includes the following steps: S1. Collect remote sensing images of ships and construct a ship detection dataset; S2. Based on adaptive feature extraction and multi-scale feature weighted selection fusion, a ship detection model is constructed. The ship detection model includes a feature extraction network, a feature selection module, a selected feature weighted fusion module, and a detection head. The feature extraction network uses the CSPDarkNet-S network. The remote-sensed ship image is input into the feature extraction network to obtain the original multi-scale feature map. The original multi-scale feature map Including the original multi-scale feature map Original multi-scale feature map Original multi-scale feature map and the original multi-scale feature map Then the original multi-scale feature map The input is fed into the feature selection module for localization enhancement, resulting in a multi-scale enhanced feature map. The multi-scale enhanced feature map Including multi-scale enhanced feature maps Multi-scale enhanced feature maps Multi-scale enhanced feature maps and multi-scale enhanced feature maps Then, the multi-scale enhanced feature map is processed. The selected feature weighted fusion module performs feature fusion to obtain a fused feature map, which is then input into the detection head to obtain the detection result. ; In step S2, the multi-scale enhanced feature map is obtained. The specific steps are as follows: Sa1, Using bilinear interpolation on the original multi-scale feature map Augmentation is performed to obtain the augmented feature map. In the adaptive feature extraction convolution module, an offset is applied to a point in the originally fixed regular grid. Augmentation is performed to obtain a more accurate output feature map, i.e., the augmented feature map. Because it is in an irregular Sampling is performed on top and Typically fractional in order, bilinear interpolation is used here for augmentation; the formula is expressed as:

[0033] In the formula, To introduce the actual sampling position after incorporating a learnable offset, , This represents the center position of the current convolution window, for example, (5,5). The nth offset in the sampling grid of the regular convolution kernel is used. For example, starting from position (5,5), move up one grid and to the right one grid, i.e., (1, 1), to position (6,6). For example, the learnable two-dimensional spatial offset of the nth sampling point; for , for , Get the nth convolution sampling point (position), plus , ; Obtain the current sampling points ; This is represented as an enumerated feature map on the original multi-scale feature map. The positions of all integral spaces in the space, It is a bilinear interpolation kernel. Represented as middle The value of the point, This is represented as an enumerated feature map on the original multi-scale feature map. The positions of all integration spaces in the interval correspond to the positions of all integration spaces in the interval. The value above; Sa2, subsequently, a weighted mechanism was introduced to augment the feature map. By performing a weighted operation, the multi-scale enhanced feature map is obtained. The formula is expressed as:

[0034] In the formula, The weight of each point, For the total number of points, For the first A modulated scalar that can be learned at each position, ranging from 0 to 1; and Both are achieved by analyzing the same original multi-scale feature map. This is achieved by applying a single convolutional layer with the same spatial resolution and dilation rate as the current convolutional layer.

[0035] In step S2, the multi-scale enhanced feature map is... The input is fed into the selected feature weighted fusion module to obtain a fused feature map. The specific steps are as follows: The selected feature weighted fusion module includes a feature filtering submodule and a feature fusion submodule; Sb1, Multi-scale Enhanced Feature Map The input feature selection submodule performs global max pooling and global average pooling along the channel dimension to obtain weights for the two channel dimensions, expressed by the formula:

[0036] In the formula, It is the feature vector obtained after global max pooling. It is the feature vector obtained after global average pooling; Indicated as to Perform a global max pooling operation; Indicated as to Perform global average pooling; Sb2, then used The activation function determines the weight value of each channel, ultimately yielding the weight of each channel, expressed by the formula:

[0037] In the formula, The final calculated weights for each channel; Represents a non-linear activation function; Sb3, Combine the obtained channel weights with the multi-scale enhanced feature map. Element-wise multiplication yields the filtered and enhanced feature maps; the formula is expressed as: ; Sb4, then used Convolution operations, since features at different scales typically have different numbers of channels, require a feature selection module to prevent this dimensionality inconsistency from hindering effective feature fusion. The convolution operation will filter and enhance the augmented feature maps. The number of channels was uniformly adjusted to 256 dimensions to obtain high-level features. The formula is expressed as: ; Sb5, followed by high-level features The input feature fusion submodule utilizes a transposed convolution with a stride of 2 and a kernel size of 3. High-level characteristics Extending this process yields the transpose feature. ; Sb6, then bilinear interpolation was used to process the transposed features. Perform upsampling or downsampling to match lower-level features With the same dimensions, features are obtained. The low-level features To compare with high-level features Lower-level features; the formula is expressed as:

[0038] In the formula, This is represented as a transposed convolution. Represented as bilinear interpolation; Sb7, then use channel attention mechanism to extract features This is converted into corresponding attention weights, which are applied to low-level features while ensuring dimensionality consistency. The low-level features are then filtered, and finally... Characteristics of high-level personnel The fusion is performed to obtain the fusion features described above; such as Figure 3 As shown.

[0039] S3. Training the model: Input the remote-sensed ship images into the ship detection model, calculate the total loss function value, perform backpropagation, and optimize the connection weights using the selected optimizer and its parameters (the selected optimizer uses the stochastic gradient descent (SGD) algorithm; the optimizer parameters are set to an initial learning rate of...). With a momentum parameter of 0.9 and a weight decay coefficient of 0.0005, the final ship detection model was obtained after multiple training rounds. The total loss function value includes classification loss, bounding box regression loss, and confidence loss; the classification loss is the binary cross-entropy loss for multiple targets, and the formula for calculating the total loss function value is:

[0040] In the formula, This is the total loss function value. For classifying losses, For bounding box regression loss, For confidence loss, and The weights for different losses are given; the formula for calculating the classification loss is:

[0041] In the formula, For the Sigmoid function, It is the natural logarithm. This represents the probability that a sample is predicted as a positive example, where the sample is an object in the image to be detected; It is the sample integral value; The The calculation formula is:

[0042] In the formula, The predicted value output by the network. These are the nearest neighbor predictions output by the network. For the actual value of the sample, For sample integral values, The integral value of the nearest sample; The The calculation formula is:

[0043] In the formula, It is the true value of the target. It is the predicted value output by the model.

[0044] In calculation After the loss, the Probiou method was also used to reduce the loss. The loss is the transformation from a horizontal bounding box (HBB) and a rotated bounding box (OBB) to a Gaussian bounding box (GBB) using the Probiou method. The transformation from HBB to GBB follows the assumption that the target region is a two-dimensional binary region with a uniform probability distribution. The formula for the mean-covariance matrix ∑ of this two-dimensional binary region distribution is:

[0045] In the formula, For the region area; Let be the mean vector of the samples. For sample values, for transpose; The binary region of the horizontal bounding box (HBB) is based on Centered on, high as , width is The mean and covariance matrix of the rectangular region are calculated using the following formulas:

[0046] Then the Parse distance The calculation is performed using the following formula:

[0047] In the formula, This represents the positional difference between the sample bounding box and the bounding box detected by the algorithm. The shape difference between the sample bounding box and the bounding box detected by the algorithm; and The calculation formula is:

[0048] In the formula, The x-coordinate of the center of the predicted bounding box, The vertical coordinate of the center of the predicted bounding box; The x-coordinate of the actual center of the annotation box. The vertical coordinate of the actual center of the annotation box; The variance of the center x-coordinate of the prediction box, The variance of the center ordinate of the prediction box, The covariance of the horizontal and vertical coordinates of the center of the predicted bounding box; Let x be the variance of the x-coordinate of the center of the true bounding box. The variance of the center ordinate of the true bounding box, The covariance of the x and y coordinates of the center of the true bounding box; However, the distance of the Barthel Since it is not a real distance, the triangle inequality cannot be satisfied; therefore, we use... Distance, the formula is:

[0049] In the formula, For Hellinger distance, The predicted bounding box and the target bounding box have the same probability distribution if and only if they are identical. The probability between the two Gaussian boxes (GBB) The calculation formula is as follows: .

[0050] S4. Input the remote sensing ship image into the final ship detection model and output the ship type and location information.

[0051] The maritime vessel identification computer based on feature extraction and feature weighted selection fusion includes a memory, a processor, and program instructions stored in the memory for the processor to run. The processor executes the program instructions to implement the steps in the maritime vessel detection method based on feature extraction and feature weighted selection fusion described above.

[0052] A computer-readable storage medium storing a computer program that implements the above-described method for detecting marine vessels based on feature extraction and feature weighted selection fusion.

[0053] Although preferred embodiments of this application have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of this application.

[0054] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and variations.

Claims

1. A method for marine vessel detection based on the fusion of feature extraction and feature weighted selection, characterized in that, Includes the following steps: S1. Collect remote sensing images of ships and construct a ship detection dataset; S2. Based on adaptive feature extraction and multi-scale feature weighted selection fusion, a ship detection model is constructed. The ship detection model includes a feature extraction network, a feature selection module, a selected feature weighted fusion module, and a detection head. The remote-sensed ship image is input into the feature extraction network to obtain the original multi-scale feature map. ; Then the original multi-scale feature map The input is fed into the feature selection module for localization enhancement, resulting in a multi-scale enhanced feature map. Then, the multi-scale enhanced feature map is processed. The selected feature weighted fusion module performs feature fusion to obtain a fused feature map, which is then input into the detection head to obtain the detection result. ; S3. Training the model: Input the remote-sensed ship image into the ship detection model, calculate the total loss function value, perform backpropagation, optimize the connection weights through the selected optimizer and corresponding optimizer parameters, and obtain the final ship detection model after training for multiple rounds. S4. Input the remote sensing ship image into the final ship detection model and output the ship type and location information; In step S2, the multi-scale enhanced feature map is... The input is fed into the selected feature weighted fusion module to obtain a fused feature map. The specific steps are as follows: The selected feature weighted fusion module includes a feature filtering submodule and a feature fusion submodule; Sb1, Multi-scale Enhanced Feature Map The input feature selection submodule performs global max pooling and global average pooling along the channel dimension to obtain weights for the two channel dimensions, expressed by the formula: ; ; In the formula, It is the feature vector obtained after global max pooling. It is the feature vector obtained after global average pooling; Indicated as to Perform a global max pooling operation; Indicated as to Perform global average pooling; Sb2, then used The activation function determines the weight value of each channel, ultimately yielding the weight of each channel, expressed by the formula: ; In the formula, The final calculated weights for each channel; Represents a non-linear activation function; Sb3, Combine the obtained channel weights with the multi-scale enhanced feature map. Element-wise multiplication yields the filtered and enhanced feature maps; the formula is expressed as: ; Sb4, then used The convolution operation transforms the filtered and enhanced feature maps into a single, more powerful array. The number of channels was uniformly adjusted to 256 dimensions to obtain high-level features. The formula is expressed as: ; Sb5, followed by high-level features The input feature fusion submodule utilizes a transposed convolution with a stride of 2 and a kernel size of 3. High-level characteristics Extending this process yields the transpose feature. ; Sb6, then bilinear interpolation was used to process the transposed features. Perform upsampling or downsampling to match lower-level features With the same dimensions, features are obtained. The low-level features To compare with high-level features Lower-level features; the formula is expressed as: ; In the formula, This is represented as a transposed convolution. Represented as bilinear interpolation; Sb7, then use channel attention mechanism to extract features This is converted into corresponding attention weights, which are applied to low-level features while ensuring dimensionality consistency. The low-level features are then filtered, and finally... Characteristics of high-level personnel The fusion process is performed to obtain the aforementioned fusion features.

2. The maritime vessel detection method based on feature extraction and feature weighted selection fusion according to claim 1, characterized in that: In step S2, the original multi-scale feature map is used. The multi-scale enhanced feature map is obtained. The specific steps are as follows: Sa1. Use bilinear interpolation in the feature extraction network to process the original multi-scale feature map. Augmentation is performed to obtain the augmented feature map. The formula is expressed as: ; In the formula, To introduce the actual sampling position after incorporating a learnable offset, , This is the center position of the current convolution window; The nth offset in the sampling grid of the regular convolution kernel; The learnable two-dimensional spatial offset of the nth sampling point; This is represented as an enumerated feature map on the original multi-scale feature map. The positions of all integral spaces in the space, It is a bilinear interpolation kernel. Represented as middle The value of the point, This is represented as an enumerated feature map on the original multi-scale feature map. The positions of all integration spaces in the interval correspond to the positions of all integration spaces in the interval. The values ​​above; Sa2, subsequently, a weighted mechanism was introduced to augment the feature map. By performing a weighted operation, the multi-scale enhanced feature map is obtained. The formula is expressed as: ; In the formula, The weight of each point, For the total number of points, For the first A learnable modulation scalar at each position, ranging from 0 to 1.

3. The maritime vessel detection method based on feature extraction and feature weighted selection fusion according to claim 1, characterized in that: In step S3, the total loss function value includes classification loss, bounding box regression loss, and confidence loss; the classification loss is the binary cross-entropy loss for multiple targets, and the formula for calculating the total loss function value is: ; In the formula, This is the total loss function value. For classifying losses, For bounding box regression loss, For confidence loss, and The weights for different losses are given; the formula for calculating the classification loss is: ; In the formula, For the Sigmoid function, It is the natural logarithm. This represents the probability that a sample is predicted as a positive example, where the sample is an object in the image to be detected; It is the sample integral value; The The calculation formula is: ; In the formula, The predicted value output by the network. These are the nearest neighbor predictions output by the network. For the actual value of the sample, For sample integral values, The integral value of the nearest sample; The The calculation formula is: ; In the formula, It is the true value of the target. It is the predicted value output by the model.

4. The ship detection method based on feature extraction and feature weighted selection fusion according to claim 3, characterized in that: In calculation After the loss, the Probiou method was also used to reduce the loss. The loss is the transformation from a horizontal bounding box (HBB) and a rotated bounding box (OBB) to a Gaussian bounding box (GBB) using the Probiou method. The transformation from HBB to GBB follows the assumption that the target region is a two-dimensional binary region with a uniform probability distribution. The formula for the mean-covariance matrix ∑ of this two-dimensional binary region distribution is: ; ; In the formula, For the region area; Let be the mean vector of the samples. For sample values, for transpose; The binary region of the horizontal bounding box (HBB) is based on Centered on, high as , width is The mean and covariance matrix of the rectangular region are calculated using the following formulas: ; Then the Parse distance The calculation is performed using the following formula: ; In the formula, This represents the positional difference between the sample bounding box and the bounding box detected by the algorithm. The shape difference between the sample bounding box and the bounding box detected by the algorithm; and The calculation formula is: ; ; In the formula, The x-coordinate of the center of the predicted bounding box, The vertical coordinate of the center of the predicted bounding box; The x-coordinate of the actual center of the annotation box. The vertical coordinate of the actual center of the annotation box; The variance of the center x-coordinate of the prediction box, The variance of the center ordinate of the prediction box, The covariance of the horizontal and vertical coordinates of the center of the predicted bounding box; Let x be the variance of the x-coordinate of the center of the true bounding box. The variance of the center ordinate of the true bounding box, The covariance of the x and y coordinates of the center of the true bounding box; However, the distance of the Barthel Since it is not a real distance, the triangle inequality cannot be satisfied; therefore, we use... Distance, the formula is: ; In the formula, For Hellinger distance, The predicted bounding box and the target bounding box have the same probability distribution if and only if they are identical. The probability between the two Gaussian boxes (GBB) The calculation formula is as follows: 。 5. The maritime vessel detection method based on feature extraction and feature weighted selection fusion according to claim 1, characterized in that: In step S2, the original multi-scale feature map Including the original multi-scale feature map Original multi-scale feature map Original multi-scale feature map and the original multi-scale feature map The multi-scale enhanced feature map Including multi-scale enhanced feature maps Multi-scale enhanced feature maps Multi-scale enhanced feature maps and multi-scale enhanced feature maps The feature extraction network used is the CSPDarkNet-S network.

6. A maritime vessel identification computer based on feature extraction and feature weighted selection fusion, comprising a memory, a processor, and program instructions stored in the memory for execution by the processor, characterized in that: The processor executes the program instructions to implement the steps of the method according to any one of claims 1 to 5.

7. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that implements the method of any one of claims 1 to 5.