Sea surface target detection and identification method and system, and readable storage medium

By using the GAN-FM model in the drone for multimodal data feature fusion and combining with the YOLOv10 detection and identification network, the problem of the joint use of multi-load data in the sea surface target detection of the drone is solved, and high-precision and real-time target detection are achieved.

CN120107542APending Publication Date: 2025-06-06CSSC SYST ENG RES INST +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411946951.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-27
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

The lack of multi-load detection and multi-modal data for unanimous use of drones during sea surface target detection, resulting in low recognition accuracy and large calculation amount, making it difficult to meet the real-time detection needs.

Method used

The improved feature fusion model based on the GAN-FM model is used to feature fusion of visible light, infrared and radar SAR images detected by drones, and a YOLOv10 detection and identification network is built to reduce inference delay and calculation through dual-label allocation strategy and lightweight model structure.

Benefits of technology

It realizes high-precision real-time detection and recognition of sea surface targets, solves the performance attenuation problem under conditions of limited computing power, and improves detection efficiency and real-time performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107542A_ABST
    Figure CN120107542A_ABST
Patent Text Reader

Abstract

The invention provides a sea surface target detection and identification method and system, and a readable storage medium, and the method comprises the steps: obtaining detection data of an unmanned plane, and carrying out the preprocessing of the detection data, the detection data comprising a visible light image, an infrared image, and a radar SAR image; constructing an improved feature fusion model based on a GAN-FM model, and performing feature fusion on the preprocessed detection data through the improved feature fusion model to obtain a fused image; and constructing a YOLOv10 detection and recognition network, and performing sea surface target detection and recognition on the fused image through the YOLOv10 detection and recognition network to obtain a recognition result. Through the technical scheme of the invention, radar, visible light, infrared and other load information can be fused for target detection, and the problem that the unmanned aerial vehicle is difficult to detect and recognize due to the fact that the unmanned aerial vehicle is seriously influenced by weather, detected target characteristics, distance and other factors is effectively solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the field of target detection technology, and in particular relates to a method and system for detecting and identifying sea surface targets, and a readable storage medium. Background Art

[0002] Since the typical use scenarios of ship-borne drones are to look down at the sea surface from a high altitude, the various payloads carried by the drones are used to detect and identify targets. Due to the long detection distance and the short multi-pass line of sight distance of high-altitude clouds, it is usually very difficult to use optoelectronic payloads alone for target detection. Similarly, due to the large number of "low, small and slow" targets on the sea surface and the strong interference of sea clutter, the effect of using radar payloads alone for target detection is usually very poor. When drones carry a single payload for sea surface target detection and identification, they have a limited range of use and it is difficult to cover the entire mission scenario. At the same time, the recognition accuracy of a single payload is poor. Taking optoelectronics as an example, it is difficult to identify target texture, color and other features when the detection distance is far, and it is difficult to distinguish between ships and large reefs on the sea surface. However, if the target speed characteristics detected by the radar are added, the target recognition can be completed more accurately.

[0003] Existing drones and drone control stations are generally lacking in high-performance computing power due to the requirements for autonomous and controllable equipment. The current mainstream detection and recognition algorithms, taking the YOLO family as an example, are difficult to achieve real-time detection based on the computing power of the control station due to the design of non-maximum suppression (NMS) modules, classifiers and regressors with the same structure, and the use of the same convolution kernel in the network structure. Therefore, a lightweight model is needed, but more pruning operations will lead to more degradation of model performance, which cannot meet the drone's needs for sea surface target detection and recognition. Summary of the invention

[0004] This application aims to solve or improve the above technical problems.

[0005] To this end, an embodiment of the present application provides a method for detecting and identifying sea surface targets.

[0006] The embodiment of the present application also provides a sea surface target detection and recognition system.

[0007] The embodiment of the present application also provides a sea surface target detection and recognition system.

[0008] The embodiment of the present application also provides a readable storage medium.

[0009] To achieve the above-mentioned purpose, an embodiment of the present application provides a method for detecting and identifying sea surface targets, including: obtaining detection data of a drone and preprocessing the detection data, the detection data including visible light images, infrared images and radar SAR images; constructing an improved feature fusion model based on a GAN-FM model, performing feature fusion on the preprocessed detection data through the improved feature fusion model to obtain a fused image, the improved feature fusion model including a generator, an infrared discriminator, a visible light discriminator and a radar SAR image discriminator; constructing a YOLOv10 detection and recognition network, performing sea surface target detection and recognition on the fused image through the YOLOv10 detection and recognition network to obtain a recognition result.

[0010] According to the sea surface target detection and recognition method provided by the present application, the detection data of the UAV is first obtained, and the detection data is preprocessed. The detection data includes visible light images, infrared images and radar SAR images. Then, an improved feature fusion model based on the GAN-FM model is constructed, and the preprocessed detection data is feature fused by the improved feature fusion model to obtain a fused image. The improved feature fusion model includes a generator, an infrared discriminator, a visible light discriminator and a radar SAR image discriminator. In order to realize the fusion of radar SAR images and optoelectronic data, the GAN-FM structure is modified, and the third radar SAR image discriminator is added. The same structure as the infrared discriminator and the visible light discriminator is adopted to realize the distinction between the fused image and the SAR image, and to maximize the retention of the intensity, phase and other features on the SAR image. For the fused features, a YOLOv10 detection and recognition network is constructed, and the non-maximum suppression strategy is replaced with a consistent dual-label allocation strategy, which solves the redundant prediction problem in the model post-processing in turn and reduces the inference delay. At the same time, the classification head structure is modified to lightweight the model. Finally, the model feature extraction module structure is modified, and a small flip module and a large convolution kernel are used to significantly reduce the training time and the time required for reasoning. A partial self-attention mechanism (PSA) is added to improve the model performance. The drone detection radar and optoelectronic data are used as input. After the above data preprocessing, feature fusion, and target detection and recognition steps, high-precision real-time detection and recognition of sea surface targets are achieved, solving the problem of lack of autonomous joint use of multi-payload detection and multi-modal data when drones detect sea surface targets. The improved feature fusion model is used to fuse features and results respectively to improve the efficiency of sea surface target detection and recognition. YOLOv10 is used to solve the problem of serious performance degradation of existing detection and recognition models under limited computing power conditions, reducing the amount of target detection and recognition calculations and improving the real-time performance of the algorithm.

[0011] In addition, the technical solution provided by this application may also have the following additional technical features:

[0012] In some technical solutions, optionally, the detection data of the UAV is obtained and the detection data is preprocessed, including: obtaining the detection data of the UAV, denoising the detection data by wavelet threshold denoising method and double-tree complex wavelet denoising method to obtain denoised data; enhancing the denoised data by multiple data enhancement methods to obtain enhanced data; and performing equalization processing on the enhanced data to obtain equalized sample data.

[0013] In this technical solution, the detection data of the drone is obtained and the detection data is preprocessed. Specifically, the detection data of the drone is first obtained, and the detection data is denoised by wavelet threshold denoising method and double-tree complex wavelet denoising method to obtain denoised data. Then, the denoised data is enhanced by multiple data enhancement methods to obtain enhanced data. Finally, the enhanced data is equalized to obtain equalized sample data. Specifically, the wavelet threshold denoising method and double-tree complex wavelet denoising method are used to denoise the radar and photoelectric data of the photoelectric and radar detection data, which can reduce the interference of sea clutter in the detection and recognition of sea surface targets and improve the detection accuracy of radar, photoelectric and other drones carrying payloads. Since there is less data on sea surface targets taken by drones, the depth detection and feature fusion models are greatly affected by the number of samples. In order to increase the number of effective samples, the denoised data is enhanced by multiple data enhancement methods to obtain enhanced data and achieve sample enrichment. In addition, the number of samples of foreground (various types of ships on the sea, unmanned boats and other targets) and background (non-target information such as sea reefs and the sea) is extremely unbalanced. To solve this problem, the enhanced data is balanced to obtain balanced sample data.

[0014] In some technical solutions, the enhanced data is optionally equalized to obtain equalized sample data, including: performing sliding window cropping on multiple samples in the enhanced data to obtain a sample set with a resolution lower than the original image; obtaining the number of target samples containing the target in the sample set, and randomly extracting samples that do not contain the target in the sample set according to the number of target samples to obtain equalized sample data.

[0015] In this technical solution, the enhanced data is equalized to obtain the equalized sample data, specifically, firstly, a plurality of samples in the enhanced data are subjected to sliding window cropping to obtain a sample set with a resolution lower than that of the original image. Then, the number of target samples containing the target in the sample set is obtained, and the samples not containing the target in the sample set are randomly sampled according to the number of target samples to obtain the equalized sample data. Specifically, by performing sliding window cropping on a plurality of samples, a large number of samples with a lower resolution than the original image are obtained, the number of samples containing the target is counted, and the samples not containing the target are randomly sampled to ensure that the ratio of the number of samples containing the target to the number of samples not containing the target is approximately one to one, thereby achieving the equalization of the samples.

[0016] In some technical solutions, optionally, the data enhancement method includes one of the following: flip transformation, rotation transformation, linear brightness shift, non-linear brightness shift, contrast map adjustment, image perspective, cropping and splicing.

[0017] In this technical solution, data enhancement methods include flip transformation, rotation transformation, linear brightness shift, nonlinear brightness shift, contrast map adjustment, image perspective or cropping and splicing.

[0018] In some technical solutions, optionally, feature fusion is performed on the preprocessed detection data by improving the feature fusion model to obtain a fused image, including: performing channel stitching operations on visible light images, infrared images and radar SAR images in matching states, extracting fusion features through a generator to generate a fused image; inputting the fused image into an infrared discriminator, a visible light discriminator and a radar SAR image discriminator respectively to distinguish the fused image from the original image; and achieving Nash equilibrium by averaging the outputs of the infrared discriminator, the visible light discriminator and the radar SAR image discriminator and batches.

[0019] In this technical solution, the feature fusion of the preprocessed detection data is performed by improving the feature fusion model to obtain a fused image. Specifically, the visible light image, infrared image and radar SAR image in the matching state are firstly subjected to channel splicing operation, and the fusion feature is extracted by the generator to generate a fused image. Then the fused image is respectively input into the infrared discriminator, the visible light discriminator and the radar SAR image discriminator to distinguish the fused image from the original image. Finally, the outputs of the infrared discriminator, the visible light discriminator, the radar SAR image discriminator and the batch are averaged to achieve Nash equilibrium. Specifically, after completing the data enhancement, the feature fusion operation is performed. The radar SAR image, infrared image and visible light image in the matching state are subjected to channel splicing operation, and the fusion feature is extracted by the U-shaped generator to generate a fused image. The fused image is respectively input into the radar, visible light and infrared discriminators to distinguish whether the fused image is the original image. By averaging the outputs of the discriminator and the batch, the Nash equilibrium is finally achieved, that is, the generator expects the discriminator to think that the generated fused image is both a visible light image and an infrared image or a radar image, thereby achieving high-precision fusion of the image.

[0020] In some technical schemes, optionally, the detection data is denoised by a wavelet threshold denoising method and a dual-tree complex wavelet denoising method to obtain denoised data, including: performing complex wavelet decomposition on the detection data; calculating the wavelet coefficient energies of multiple scales, obtaining the correlation coefficient energies of the corresponding scales, normalizing the correlation coefficients of the wavelet coefficients, and obtaining a coefficient threshold; determining whether the wavelet coefficient is less than the coefficient threshold; if so, retaining the wavelet coefficient and reconstructing the retained wavelet coefficient.

[0021] In this technical solution, the detection data is denoised by wavelet threshold denoising method and double-tree complex wavelet denoising method to obtain denoised data. Specifically, the detection data is first decomposed by complex wavelet. Then the wavelet coefficient energy of multiple scales is calculated, the correlation coefficient energy of the corresponding scale is obtained, the correlation coefficient of the wavelet coefficient is normalized, and the coefficient threshold is obtained. Then it is judged whether the wavelet coefficient is less than the coefficient threshold. If so, the wavelet coefficient is retained and the retained wavelet coefficient is reconstructed. Specifically, since the sea surface target information detected by the actual UAV payload is usually subject to transmission interference, sea clutter interference, data compression and other problems, it contains extensive noise, especially radar data. Since the scattering characteristics of sea clutter are very close to those of targets such as ships in Doppler frequency, the radar detection performance is greatly interfered by sea clutter. In order to solve this problem, the double-tree complex wavelet change is applied to suppress sea clutter in view of the characteristics of sea clutter with wide spectrum width, unstable distribution and weak correlation. The dual-tree complex wavelet transform decomposes and reconstructs complex signals through two different sets of low-pass filters and high-pass filters. The two sets of filters form the real tree and imaginary tree of the transform, and constitute a Hilbert transform pair. For radar signals, complex wavelet decomposition is performed. After the decomposition is completed, the energy of the wavelet coefficients at each scale is calculated, and the energy of the correlation coefficient of the corresponding scale is obtained. The correlation coefficient of the wavelet coefficient is normalized, and the value is used as the threshold to consider that the wavelet coefficients less than the threshold are retained, and the retained wavelet coefficients are reconstructed to achieve sea clutter suppression.

[0022] In some technical solutions, optionally, the detection data includes photoelectric detection data and radar detection data.

[0023] In this technical solution, the detection data includes photoelectric detection data and radar detection data.

[0024] The embodiment of the present application provides a sea surface target detection and recognition system, including: an acquisition module, used to acquire detection data of a drone and preprocess the detection data, the detection data including visible light images, infrared images and radar SAR images; a feature fusion module, used to construct an improved feature fusion model based on a GAN-FM model, and perform feature fusion on the preprocessed detection data through the improved feature fusion model to obtain a fused image, the improved feature fusion model including a generator, an infrared discriminator, a visible light discriminator and a radar SAR image discriminator; a detection and recognition module, used to construct a YOLOv10 detection and recognition network, and perform sea surface target detection and recognition on the fused image through the YOLOv10 detection and recognition network to obtain a recognition result.

[0025] The sea surface target detection and recognition system provided by the present application includes an acquisition module, a feature fusion module and a detection and recognition module. Among them, the acquisition module is used to acquire the detection data of the drone and preprocess the detection data, and the detection data includes visible light images, infrared images and radar SAR images. The feature fusion module is used to construct an improved feature fusion model based on the GAN-FM model, and the preprocessed detection data is feature-fused by the improved feature fusion model to obtain a fused image. The improved feature fusion model includes a generator, an infrared discriminator, a visible light discriminator and a radar SAR image discriminator. The detection and recognition module is used to construct a YOLOv10 detection and recognition network, and the fused image is subjected to sea surface target detection and recognition by the YOLOv10 detection and recognition network to obtain a recognition result. Taking the drone detection radar and optoelectronic data as input, after the above-mentioned data preprocessing, feature fusion, target detection and recognition steps, high-precision real-time detection and recognition of sea surface targets is achieved, which solves the problem of lack of autonomous joint use of multi-payload detection and multi-modal data when drones perform sea surface target detection. The improved feature fusion model is used to fuse the features and results respectively to improve the efficiency of sea surface target detection and recognition. YOLOv10 is used to solve the problem of serious performance degradation of existing detection and recognition models under limited computing power conditions, which reduces the calculation amount of target detection and recognition and improves the real-time performance of the algorithm.

[0026] An embodiment of the present application provides a sea surface target detection and identification system, including: a memory and a processor, wherein the memory stores a program or instruction that can be run on the processor, and when the processor executes the program or instruction, it implements the sea surface target detection and identification method of any one of the technical solutions of the first aspect, and thus has the technical effect of any of the technical solutions of the first aspect above, which will not be repeated here.

[0027] An embodiment of the present application provides a readable storage medium on which a program or instruction is stored. When the program or instruction is executed by a processor, the steps of the sea surface target detection and identification method of any one of the technical solutions of the first aspect are implemented, and thus the technical effects of any of the technical solutions of the first aspect are obtained, which will not be repeated here.

[0028] Additional aspects and advantages of the present application will become apparent in the following description or will be learned through practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] The above and / or additional aspects and advantages of the present application will become apparent and easily understood from the description of the embodiments in conjunction with the following drawings, in which:

[0030] Figure 1 A schematic diagram of the steps of a method for detecting and identifying sea surface targets according to an embodiment of the present application;

[0031] Figure 2A schematic diagram of the steps of a method for detecting and identifying sea surface targets according to an embodiment of the present application;

[0032] Figure 3 A schematic diagram of the steps of a method for detecting and identifying sea surface targets according to an embodiment of the present application;

[0033] Figure 4 A schematic diagram of the steps of a method for detecting and identifying sea surface targets according to an embodiment of the present application;

[0034] Figure 5 A schematic diagram of the steps of a method for detecting and identifying sea surface targets according to an embodiment of the present application;

[0035] Figure 6 This is a schematic block diagram of the structure of a sea surface target detection and recognition system according to an embodiment of the present application;

[0036] Figure 7 This is a schematic block diagram of the structure of a sea surface target detection and recognition system according to an embodiment of the present application;

[0037] Figure 8 A schematic diagram of dual-tree complex wavelet decomposition of a method for detecting and identifying sea targets according to an embodiment of the present application;

[0038] Fig. 9 A schematic diagram of dual-tree complex wavelet decomposition of a method for detecting and identifying sea targets according to an embodiment of the present application;

[0039] FIG10( a ) is an original image of a method for detecting and identifying sea surface targets according to an embodiment of the present application;

[0040] FIG10( b ) is a schematic diagram of a stitching operation of a method for detecting and identifying sea surface targets according to an embodiment of the present application;

[0041] FIG10( c ) is a schematic diagram of the rotation operation of a method for detecting and identifying sea surface targets according to an embodiment of the present application;

[0042] FIG. 11( a ) is a generator structure of a method for detecting and identifying sea targets according to an embodiment of the present application;

[0043] FIG11( b ) is a full-scale skip connection structure of a DEB decoding module of a method for detecting and identifying sea targets according to an embodiment of the present application;

[0044] Fig.12 A Markov discriminator structure diagram of a method for detecting and identifying sea targets according to an embodiment of the present application;

[0045] Fig.13 A schematic diagram of dual label allocation of a method for detecting and identifying sea targets according to an embodiment of the present application;

[0046] Fig.14A schematic diagram of the CIB structure of a method for detecting and identifying sea targets according to an embodiment of the present application;

[0047] Fig.15 A schematic diagram of the PSA structure of a method for detecting and identifying sea surface targets according to an embodiment of the present application.

[0048] in, Figure 6 and Figure 7 The corresponding relationship between the reference numerals and component names in the figure is:

[0049] 10: sea surface target detection and recognition system; 110: acquisition module; 120: feature fusion module; 130: detection and recognition module; 20: sea surface target detection and recognition system; 300: memory; 400: processor. DETAILED DESCRIPTION

[0050] In order to make the purpose, technical solutions and advantages of this application clearer, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the drawings in the embodiments of this application. Obviously, the described embodiments are part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0051] Refer to the following Figures 1 to 15 Describe the sea surface target detection and recognition method and system, and readable storage medium of some embodiments of the present application.

[0052] like Figure 1 As shown, the embodiment of the first aspect of the present application provides a method for detecting and identifying sea surface targets, comprising the following steps:

[0053] Step S102: Acquire detection data of the UAV and pre-process the detection data, the detection data including visible light images, infrared images and radar SAR images;

[0054] Step S104: constructing an improved feature fusion model based on the GAN-FM model, performing feature fusion on the preprocessed detection data through the improved feature fusion model to obtain a fused image, wherein the improved feature fusion model includes a generator, an infrared discriminator, a visible light discriminator, and a radar SAR image discriminator;

[0055] Step S106: construct a YOLOv10 detection and recognition network, and perform sea surface target detection and recognition on the fused image through the YOLOv10 detection and recognition network to obtain a recognition result.

[0056] According to the sea surface target detection and recognition method provided by this embodiment, the detection data of the drone is first obtained, and the detection data is preprocessed. The detection data includes visible light images, infrared images and radar SAR images. Then, an improved feature fusion model based on the GAN-FM model is constructed, and the preprocessed detection data is feature fused by the improved feature fusion model to obtain a fused image. The improved feature fusion model includes a generator, an infrared discriminator, a visible light discriminator and a radar SAR image discriminator. In order to realize the fusion of radar SAR images and optoelectronic data, the GAN-FM structure is modified, and the third radar SAR image discriminator is added. The same structure is adopted as the infrared discriminator and the visible light discriminator to realize the distinction between the fused image and the SAR image, and the intensity, phase and other features on the SAR image are retained to the maximum extent. For the fused features, a YOLOv10 detection and recognition network is constructed, and the non-maximum suppression strategy is replaced with a consistent dual-label allocation strategy, which solves the redundant prediction problem in the model post-processing in turn and reduces the inference delay. At the same time, the classification head structure is modified to lightweight the model. Finally, the model feature extraction module structure is modified, and a small flip module and a large convolution kernel are used to significantly reduce the training time and the time required for reasoning. A partial self-attention mechanism (PSA) is added to improve the model performance. The drone detection radar and optoelectronic data are used as input. After the above data preprocessing, feature fusion, and target detection and recognition steps, high-precision real-time detection and recognition of sea surface targets are achieved, solving the problem of lack of autonomous joint use of multi-payload detection and multi-modal data when drones detect sea surface targets. The improved feature fusion model is used to fuse features and results respectively to improve the efficiency of sea surface target detection and recognition. YOLOv10 is used to solve the problem of serious performance degradation of existing detection and recognition models under limited computing power conditions, reducing the amount of target detection and recognition calculations and improving the real-time performance of the algorithm.

[0057] Among them, the GAN-FM model is a generative adversarial network (GAN) model for the fusion of infrared and visible light images. It combines full-scale skip connections and dual Markovian discriminators to effectively fuse the thermal information of infrared images and the texture details of visible light images. YOLOv10 is a real-time target detection method developed based on the Ultralytics Python package, which aims to address the deficiencies of previous YOLO versions in post-processing and model architecture. By eliminating non-maximum suppression (NMS) and optimizing various model components, YOLOv10 achieves state-of-the-art performance while significantly reducing computational overhead.

[0058] like Figure 2 As shown, according to a method for detecting and identifying sea targets in an embodiment of the present application, the detection data of the UAV is obtained and the detection data is preprocessed, including the following steps:

[0059] Step S202: Acquire the detection data of the UAV, and perform noise reduction on the detection data by using the wavelet threshold noise reduction method and the double-tree complex wavelet noise reduction method to obtain noise-reduced data;

[0060] Step S204: performing data enhancement on the denoised data through multiple data enhancement methods to obtain enhanced data;

[0061] Step S206: performing equalization processing on the enhanced data to obtain equalized sample data.

[0062] In this embodiment, the detection data of the drone is obtained, and the detection data is preprocessed. Specifically, the detection data of the drone is first obtained, and the detection data is denoised by wavelet threshold denoising method and double-tree complex wavelet denoising method to obtain denoised data. Then, the denoised data is enhanced by multiple data enhancement methods to obtain enhanced data. Finally, the enhanced data is equalized to obtain equalized sample data. Specifically, the wavelet threshold denoising method and double-tree complex wavelet denoising method are used to denoise the radar and photoelectric data of the photoelectric and radar detection data, which can reduce the interference of sea clutter in the detection and recognition of sea surface targets, and improve the detection accuracy of radar, photoelectric and other drones carrying payloads. Since the data of sea surface targets taken by drones is relatively small, the depth detection and feature fusion models are greatly affected by the number of samples. In order to increase the number of effective samples, the denoised data is enhanced by multiple data enhancement methods to obtain enhanced data and achieve sample enrichment. In addition, the number of samples of foreground (various types of ships on the sea, unmanned boats and other targets) and background (non-target information such as sea reefs and the sea) is extremely unbalanced. To solve this problem, the enhanced data is balanced to obtain balanced sample data.

[0063] like Figure 3 As shown, according to a method for detecting and identifying sea targets in an embodiment of the present application, the enhanced data is equalized to obtain equalized sample data, including the following steps:

[0064] Step S302: performing sliding window cropping on multiple samples in the enhanced data to obtain a sample set with a resolution lower than the original image;

[0065] Step S304: Obtain the number of target samples containing the target in the sample set, and randomly extract samples that do not contain the target in the sample set according to the number of target samples to obtain balanced sample data.

[0066] In this embodiment, the enhanced data is subjected to equalization processing to obtain equalized sample data, specifically, firstly, a plurality of samples in the enhanced data are subjected to sliding window cropping to obtain a sample set with a resolution lower than the original image. Then, the number of target samples containing the target in the sample set is obtained, and samples not containing the target in the sample set are randomly sampled according to the number of target samples to obtain equalized sample data. Specifically, by performing sliding window cropping on a plurality of samples, a large number of samples with a lower resolution than the original image are obtained, the number of samples containing the target is counted, and samples not containing the target are randomly sampled to ensure that the ratio of the number of samples containing the target to the number of samples not containing the target is approximately one to one, thereby achieving equalization processing of the samples.

[0067] In some embodiments, optionally, the data enhancement method includes one of the following: flip transformation, rotation transformation, linear brightness shift, non-linear brightness shift, contrast map adjustment, image perspective, cropping and splicing.

[0068] like Figure 4 As shown, according to a method for detecting and identifying sea targets in an embodiment of the present application, the feature fusion model is improved to fuse the preprocessed detection data to obtain a fused image, including the following steps:

[0069] Step S402: performing channel stitching operation on the visible light image, infrared image and radar SAR image in a matching state, extracting fusion features through a generator, and generating a fused image;

[0070] Step S404: inputting the fused image into the infrared discriminator, the visible light discriminator and the radar SAR image discriminator respectively to distinguish the fused image from the original image;

[0071] Step S406: A Nash equilibrium is achieved by averaging the outputs of the infrared discriminator, the visible light discriminator, the radar SAR image discriminator and the batch.

[0072] In this embodiment, the preprocessed detection data is feature fused by improving the feature fusion model to obtain a fused image. Specifically, the visible light image, infrared image and radar SAR image in the matching state are firstly subjected to channel splicing operation, and the fusion feature is extracted by the generator to generate a fused image. Then the fused image is respectively input into the infrared discriminator, the visible light discriminator and the radar SAR image discriminator to distinguish the fused image from the original image. Finally, the outputs of the infrared discriminator, the visible light discriminator and the radar SAR image discriminator and the batch are averaged to achieve Nash equilibrium. Specifically, after completing the data enhancement, the feature fusion operation is performed. The radar SAR image, infrared image and visible light image in the matching state are subjected to channel splicing operation, and the fusion feature is extracted by the U-shaped generator to generate a fused image. The fused image is respectively input into the radar, visible light and infrared discriminators to distinguish whether the fused image is the original image. By averaging the outputs of the discriminator and the batch, the Nash equilibrium is finally achieved, that is, the generator expects the discriminator to think that the generated fused image is both a visible light image and an infrared image or a radar image, thereby achieving high-precision fusion of the image.

[0073] like Figure 5 As shown, according to a method for detecting and identifying sea targets in an embodiment of the present application, the detection data is denoised by a wavelet threshold denoising method and a dual-tree complex wavelet denoising method to obtain denoised data, including the following steps:

[0074] Step S502: performing complex wavelet decomposition on the detection data;

[0075] Step S504: Calculate the wavelet coefficient energies of multiple scales, obtain the correlation coefficient energies of the corresponding scales, normalize the correlation coefficients of the wavelet coefficients, and obtain the coefficient thresholds;

[0076] Step S506: determine whether the wavelet coefficient is less than the coefficient threshold;

[0077] Step S508: If yes, the wavelet coefficients are retained and the retained wavelet coefficients are reconstructed.

[0078] In this embodiment, the detection data is denoised by wavelet threshold denoising method and double-tree complex wavelet denoising method to obtain denoised data, specifically, the detection data is firstly subjected to complex wavelet decomposition. Then the wavelet coefficient energy of multiple scales is calculated, the correlation coefficient energy of the corresponding scale is obtained, the correlation coefficient of the wavelet coefficient is normalized, and the coefficient threshold is obtained. Then it is determined whether the wavelet coefficient is less than the coefficient threshold. If so, the wavelet coefficient is retained and the retained wavelet coefficient is reconstructed. Specifically, since the sea surface target information detected by the actual UAV payload is usually subject to transmission interference, sea clutter interference, data compression and other problems, it contains extensive noise, especially radar data, because the sea clutter scattering characteristics are very close to targets such as ships in Doppler frequency, resulting in radar detection performance being greatly interfered by sea clutter. To solve this problem, the double-tree complex wavelet change is applied to suppress sea clutter in view of the characteristics of sea clutter having a wide spectrum width, unstable distribution and weak correlation. The dual-tree complex wavelet transform decomposes and reconstructs complex signals through two different sets of low-pass filters and high-pass filters. The two sets of filters form the real tree and imaginary tree of the transform, and constitute a Hilbert transform pair. For radar signals, complex wavelet decomposition is performed. After the decomposition is completed, the energy of the wavelet coefficients at each scale is calculated, and the energy of the correlation coefficient of the corresponding scale is obtained. The correlation coefficient of the wavelet coefficient is normalized, and the value is used as the threshold to consider that the wavelet coefficients less than the threshold are retained, and the retained wavelet coefficients are reconstructed to achieve sea clutter suppression.

[0079] In some embodiments, optionally, the detection data includes photoelectric detection data and radar detection data.

[0080] like Figure 6 As shown, an embodiment of the second aspect of the present application provides a sea surface target detection and recognition system 10, including: an acquisition module 110, used to acquire the detection data of the UAV and preprocess the detection data, the detection data including visible light images, infrared images and radar SAR images; a feature fusion module 120, used to construct an improved feature fusion model based on the GAN-FM model, and perform feature fusion on the preprocessed detection data through the improved feature fusion model to obtain a fused image, the improved feature fusion model includes a generator, an infrared discriminator, a visible light discriminator and a radar SAR image discriminator; a detection and recognition module 130, used to construct a YOLOv10 detection and recognition network, and perform sea surface target detection and recognition on the fused image through the YOLOv10 detection and recognition network to obtain a recognition result.

[0081] The sea surface target detection and recognition system 10 provided in this embodiment includes an acquisition module 110, a feature fusion module 120 and a detection and recognition module 130. Among them, the acquisition module 110 is used to acquire the detection data of the drone and preprocess the detection data, and the detection data includes visible light images, infrared images and radar SAR images. The feature fusion module 120 is used to construct an improved feature fusion model based on the GAN-FM model, and the preprocessed detection data is feature-fused by the improved feature fusion model to obtain a fused image. The improved feature fusion model includes a generator, an infrared discriminator, a visible light discriminator and a radar SAR image discriminator. The detection and recognition module 130 is used to construct a YOLOv10 detection and recognition network, and the fused image is subjected to sea surface target detection and recognition by the YOLOv10 detection and recognition network to obtain a recognition result. Taking the drone detection radar and optoelectronic data as input, after the above-mentioned data preprocessing, feature fusion, target detection and recognition steps, high-precision real-time detection and recognition of sea surface targets is achieved, which solves the problem of lack of autonomous joint use of multi-payload detection and multi-modal data when drones perform sea surface target detection. The improved feature fusion model is used to fuse the features and results respectively to improve the efficiency of sea surface target detection and recognition. YOLOv10 is used to solve the problem of serious performance degradation of existing detection and recognition models under limited computing power conditions, which reduces the calculation amount of target detection and recognition and improves the real-time performance of the algorithm.

[0082] like Figure 7 As shown, an embodiment of the third aspect of the present application provides a sea surface target detection and identification system 20, including: a memory 300 and a processor 400, wherein the memory 300 stores a program or instruction that can be run on the processor 400, and when the processor 400 executes the program or instruction, it implements the steps of the sea surface target detection and identification method of any one of the embodiments of the first aspect, and thus has the technical effect of any one of the embodiments of the first aspect mentioned above, which will not be repeated here.

[0083] The embodiment of the fourth aspect of the present application provides a readable storage medium on which a program or instruction is stored. When the program or instruction is executed by a processor, the steps of the sea surface target detection and identification method of any one of the embodiments of the first aspect are implemented, and thus it has the technical effect of any one of the embodiments of the first aspect mentioned above, which will not be repeated here.

[0084] like Figures 8 to 15As shown, according to a specific embodiment of the sea surface target detection and recognition method provided by the present application, for the photoelectric and radar images collected by the drone at high altitude, the method can perform high-quality fusion of the multi-modal load information detected by the drone, and simultaneously apply the improved GAN-FM feature layer fusion method to solve the problem of poor target detection and recognition effect caused by the limited detection performance of a single sensor data. For the data after the feature layer fusion, the latest detection and recognition model YOLOv10 of the YOLO series is applied to the sea surface target detection and recognition. Through the consistent double assignment strategy, lightweight classification head, optimized downsampling method and other means, under the premise of ensuring the basic performance of the model, the reasoning delay is greatly reduced, so that the drone control station with insufficient computing power can bear the model operation requirements, so that the method has the possibility of installation. At the same time, due to the fusion of radar, visible light, infrared and other load information for target detection, it can effectively adapt to the problem of detection and recognition difficulties caused by the serious influence of weather, detection target characteristics, distance and other factors on drones.

[0085] The sea surface target detection and recognition method provided in this embodiment first performs feature fusion on multi-modal payload data, and then applies YOLOv10 to perform target detection to achieve high-precision sea surface target detection and recognition. The specific method flow is as follows:

[0086] 1) Data preprocessing:

[0087] In the maritime battlefield environment, the detection and identification of sea surface targets is usually interfered by sea clutter. In order to improve the detection accuracy of radar, optoelectronic and other UAV payloads, it is necessary to carry out sea clutter noise reduction technology research to achieve the purpose of data noise reduction and purification. To solve the above problems, the wavelet threshold noise reduction method and the double-tree complex wavelet noise reduction method are used to reduce the noise of radar and optoelectronic data.

[0088] Since there is less data on sea surface targets taken by drones, the depth detection and feature fusion models are greatly affected by the number of samples. In order to increase the number of effective samples, various data enhancement methods such as flip transformation, rotation transformation, linear brightness offset, nonlinear brightness offset, contrast image adjustment, image perspective, cropping and splicing are applied to enrich the samples. In addition, due to the characteristics of the naval battlefield, the number of samples of foreground (various types of ships, unmanned boats and other targets on the sea surface) and background (non-target information such as sea surface reefs and the sea) is extremely unbalanced. To solve this problem, a large number of samples with lower resolution than the original image are obtained by sliding window cropping of multiple samples, and the number of samples containing targets is counted. Samples without targets are randomly extracted to ensure that the ratio of samples containing targets to samples without targets is roughly one to one, so as to achieve balanced processing of samples.

[0089] 2) Feature-level fusion based on improved GAN-FM:

[0090] The information obtained by a single sensor is very limited and is affected by its own quality and performance. Therefore, drones are usually equipped with a large number of different types of sensors to meet the needs of detection and data collection. If the information collected by each sensor is processed separately and in isolation, it will not only increase the workload of information processing, but also cut off the inherent connection between the information of each sensor, lose the relevant environmental characteristics that may be contained in the organic combination of information, cause a waste of information resources, and may even lead to wrong decisions. In order to solve the above problems, the improved GAN-FM method is applied to realize feature-level fusion.

[0091] For the visible light images, infrared images, and radar SAR images of the drone after noise reduction, the first thing to do is to perform a registration operation, that is, it is necessary to ensure that the viewing angles of the three payloads are consistent and the shooting objects are consistent, so that feature fusion makes sense. Considering that in actual use, visible light and infrared are usually installed on the same optoelectronic payload, their images are registered. When the radar performs SAR imaging, it usually does not detect the same target as the optoelectronics, so it is necessary to select the registered data through the payload internal parameters during training. However, in actual use, it is usually after the radar finds a suspicious target that it guides the optoelectronics to perform collaborative detection of the suspicious target. In this process, the payload completes the automatic registration.

[0092] The traditional GAN-FM feature fusion model consists of 1 generator and 2 discriminators. The generator adopts the idea of ​​full-scale connection and draws on the network structure of U-net to extract multi-scale and hierarchical features of the input multi-channel data. The two discriminators have the same structure and consist of five convolutional layers. The first four layers use the ReLU activation function, while the last layer uses the tanh activation function to distinguish the fused image from the source image (the two discriminators distinguish infrared and visible light images respectively).

[0093] In order to realize the fusion of radar SAR images and optoelectronic data, the GAN-FM structure is modified and a third radar SAR image discriminator is added. The same structure as the infrared and visible light discriminators is adopted to distinguish the fused image from the SAR image and retain the intensity, phase and other features of the SAR image to the maximum extent.

[0094] 3) Sea surface target detection based on YOLOv10:

[0095] For the fused features, a YOLOv10 detection and recognition network is constructed, and the non-maximum suppression (NMS) strategy is replaced with a consistent dual-label assignment strategy, which solves the redundant prediction problem in the model post-processing and reduces the inference delay. At the same time, the classification head structure is modified to lightweight the model. Finally, the model feature extraction module structure is modified, and a small flip module (CIB module) and a large convolution kernel are used to significantly reduce the training time and the time required for inference. The partial self-attention mechanism (PSA) is added to improve the model performance.

[0096] The drone detection radar and optoelectronic data are used as input, and after the above-mentioned data noise reduction, feature fusion, and target detection and recognition steps, high-precision real-time detection and recognition of sea surface targets can be achieved.

[0097] The specific implementation is as follows:

[0098] Since the sea surface target information detected by the actual UAV payload is usually subject to transmission interference, sea clutter interference, data compression and other problems, it contains extensive noise, especially radar data. Since the scattering characteristics of sea clutter are very close to those of targets such as ships in Doppler frequency, the radar detection performance is greatly affected by sea clutter. To solve this problem, the dual-tree complex wavelet transform is applied to suppress sea clutter, considering that sea clutter has wide spectrum width, unstable distribution and weak correlation.

[0099] The dual-tree complex wavelet transform decomposes and reconstructs complex signals through two different sets of low-pass filters and high-pass filters. The two sets of filters form the real tree and imaginary tree of the transform and constitute a Hilbert transform pair. Let h 0 (n) and h 1 (n) are the low-pass filter and high-pass filter of the real part tree 1, g 0 (n) and g 1 (n) are the low pass filter and high pass filter of imaginary tree 2, and the real part h 0 (n) and h 1 (n) corresponds to the scaling function φ h (t) and the wavelet function ψ h (t) is defined as:

[0100]

[0101] and the imaginary part g 0 (n) and g 1 (n) The corresponding scaling function and wavelet function are:

[0102]

[0103] The wavelet functions of tree 1 and tree 2 constitute a Hilbert transform pair:

[0104] ψ g (t) = H{ψ h (t)};

[0105] For radar signals, complex wavelet decomposition is performed, such as Figure 8As shown in the figure, after the decomposition is completed, the energy of the wavelet coefficients at each scale is calculated, and the energy of the correlation coefficients of the corresponding scales is obtained. The correlation coefficients of the wavelet coefficients are normalized, and the value is used as the threshold to consider that the wavelet coefficients less than the threshold are retained. The retained wavelet coefficients are reconstructed to achieve sea clutter suppression. The specific process is as follows: Fig. 9 shown.

[0106] After data denoising, data enhancement is performed on the data, such as cropping, rotation, size change, contrast change, saturation change, etc., to increase the amount of data. Some operations are shown in Figure 10(a), Figure 10(b), and Figure 10(c). Note that for the infrared, visible light, and radar images of the same target, the same parameter changes should be performed to ensure that the bound multimodal data remains matched after data enhancement.

[0107] After completing the data enhancement, the feature fusion operation is performed. The channel splicing operation is performed on the matching radar SAR, infrared, and visible light images. Through the U-shaped generator, as shown in Figure 11(a) and Figure 11(b), the fusion feature extraction is performed and the fused image is generated. The fused image is input into the radar, visible light, and infrared discriminators respectively to distinguish whether the fused image is the original image. The three Markov discriminators have the same structure as shown in Fig.12 As shown in the figure, the loss function includes 1 generator, 1 infrared discriminator, 1 visible light discriminator, and 1 radar SAR discriminator, a total of 4 parts. Generator loss function L G , infrared discriminator loss function L Dir , visible light discriminator loss function L Dvi , radar discriminator loss function L DSAR .L G =L adv +λL con , where L adv To combat the loss L con is the content loss function, which is used to impose additional restrictions on the generator, and λ is the weight coefficient, which is used to control the importance of the two-part loss. The adversarial loss is used to guide the generator to produce a real fusion result to deceive the two discriminators, which can be defined as:

[0108] L adv =E(log(1-D vi (I f )))+E(log(1-D ir (I f )))+E(log(1-D SAR (I f ));

[0109] Content loss L con It consists of two parts:

[0110] L con =L grad +βL in ;

[0111] Where L in Indicates strength loss, L grad represents the gradient loss, and β is the weight coefficient, which is used to control the importance of the two parts of the loss. The formula is as follows:

[0112]

[0113] in, represents the Frobenius norm, ξ is a positive parameter that controls the weights of the two terms.

[0114] The gradient loss measures the degree of texture preservation, and the formula is as follows:

[0115]

[0116] where |·| represents the absolute value, |·| 1 represents the 1 norm, This is the Laplacian gradient operation.

[0117] The discriminator loss function is as follows:

[0118]

[0119] By averaging the outputs of the discriminator and the batch, a Nash equilibrium is finally reached, that is, the generator expects the discriminator to believe that the generated fused image is both a visible light image and an infrared image or a radar image, thereby achieving high-precision image fusion.

[0120] After fusion is completed, YOLOv10 is used to detect and identify the target on the fused image. YOLOv10 uses a dual-label allocation strategy to replace the non-maximum suppression strategy to solve the redundant prediction problem in model post-processing and reduce inference latency. The schematic diagram of the dual-label allocation strategy is shown in Fig.13 As shown in the figure. Label assignment is obtained by introducing another one-to-one branch with the same structure and optimization objective as the original one-to-many branch. The two branches are optimized together with the model, so that the framework network and the classification head are supervised by the one-to-many task at the same time. During the inference process, only the one-to-one branch is used for prediction, and the optimal match is used when the one-to-one branch makes predictions, which greatly shortens the training and inference time.

[0121] At the same time, in order to solve the problem of model engineering application, a small inversion module (CIB) is used to replace the basic convolution block structure in the network layer with a higher inherent order, which greatly reduces the model inference time. The CIB structure is as follows: Fig.14As shown. While ensuring the lightweight of the model, the partial self-attention mechanism (PSA) is introduced to improve the model performance. The schematic diagram is shown in Fig.15 shown.

[0122] Through the above process, it is possible to use drones to detect heterogeneous data and achieve high-precision detection and identification capabilities for sea surface targets.

[0123] The algorithm platform requirements are as follows:

[0124] The programming software used is PyCharm 2021.3.1 (Community Edition), Anaconda 4.8.3 and Python 3.9. The selected cuda acceleration version is 11.2. The Python libraries used include: numpy 1.19.4, Pillow 8.0.0, matplotlib 3.3.2, scikit-learn 0.23.2, Pytorch 1.12, etc.

[0125] The hardware platform requirements are as follows:

[0126] The processor computing power must not be lower than AMD Ryzen 9 5900HS with Radeon Graphics 3.30GHz;

[0127] The graphics card should be no less than NVIDIA GeForce RTX 3060;

[0128] Memory should be no less than 16GB;

[0129] Solid-state drives should be no less than 1TB;

[0130] The operating system is Windows 10 with NVIDIA CUDA Drivers and Toolkit.

[0131] In summary, the beneficial effects of the embodiments of the present application are:

[0132] 1. The problem of lack of autonomous joint use of multi-payload detection and multi-modal data when UAVs detect sea surface targets has been solved.

[0133] 2. The GAN-FM model structure has been improved to realize the fusion of SAR images, visible light images, and infrared images.

[0134] 3. Apply the improved GAN-FM to fuse features and results respectively to improve the efficiency of sea surface target detection and recognition.

[0135] 4. The application of YOLOv10 solves the problem of serious performance degradation of existing detection and recognition models under conditions of limited computing power, reduces the amount of target detection and recognition calculations, and improves the real-time performance of the algorithm.

[0136] In this application, the terms "first", "second", and "third" are used for descriptive purposes only and are not to be understood as indicating or implying relative importance; the term "plurality" refers to two or more, unless otherwise expressly defined. Terms such as "installed", "connected", "connected", and "fixed" should be understood in a broad sense. For example, "connected" can be a fixed connection, a detachable connection, or an integral connection; "connected" can be a direct connection or an indirect connection through an intermediate medium. For those of ordinary skill in the art, the specific meanings of the above terms in this application can be understood according to the specific circumstances.

[0137] In the description of the present application, it should be understood that the terms "upper", "lower", "front", "rear", etc., indicating directions or positional relationships are based on the directions or positional relationships shown in the accompanying drawings, and are only for the convenience of describing the present application and simplifying the description, rather than indicating or implying that the system or module referred to must have a specific direction, be constructed and operated in a specific orientation, and therefore, should not be understood as a limitation on the present application.

[0138] In the description of this specification, the description of the terms "one embodiment", "some embodiments", "specific embodiments", etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.

[0139] The above are only preferred embodiments of the present application and are not intended to limit the present application. For those skilled in the art, the present application may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. A method for detecting and identifying sea surface targets, characterized in that: include: Acquire detection data of the UAV and pre-process the detection data, wherein the detection data includes visible light images, infrared images, and radar SAR images; Constructing an improved feature fusion model based on the GAN-FM model, performing feature fusion on the preprocessed detection data through the improved feature fusion model to obtain a fused image, wherein the improved feature fusion model includes a generator, an infrared discriminator, a visible light discriminator, and a radar SAR image discriminator; A YOLOv10 detection and recognition network is constructed, and sea surface target detection and recognition is performed on the fused image through the YOLOv10 detection and recognition network to obtain a recognition result.

2. The method for detecting and identifying sea targets according to claim 1, characterized in that: The acquiring of the detection data of the UAV and preprocessing the detection data include: Acquire detection data of the UAV, and perform noise reduction on the detection data by using a wavelet threshold noise reduction method and a double-tree complex wavelet noise reduction method to obtain noise-reduced data; Performing data enhancement on the noise reduction data by using multiple data enhancement methods to obtain enhanced data; The enhanced data is subjected to equalization processing to obtain equalized sample data.

3. The method for detecting and identifying sea targets according to claim 2, characterized in that: The performing equalization processing on the enhanced data to obtain equalized sample data includes: Performing sliding window cropping on multiple samples in the enhanced data to obtain a sample set with a resolution lower than the original image; The number of target samples containing the target in the sample set is obtained, and samples not containing the target in the sample set are randomly sampled according to the number of target samples to obtain balanced sample data.

4. The method for detecting and identifying sea targets according to claim 2, characterized in that: The data enhancement method includes one of the following: flip transformation, rotation transformation, linear brightness shift, non-linear brightness shift, contrast image adjustment, image perspective, and cropping and splicing.

5. The method for detecting and identifying sea surface targets according to any one of claims 1 to 4, characterized in that: The step of performing feature fusion on the preprocessed detection data by using the improved feature fusion model to obtain a fused image includes: Performing a channel stitching operation on the visible light image, the infrared image, and the radar SAR image in a matching state, extracting fusion features through a generator, and generating a fused image; Inputting the fused image into an infrared discriminator, a visible light discriminator and a radar SAR image discriminator respectively to distinguish the fused image from the original image; Nash equilibrium is achieved by averaging the outputs of the infrared discriminator, the visible light discriminator, the radar SAR image discriminator and the batch.

6. The method for detecting and identifying sea targets according to claim 2, characterized in that: The noise reduction of the detection data by using the wavelet threshold noise reduction method and the double-tree complex wavelet noise reduction method to obtain the noise reduction data includes: Performing complex wavelet decomposition on the detection data; Calculate the energy of wavelet coefficients at multiple scales, obtain the energy of the correlation coefficients at the corresponding scales, normalize the correlation coefficients of the wavelet coefficients, and obtain the coefficient threshold; Determining whether the wavelet coefficient is less than the coefficient threshold; If so, the wavelet coefficients are retained and the retained wavelet coefficients are reconstructed.

7. The method for detecting and identifying sea targets according to claim 6, characterized in that: The detection data includes photoelectric detection data and radar detection data.

8. A sea surface target detection and recognition system, characterized in that: include: An acquisition module (110) is used to acquire detection data of the UAV and pre-process the detection data, wherein the detection data includes visible light images, infrared images and radar SAR images; A feature fusion module (120) is used to construct an improved feature fusion model based on the GAN-FM model, and to perform feature fusion on the pre-processed detection data through the improved feature fusion model to obtain a fused image, wherein the improved feature fusion model includes a generator, an infrared discriminator, a visible light discriminator, and a radar SAR image discriminator; The detection and recognition module (130) is used to construct a YOLOv10 detection and recognition network, and perform sea surface target detection and recognition on the fused image through the YOLOv10 detection and recognition network to obtain a recognition result.

9. A sea surface target detection and recognition system, characterized in that: include: A memory (300) and a processor (400), wherein the memory (300) stores a program or instruction that can be run on the processor (400), and when the processor (400) executes the program or the instruction, the steps of the sea surface target detection and recognition method according to any one of claims 1 to 7 are implemented.

10. A readable storage medium having a program or instruction stored thereon, characterized in that: When the program or the instruction is executed by the processor, the steps of the sea surface target detection and recognition method as described in any one of claims 1 to 7 are implemented.