Synthetic aperture sonar image recognition method and system
Through the combination of Real-ESRGAN and SimCLR models, super-resolution reconstruction and feature extraction of high-quality sonar images are achieved, solving the problem of scarcity and blur of samples in synthetic aperture sonar image recognition, and improving the recognition accuracy and generalization performance.
Patent Information
- Application Number
- CN202510306230.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-14
- Publication Date
- 2025-08-05
- Estimated Expiration
- 2045-03-14
AI Technical Summary
In the prior art, synthetic aperture sonar image recognition has scarce samples and blurred samples, resulting in low classification model recognition accuracy, especially under the conditions of unbalanced sparse data sets, which is prone to overfitting, affecting the classification accuracy in actual applications.
Super-resolution reconstruction is used for the Real-ESRGAN model, combined with the fusion attention mechanism and transfer learning of the SimCLR model, features are extracted through the channel attention and spatial attention mechanism to realize the recognition and classification of high-quality sonar images.
It improves the accuracy of sonar image recognition and the generalization ability of the model, effectively solves the problem of sample scarcity and fuzziness, and improves the recognition accuracy of the classification model.
Smart Images

Figure CN120431361A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer vision, and in particular relates to a synthetic aperture sonar image recognition method and system. Background Art
[0002] Synthetic Aperture Sonar (SAS) is a leading method for long-range, high-resolution underwater imaging. Small-sample recognition based on SAS images holds significant application value. Due to its high-precision underwater imaging capabilities and high detection efficiency, SAS is widely used in a variety of marine applications, including underwater resource exploration and development, military geographic information acquisition, underwater battlefield surveillance and intelligence gathering, and underwater early warning. Because SAS is not limited by narrow azimuth beams, it can operate at lower frequencies and has a long detection range. It can effectively distinguish mines from surrounding rocks and other unidentified objects in complex seabed environments, providing precise target information for subsequent mine clearance operations. As a result, SAS is gaining increasing attention in the field of marine exploration. With the advancement of computer vision technology, deep learning models can accurately and efficiently identify sonar target images. However, due to limitations in sonar imaging mechanisms and acquisition costs, SAS images have fewer target pixels and a limited sample size compared to optical images.
[0003] To address the issue of small sample sizes, several approaches have been proposed, including increasing the sample size of target images by applying scattering and geometric transformations to sonar images; using CGANs, a data-driven model, to generate realistic side-scan sonar measurement data from environmental inputs; increasing the sample size by fusing optical target images with existing sonar background images; and enhancing the sample size of all side-scan sonar images using techniques such as image transformation, noise addition, and synthetic data generation. Finally, the SGAN model has been proposed, effectively enhancing the sample size. While these methods, such as data augmentation, automatic data expansion, and synthesizing new samples using simulation software or generative models, can increase the sample size of sonar target images, they all rely on raw sonar images, and problems such as blurry, low-quality, and inconspicuous regions of interest remain unresolved. This makes effective feature extraction difficult during recognition and results in low recognition accuracy.
[0004] To address the issue of too few pixels, a method was proposed to use style transfer to transfer an optical target image to a sonar background image to generate a simulated image. Furthermore, an improved CycleGAN network was proposed to successfully generate high-quality sonar images using the optical image as a guide image. A rasterization-based forward-looking sonar image simulator was proposed, combining ray tracing and rasterization techniques while incorporating the effects of sound pressure attenuation and multipath propagation into the model to improve the realism of the simulated sound images. Finally, a Transformer-based generative adversarial network was proposed for sonar image despeckling, enhancing the clarity and quality of sonar images. While these methods improve sonar image quality by synthesizing pseudo-sonar samples through the simulator and using a generative adversarial network to denoise the blurred image, they still suffer from domain differences in texture and subtle features, failing to alleviate the few-shot problem. Furthermore, due to uneven data distribution and lack of clear intra-class differences, classification bias is common in the classification of imbalanced sonar datasets, whereby classes with many samples are more likely to be identified as correct, replacing classes with few samples. This phenomenon not only weakens the model's generalization performance but also severely impacts its classification accuracy in practical applications. At the same time, under the conditions of unbalanced sparse data samples, directly using deep convolutional neural networks for sonar image recognition will produce serious overfitting, further leading to poor recognition accuracy. Summary of the Invention
[0005] In view of this, embodiments of the present invention provide a synthetic aperture sonar image recognition method and system for solving the current problem of low recognition accuracy of classification models caused by sample scarcity and sample ambiguity in sonar image recognition.
[0006] In a first aspect of an embodiment of the present invention, a synthetic aperture sonar image recognition method is provided, comprising: Acquire synthetic aperture sonar images; The sonar image is preprocessed using the Real-ESRGAN model based on deep learning to obtain high-quality reconstructed sonar images. Inputting the high-quality sonar image into the SimCLR model, and performing feature extraction on the sonar image using a fusion attention module in the SimCLR model encoder, wherein the fusion attention module includes a channel attention mechanism and a spatial attention mechanism; Sonar images are identified and classified through transfer learning based on extracted features.
[0007] In a second aspect of an embodiment of the present invention, a synthetic aperture sonar image recognition system is provided, comprising: An image acquisition module, used to acquire synthetic aperture sonar images; The image reconstruction module is used to preprocess sonar images using the Real-ESRGAN model based on deep learning to obtain high-quality reconstructed sonar images; A feature extraction module is used to input the high-quality sonar image into the SimCLR model and perform feature extraction on the sonar image through a fusion attention module in the SimCLR model encoder, wherein the fusion attention module includes a channel attention mechanism and a spatial attention mechanism; The image recognition module is used to identify and classify sonar images through transfer learning based on extracted features.
[0008] In a third aspect of an embodiment of the present invention, an electronic device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable by the processor, wherein the processor implements the steps of the method described in the first aspect of the embodiment of the present invention when executing the computer program.
[0009] In a fourth aspect of an embodiment of the present invention, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the method provided in the first aspect of the embodiment of the present invention are implemented.
[0010] In an embodiment of the present invention, a Real-ESRGAN model is adopted and a super-resolution reconstruction algorithm is introduced to achieve pixel expansion and structure completion of blurred targets. The contrastive learning mechanism is utilized and the SimCLR model is adopted to achieve high-accuracy classification of small sample images after super-resolution reconstruction. At the same time, a fusion attention module is introduced within the contrastive learning framework, thereby reducing the parameters of the classification model while improving the recognition accuracy of sonar images and enhancing the generalization ability of the classification model. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments or descriptions of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0012] Figure 1 A schematic flow chart of a synthetic aperture sonar image recognition method provided by one embodiment of the present invention; Figure 2 A schematic diagram of the SimCLR model structure provided by one embodiment of the present invention; Figure 3 A schematic structural diagram of a synthetic aperture sonar image recognition system provided by one embodiment of the present invention; Figure 4The present invention provides a schematic structural diagram of an electronic device according to an embodiment of the present invention. DETAILED DESCRIPTION
[0013] In order to make the purpose, features, and advantages of the present invention more obvious and easy to understand, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described below are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.
[0014] It should be understood that the terms "including" and similar expressions in the specification, claims, and drawings of the present invention are intended to cover non-exclusive inclusions. For example, a process, method, system, or apparatus comprising a series of steps or units is not limited to the listed steps or units. Furthermore, the terms "first" and "second" are used to distinguish between different objects and are not intended to describe a specific order.
[0015] See also Figure 1 , a flow chart of a synthetic aperture sonar image recognition method provided by an embodiment of the present invention includes: S101, acquiring a synthetic aperture sonar image; Synthetic aperture sonar (SAS) is a two-dimensional imaging sonar that uses a uniformly moving acoustic array to form a large virtual synthetic aperture, thereby improving the sonar's lateral resolution. The two-dimensional images captured by SAS can be used for target detection and recognition.
[0016] Among them, the synthetic aperture sonar images are cleaned, duplicate data is removed, missing, erroneous and abnormal data are processed, and data format conversion is performed.
[0017] S102, preprocessing the sonar image using the Real-ESRGAN model based on deep learning to obtain a reconstructed high-quality sonar image; Real-ESRGAN (Realistic Enhanced Super-Resolution Generative Adversarial Networks) is a super-resolution reconstruction framework based on deep learning. It introduces generative adversarial networks (GANs) into super-resolution reconstruction, making the generated images more realistic in visual effects.
[0018] The Real-ESRGAN model training in this embodiment is divided into two stages: the first stage uses the mean absolute error loss L1 loss to train a peak signal-to-noise ratio-oriented model, and the second stage uses a combination of L1 loss, perceptual loss, and generative adversarial network loss GAN loss to train Real-ESRGAN.
[0019] The total loss function is: ; Where λ and γ are weight coefficients, and λ=1, γ=0.1; ; Where, represents the true value, represents the predicted value, n represents the number of samples; ; Where, represents the features extracted from the pre-trained network, represents the characteristics of the target image, Represents the features of the generated image; ; Among them, BCELoss represents binary cross entropy loss, D represents the discriminator, To generate the output of the model, ones represents the full 1 matrix of the real image label.
[0020] Optionally, Real-ESRGAN model training includes: Performing degradation processing on a preset proportion of the training set images, wherein the degradation processing includes one or more of blurring, resolution downsampling, random noise, and image compression; The generator reconstructs the degraded image, mixes the reconstructed image with the real image in the training set and inputs it into the discriminator, which then judges the mixed image. The Real-ESRGAN model is optimized through adversarial training between the generator and the discriminator.
[0021] During training, Real-ESRGAN uses a series of complex image degradation techniques, such as blurring, resolution downsampling, random noise addition, and image compression, to maximize the reproduction of real-world scene information and reconstruct a set of images with multiple distortion factors and low resolution for training. The n-order model in the Real-ESRGAN model contains n repeated degradation processes, each of which follows the classic model: ; Where D represents the entire degradation process, y represents the high-definition image, and x represents the low-definition image after degradation.
[0022] The blur operation is usually represented by a linear blur filter convolution, and for a Gaussian blur kernel k with a kernel size of 2t+1, its element values are sampled from a Gaussian distribution.
[0023] In the Real-ESRGAN model, the generator's primary task is to amplify and repair degraded low-resolution images. After low-quality input data is fed into the generator, the reconstructed high-quality image is fused with its corresponding real-world image to produce a more realistic, high-quality image that resembles actual scene conditions. Finally, the reconstructed high-quality image is fed into the discriminator for evaluation.
[0024] The Real-ESRGAN model uses a spectrally normalized U-Net as the discriminator, which excels in reducing artifacts and enhancing detail. The core of the U-Net discriminator is to distinguish which of a pair of input images is real and which is reconstructed by the generator. Through adversarial training between the generator and the discriminator, Real-ESRGAN continuously optimizes its reconstruction algorithm, making the generated high-resolution images increasingly difficult for the discriminator to distinguish, thereby gradually approaching or even exceeding the quality of real images.
[0025] S103: Input the high-quality sonar image into the SimCLR model, and perform feature extraction on the sonar image using a fusion attention module in the SimCLR model encoder, wherein the fusion attention module includes a channel attention mechanism and a spatial attention mechanism; SimCLR (Simple Contrastive Learning of Visual Representations) is a self-supervised contrastive learning framework that focuses on learning high-quality image representations in an unsupervised manner. A SimCLR model typically consists of a data augmentation module, an encoder, and a loss function.
[0026] In some embodiments, the encoder is a Resnet-18 network.
[0027] The SimCLR model itself has a data enhancement module, and there is no need to perform additional data enhancement on the dataset. In this module, two different data enhancement methods are randomly selected for each data sample, such as cropping, brightness change, etc. A fusion attention module is embedded in the encoder, which can effectively extract channel and spatial dimension features, suppress redundant information, and reduce model complexity. Based on the similarity of image feature vectors, the contrast loss is calculated, and the InfoNCE loss function is used to enhance the similarity between similar samples and obtain a more efficient representation method. Figure 2 As shown in the figure, after the sonar image reconstructed by the Real-ESRGAN model is input into SimCLR, two different data enhancement methods are randomly selected to obtain two related images. x i 、 x j , the two images are fed into the encoder, and the result is a vector converted into the same dimensional space h i and h j , the image feature value z can be extracted by fusing the attention module i 、 z j , the feature vector in the encoder can be input into the downstream classification task to identify and classify images based on transfer learning.
[0028] The fused attention module includes both channel-based and spatial-based attention mechanisms. This channel-based attention mechanism, known as Efficient Channel Attention (ECA), inherits the core concept of the SE module (SENet) to improve model performance by modeling the importance of feature channels, while also streamlining and innovating on this foundation. Through adaptive adjustment of one-dimensional convolutions, it intelligently focuses on the channel feature combinations that most significantly impact model performance, dynamically adjusting the weights of different channels to enhance important feature channels and suppress redundant information.
[0029] The Spatial Attention Mechanism (SAM) prioritizes information in the spatial dimension of feature maps, particularly focusing on interactions between different pixels. This mechanism enables the model to consciously focus on locations rich in key feature information and ignore areas with less significant contributions. By dynamically adjusting the importance of each spatial location in the input image, it not only improves the relevance and effectiveness of feature representation, but also effectively reduces interference from irrelevant information. Using the spatial attention mechanism with adaptive multi-scale convolution kernels (with kernel sizes of 3, 5, or 7), it not only effectively extracts target features at different scales but also enhances the model's adaptability.
[0030] By combining the spatial attention mechanism with adaptive multi-scale convolution kernels with an efficient channel attention module, it can not only capture key area information, but also effectively enhance the feature contribution of important channels, achieving multi-dimensional attention enhancement, thereby improving model performance and better solving image recognition accuracy problems in complex backgrounds.
[0031] During SimCLR model training, model performance is evaluated based on the Top-1 accuracy function and the InfoNCE loss function. Top-1 accuracy indicates the proportion of the correct category for which the model predicts the highest probability.
[0032] Optionally, the encoder of the SimCLR model is connected to a classification head for the image classification task, where the classification head contains one or more fully connected layers for mapping features to the category space.
[0033] The classification head is the part of the model that specifically maps the learned feature representations to category predictions. It is usually located at the end of the model, immediately following the feature extraction layer.
[0034] S104: Identify and classify the sonar image through transfer learning based on the extracted features.
[0035] The features extracted by the fused attention module are used through transfer learning to perform downstream classification tasks on sonar image data. Transfer learning is a machine learning method that applies models or knowledge from a source task to a target task. In this embodiment, features extracted using an encoder are directly applied to an image classification model in the target domain through transfer learning to perform image recognition and classification. For example, the encoder of the SimCLR model is connected to the classification head of a mature classification model to classify target features in the image.
[0036] In this implementation, the Real-ESRGAN model based on deep learning is used to preprocess sonar images to obtain high-quality sonar images. The reconstructed high-quality sonar images are then extracted through the fusion attention module in SimCLR, and the sonar images are recognized and classified based on transfer learning, which can effectively improve the recognition accuracy and generalization performance of the model.
[0037] To address image blur, a super-resolution reconstruction algorithm was introduced to achieve pixel expansion and structural completion of blurred images. To address sample scarcity, a small-sample classification mechanism based on contrastive learning was used to achieve high-accuracy classification of small-sample sonar images reconstructed after super-resolution. To address the long-tail distribution problem, a fusion attention module was introduced within the contrastive learning framework to focus on regions of interest, further improving sonar image classification performance while reducing model parameters. This improved the existing recognition accuracy of small-sample sonar images, significantly increasing recognition accuracy compared to traditional convolutional neural network classification models.
[0038] In some embodiments, during SimCLR model training, supervised feature extraction is first performed on the target data, and the extracted features are then used in downstream classification tasks. After obtaining the encoder trained through contrastive learning, the final fully connected layer is replaced with the classification head for the classification task.
[0039] As can be understood, to verify the effectiveness of supervised learning for feature extraction, the encoder's convolutional layers were frozen, meaning only the classification head was trained without the convolutional neural network. Without the improved attention mechanism, the encoder-frozen model achieved 72.1% accuracy on the training set, despite relatively high loss on the training set. However, its accuracy on the validation set was consistent with the training set and showed an upward trend. This demonstrates that supervised contrastive learning can leverage the superior feature extraction capabilities of upstream tasks for downstream transfer learning, significantly improving performance in image classification tasks. The model's suboptimal performance on the training set was attributed to the frozen convolutional layers, which prevented parameter updates and resulted in a significant loss contamination of the input classification feature vectors. Subsequently, the encoder's convolutional layers were unfrozen and training experiments were conducted. The model achieved an accuracy of 79.2% on the validation set. To further improve classification accuracy, a hybrid attention mechanism and the SimCLR algorithm were applied for feature extraction. On the test set, the hybrid attention mechanism and SimCLR algorithm were used for target image recognition, achieving an accuracy of 87.5%, further improving model performance. Overall, all models observed a rapid reduction and stabilization of loss during the initial phase. The second stage may initially show significant loss fluctuations, but these fluctuations eventually stabilize. Notably, the super-resolution reconstruction + hybrid attention + SimCLR model exhibits the most stable loss rate and the smallest overall variance in both training stages, indicating a robust and consistent training process.
[0040] To address the scarcity, poor quality, and uneven classification of sonar image samples, a method for image recognition that combines super-resolution reconstruction with supervised contrastive learning, leveraging the superior image processing capabilities of super-resolution reconstruction technology and the small-sample recognition capabilities of supervised contrastive learning, was developed. This method generates datasets of higher quality than the original sonar images, effectively improving sonar image recognition accuracy under small-sample conditions. Comparative experiments show that the introduction of supervised contrastive learning effectively suppresses model overfitting. Experimental results demonstrate that the combined super-resolution reconstruction and supervised contrastive learning method improves recognition accuracy by 4.17% compared to existing DCNNs, particularly in the task of classifying synthetic aperture sonar images.
[0041] It should be understood that the sequence numbers of the steps in the above embodiments do not imply a specific order of execution; the order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0042] Figure 3 A schematic diagram of the structure of a synthetic aperture sonar image recognition system provided in an embodiment of the present invention, the system comprising: An image acquisition module 310 is used to acquire synthetic aperture sonar images; An image reconstruction module 320 is configured to pre-process the sonar image using a Real-ESRGAN model based on deep learning to obtain a reconstructed high-quality sonar image; The preprocessing of the sonar image using the Real-ESRGAN model based on deep learning to obtain a high-quality sonar image includes: Performing degradation processing on a preset proportion of the training set images, wherein the degradation processing includes one or more of blurring, resolution downsampling, random noise, and image compression; The generator reconstructs the degraded image, mixes the reconstructed image with the real image in the training set and inputs it into the discriminator, which then judges the mixed image. The Real-ESRGAN model is optimized through adversarial training between the generator and the discriminator.
[0043] A feature extraction module 330 is configured to input the high-quality sonar image into the SimCLR model and perform feature extraction on the sonar image using a fusion attention module in the SimCLR model encoder, wherein the fusion attention module includes a channel attention mechanism and a spatial attention mechanism; The image recognition module 340 is used to recognize and classify sonar images through transfer learning based on the extracted features.
[0044] Optionally, the feature extraction module 330 further includes: The model evaluation unit is used to evaluate model performance based on the Top-1 accuracy loss function and the InfoNCE loss function during SimCLR model training.
[0045] The encoder of the SimCLR model is connected to the classification head of the classification task, and the classification head contains one or more fully connected layers for mapping features to the category space.
[0046] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described systems and modules can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0047] Figure 4 FIG. 1 is a schematic diagram of the structure of an electronic device provided by an embodiment of the present invention. The electronic device is used for synthetic aperture sonar image recognition and classification. Figure 4 As shown, the electronic device 4 of this embodiment includes: a memory 410, a processor 420 and a system bus 430, wherein the memory 410 includes an executable program 4101 stored thereon. It can be understood by those skilled in the art that Figure 4The electronic device structure shown in the figure does not constitute a limitation to the electronic device, and may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently.
[0048] The following combination Figure 4 A detailed introduction to the various components of electronic equipment: Memory 410 can be used to store software programs and modules. Processor 420 executes the software programs and modules stored in memory 410 to perform various functional applications and data processing of the electronic device. Memory 410 may primarily include a program storage area and a data storage area. The program storage area may store an operating system and at least one application required for a function (such as sound playback or image playback). The data storage area may store data generated based on the use of the electronic device (such as cached data). Memory 410 may also include high-speed random access memory and non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state memory device.
[0049] Memory 410 includes an executable program 4101 for the interface generation method. The executable program 4101 can be divided into one or more modules / units. These modules / units are stored in memory 410 and executed by processor 420 to implement sonar image recognition, etc. The one or more modules / units can be a series of computer program instruction segments capable of performing specific functions, and these instruction segments are used to describe the execution process of the executable program 4101 in the electronic device 4. For example, the executable program 4101 can be divided into functional modules such as an image acquisition module, an image reconstruction module, a feature extraction module, and an image recognition module.
[0050] Processor 420 is the control center of the electronic device, connecting the various components of the entire electronic device using various interfaces and lines. By running or executing software programs and / or modules stored in memory 410 and accessing data stored in memory 410, it performs various functions of the electronic device and processes data, thereby monitoring the overall status of the electronic device. Optionally, processor 420 may include one or more processing units; preferably, processor 420 may integrate an application processor and a modem processor, wherein the application processor primarily processes the operating system, application programs, etc., and the modem processor primarily handles wireless communications. It is understood that the modem processor described above may not be integrated into processor 420.
[0051] The system bus 430 connects the various functional components within the computer and can transmit data, address information, and control information. It can be a PCI bus, an ISA bus, a CAN bus, or other types. Instructions from the processor 420 are transmitted to the memory 410 via the bus, and the memory 410 feeds data back to the processor 420. The system bus 430 is responsible for the exchange of data and instructions between the processor 420 and the memory 410. Of course, the system bus 430 can also connect to other devices, such as network interfaces and display devices.
[0052] In an embodiment of the present invention, the executable program executed by the processing 420 included in the electronic device includes: Acquire synthetic aperture sonar images; The sonar image is preprocessed using the Real-ESRGAN model based on deep learning to obtain high-quality reconstructed sonar images. Inputting the high-quality sonar image into the SimCLR model, and performing feature extraction on the sonar image using a fusion attention module in the SimCLR model encoder, wherein the fusion attention module includes a channel attention mechanism and a spatial attention mechanism; Sonar images are identified and classified through transfer learning based on extracted features.
[0053] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described systems, devices, and modules can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0054] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant description of other embodiments.
[0055] As described above, the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that the technical solutions described in the above embodiments can still be modified, or some of the technical features thereof can be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A synthetic aperture sonar image recognition method, characterized in that: include: Acquire synthetic aperture sonar images; The sonar image is preprocessed using the Real-ESRGAN model based on deep learning to obtain high-quality reconstructed sonar images. Inputting the high-quality sonar image into the SimCLR model, and performing feature extraction on the sonar image using a fusion attention module in the SimCLR model encoder, wherein the fusion attention module includes a channel attention mechanism and a spatial attention mechanism; Sonar images are identified and classified through transfer learning based on extracted features.
2. The method according to claim 1, characterized in that The sonar image is preprocessed using the deep learning-based Real-ESRGAN model to obtain a reconstructed high-quality sonar image, including: Performing degradation processing on a preset proportion of the training set images, wherein the degradation processing includes one or more of blurring, resolution downsampling, random noise, and image compression; The generator reconstructs the degraded image, mixes the reconstructed image with the real image in the training set and inputs it into the discriminator, which then judges the mixed image. The Real-ESRGAN model is optimized through adversarial training between the generator and the discriminator.
3. The method according to claim 1, characterized in that Before inputting the high-quality sonar image into the SimCLR model, the method further includes: In SimCLR model training, the model performance is evaluated based on the Top-1 accuracy function and the InfoNCE loss function.
4. The method according to claim 1, wherein Before inputting the high-quality sonar image into the SimCLR model, the method further includes: Connect the encoder of the SimCLR model to the classification head of the classification task, which contains one or more fully connected layers for mapping features to the category space.
5. A synthetic aperture sonar image recognition system, characterized in that: include: An image acquisition module, used to acquire synthetic aperture sonar images; The image reconstruction module is used to preprocess sonar images using the Real-ESRGAN model based on deep learning to obtain high-quality reconstructed sonar images; A feature extraction module is used to input the high-quality sonar image into the SimCLR model and perform feature extraction on the sonar image through a fusion attention module in the SimCLR model encoder, wherein the fusion attention module includes a channel attention mechanism and a spatial attention mechanism; The image recognition module is used to identify and classify sonar images through transfer learning based on extracted features.
6. The system according to claim 5, characterized in that The sonar image is preprocessed using the deep learning-based Real-ESRGAN model to obtain high-quality sonar images, including the following steps: Performing degradation processing on a preset proportion of the training set images, wherein the degradation processing includes one or more of blurring, resolution downsampling, random noise, and image compression; The generator reconstructs the degraded image, mixes the reconstructed image with the real image in the training set and inputs it into the discriminator, which then judges the mixed image. The Real-ESRGAN model is optimized through adversarial training between the generator and the discriminator.
7. The system according to claim 5, characterized in that The feature extraction module also includes: The model evaluation unit is used to evaluate model performance based on the Top-1 accuracy function and InfoNCE loss function during SimCLR model training.
8. The system according to claim 5, wherein: Before inputting the high-quality sonar image into the SimCLR model, the method further includes: Connect the encoder of the SimCLR model to the classification head of the classification task, which contains one or more fully connected layers for mapping features to the category space.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the steps of the synthetic aperture sonar image recognition method according to any one of claims 1 to 4 are implemented.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed, the steps of the synthetic aperture sonar image recognition method according to any one of claims 1 to 4 are implemented.
Citation Information
Patent Citations
Sonar image recognition method and device, electronic equipment and storage medium
CN113807324A
Fabric image data enhancement method based on deep learning
CN118314063A
Contrast learning fused generative adversarial network sonar image denoising method and device
CN119599900A