A sonar target recognition system based on residual network

By integrating multiple types of spectrum information and sonar propagation physical parameters into the residual network sonar target recognition system, the problem of insufficient sonar recognition accuracy in complex ocean environments is solved, and efficient feature extraction and stable target recognition are achieved.

CN120428212BActive Publication Date: 2025-09-12SHENYANG LIAOHAI EQUIP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510928206.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-07
Publication Date
2025-09-12
Estimated Expiration
2045-07-07

AI Technical Summary

Technical Problem

Existing sonar target recognition methods lack recognition accuracy in complex ocean environments, and their reliance on high-quality labeled data leads to poor training results. They make it difficult to effectively extract the complex time-frequency distribution characteristics in sonar images and adapt to changes in echo energy attenuation and frequency attenuation.

Method used

A sonar target recognition system based on residual networks is adopted to integrate multiple types of spectral information and sonar propagation physical parameters. Through spectrum enhancement and label credibility scoring mechanism, a multi-channel spectrum image is constructed, and feature extraction and joint training are performed to improve the robustness and generalization ability of the model.

Benefits of technology

It improves the accuracy and robustness of sonar target recognition, enhances the model's ability to learn weakly labeled samples, and improves recognition accuracy and the ability to adapt to complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120428212B_ABST
    Figure CN120428212B_ABST
Patent Text Reader

Abstract

The present invention discloses a sonar target recognition system based on a residual network, comprising: a sonar preprocessing module for collecting and preprocessing raw sonar echo signals; a spectrum construction module for generating multi-channel spectrum images; a spectrum enhancement module for performing spectrum enhancement processing; a residual feature extraction module for extracting sonar target feature representations; a label screening module for calculating label credibility scores and constructing a label training set; a joint training module for merging the label training set with real annotated samples to form a joint training set and performing iterative optimization of the SonarResBlock model; and a recognition output module for performing inference recognition using the updated model and outputting the category label of the sonar target. The present invention integrates multi-category spectrum construction with a residual attention mechanism to achieve accurate sonar target recognition, with the advantages of comprehensive modeling, strong robustness, and efficient training.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of intelligent sonar recognition, and in particular to a sonar target recognition system based on a residual network. Background Art

[0002] Currently, in the field of sonar target recognition, a model based on traditional feature extraction and classification methods is generally used for processing, that is, the sonar echo signal is first converted into a spectrum diagram through short-time Fourier transform and wavelet transform, and then support vector machine, random forest or shallow neural network is used for target recognition. This type of method is relatively simple in structure and has a low degree of dependence on data. However, due to the inability to effectively extract the complex time-frequency distribution characteristics in sonar images, especially in non-ideal ocean environments such as multipath interference, severe echo attenuation, and unstable signal-to-noise ratio, its recognition accuracy and generalization ability have obvious bottlenecks. In order to improve the recognition performance, some studies introduced convolutional neural networks for automatic feature extraction, but most of them used single-channel spectrum as input, ignoring multidimensional features such as envelope diagram, amplitude and phase information, as well as the influence of physical conditions of sonar propagation on image characteristics. As a result, the model has poor adaptability to changes in echo energy attenuation and frequency attenuation, and it is still difficult to work stably in complex environments.

[0003] In addition, existing methods generally rely on high-quality labeled data for supervised learning. In the sonar field, the high cost of manual labeling and large subjective errors lead to limited available samples, which further restricts the training and performance of the model. Some labeling methods attempt to alleviate this problem, but do not consider the dynamic evaluation of label credibility and sample screening mechanism, which easily introduces erroneous labels to interfere with training and affect model stability and convergence. Therefore, there is an urgent need for a sonar recognition method that integrates multi-channel spectral information, combines the physical mechanism of sonar propagation, and introduces a dynamic label credibility screening strategy to solve the problems of insufficient recognition accuracy and high risk of label introduction in complex environments in existing technologies. Summary of the Invention

[0004] One purpose of the present invention is to propose a sonar target recognition system based on residual network. The present invention integrates multi-class spectrum construction, spectrum enhancement and residual attention mechanism, and realizes accurate recognition of sonar targets in complex ocean environments by introducing sonar propagation physical parameters and label credibility scoring method. This method not only improves the physical consistency of feature extraction, but also enhances the model's learning ability for weakly labeled samples. It has the significant advantages of comprehensive modeling, strong robustness and efficient training.

[0005] A sonar target recognition system based on a residual network according to an embodiment of the present invention includes:

[0006] Sonar preprocessing module, used to collect raw sonar echo signals and perform preprocessing;

[0007] The spectrum construction module is used to convert the pre-processed sonar echo signal into multiple types of spectrum images, and perform weighted fusion on them in combination with the physical parameters of sonar propagation to generate multi-channel spectrum images;

[0008] A spectrum enhancement module is used to perform spectrum enhancement processing on the multi-channel spectrum image, wherein the spectrum enhancement includes logarithmic spectrum compression, pseudo color mapping and local contrast enhancement;

[0009] The residual feature extraction module is used to input the enhanced spectrum image into the SonarResBlock model to extract the sonar target feature representation;

[0010] The label filtering module is used to calculate the label credibility score based on the sonar target feature representation and construct a label training set for label samples with scores greater than the label filtering threshold;

[0011] The joint training module is used to merge the labeled training set with the real annotated samples to form a joint training set and perform iterative optimization of the SonarResBlock model;

[0012] The recognition output module is used to perform inference recognition using the updated model and output the category label of the sonar target.

[0013] Optionally, modules can be connected using the following methods:

[0014] Collecting original sonar echo signals and preprocessing the original sonar echo signals to obtain preprocessed sonar echo signals;

[0015] The spectral features of the pre-processed sonar echo signal are constructed to generate time-frequency diagrams, amplitude-phase diagrams, and envelope spectrograms. Combined with the physical parameters of sonar propagation, the three types of spectrum diagrams are fused according to channel weighting to construct a multi-channel spectrum image.

[0016] Performing spectrum enhancement processing on the multi-channel spectrum image to obtain an enhanced spectrum image;

[0017] The enhanced spectrum image is input into the SonarResBlock model to extract the key target features in the enhanced spectrum image and obtain the sonar target feature representation;

[0018] Based on the sonar target feature representation, a classification head is used for inference. The label confidence score is calculated based on the category probability distribution, the category KL divergence of the prediction results, and the category frequency offset. A label training set is constructed for label samples with label confidence scores greater than the label screening threshold. This is then merged with the real labeled samples to form a joint training set, and the SonarResBlock model is iteratively optimized.

[0019] The updated SonarResBlock model is used to perform recognition reasoning on the enhanced spectrum image, and the classification head is used to reason about the sonar target feature representation and output the target category label.

[0020] Optionally, the preprocessing includes noise suppression, amplitude normalization and time window truncation.

[0021] Optionally, the sonar propagation physical parameters include propagation delay, frequency attenuation factor and echo energy attenuation.

[0022] Optionally, the constructing of a multi-channel spectrum image includes the following specific steps:

[0023] Perform short-time Fourier transform on the preprocessed sonar echo signal to generate the corresponding time-frequency diagram;

[0024] Perform Hilbert-Huang transform on the pre-processed sonar echo signal to construct the envelope spectrum;

[0025] The frequency domain complex spectrum is obtained by fast Fourier transform on the preprocessed sonar echo signal, and the corresponding amplitude spectrum and phase spectrum are extracted based on the frequency domain complex spectrum. The amplitude spectrum and phase spectrum are then channel-joined to construct the amplitude-phase diagram.

[0026] Normalize the sonar propagation physical parameters to unitless values. For the time-frequency diagram, envelope spectrum, and amplitude-phase diagram, adjust the channel fusion weights based on the normalized sonar propagation physical parameters:

[0027] ;

[0028] in, represents the weighting coefficient of the time-frequency graph channel, represents the weighting coefficient of the envelope spectrum channel, represents the weighting coefficient of the amplitude-phase image channel, Indicates the maximum attenuation reference value, represents the frequency attenuation factor, Indicates target distance and frequency The echo energy attenuation under Indicates frequency, represents the propagation delay;

[0029] The three types of spectrum images are weighted and combined according to the channel dimension to form a multi-channel spectrum image:

[0030] ;

[0031] in, Indicates the Multi-channel spectrum image of the frame, Represents the normalization operation, represents the channel-level concatenation operation, represents the time-frequency diagram, represents the envelope spectrum, Represents the amplitude phase diagram.

[0032] Optionally, obtaining the enhanced spectrum image includes the following specific steps:

[0033] Perform channel-independent logarithmic spectrum compression on the multi-channel spectrum image to obtain a logarithmic compressed image:

[0034] ;

[0035] in, Indicates the Frame in channel , pixel position The logarithmic compression result on , Represents the multi-channel spectrum image Frame in channel , pixel position Pixel intensity on ;

[0036] Performing a pseudo color mapping operation on the logarithmic compressed image to map the grayscale value into a pseudo color image in the RGB color space to obtain a pseudo color spectrum image, wherein the pseudo color spectrum image is an RGB image with three channels;

[0037] Perform local contrast enhancement on the pseudo-color spectrum image using a weighted local statistical contrast enhancement function:

[0038] ;

[0039] in, Indicates the Frame at pixel location The contrast enhancement value on Represents the pseudo-color spectrum image Frame at pixel location The RGB value on Indicates pixel position The local neighborhood mean centered at represents the corresponding local neighborhood standard deviation, It represents the stability constant introduced to prevent the denominator from being zero;

[0040] Normalization is performed on the image after local contrast enhancement to obtain an enhanced spectrum image.

[0041] Optionally, obtaining the sonar target feature representation includes the following specific steps:

[0042] The enhanced spectrum image is input into the SonarResBlock model for feature extraction, wherein the SonarResBlock model is composed of a sonar-adapted residual module;

[0043] The sonar-adaptive residual module includes a main branch and an identity mapping branch, wherein the main branch sequentially includes a first convolutional layer, a batch normalization layer, a ReLU activation function, a second convolutional layer, and a sonar propagation perception attention layer, and the identity mapping branch includes an identity connection or Convolutional connections:

[0044] When the sizes of the main branch output and the enhanced spectrum image are inconsistent, the identity mapping branch selects Convolutional connections;

[0045] When the size of the main branch output and the enhanced spectrum image are consistent, the identity mapping branch selects the identity connection;

[0046] The convolution structure of the main branch adopts a frequency-time bidirectional perception strategy, and the convolution kernel size of the first convolution layer is set to , used to extract contextual information in the frequency dimension, the convolution kernel size of the second convolution layer is set to , used to extract the local continuous change features of the time dimension. The frequency dimension and time dimension correspond to the height and width directions of the enhanced spectrum image respectively;

[0047] Output the feature map after the second convolution:

[0048] ;

[0049] ;

[0050] in, Represents the feature map after the first convolution process, represents the feature map after the second convolution, 、 Represent the first and second convolution operations respectively, represents the batch normalization function, represents the rectified linear unit activation function;

[0051] The feature map after the second convolution is input to the sonar propagation perception attention layer for channel response regulation. The sonar propagation perception attention layer uses the propagation delay, frequency attenuation factor and echo energy attenuation to construct the channel attention weight:

[0052] ;

[0053] in, Indicates the The attention weight of the channel, represents the Sigmoid activation function, represents the frequency attenuation factor, represents the propagation delay, Indicates target distance and frequency The echo energy attenuation under Indicates the maximum energy attenuation reference value, 、 、 Represents the weighting coefficient of each physical quantity;

[0054] The propagation delay, frequency attenuation factor and echo energy attenuation adopt normalized unitless values;

[0055] Apply the attention weight to the feature map after the second convolution to obtain the attention-weighted feature map:

[0056] ;

[0057] in, Indicates the Channel, pixel position The attention-weighted feature value on , Indicates the Channel, pixel position The eigenvalue after the second convolution;

[0058] The attention-weighted feature map is element-wise added to the output of the identity mapping branch to obtain the output feature map of the sonar-adaptive residual module, which serves as the input of the next round of sonar-adaptive residual module and finally outputs the sonar target feature representation.

[0059] Optionally, the iterative optimization of the SonarResBlock model includes the following specific steps:

[0060] Inputting the sonar target feature representation into a classification head to perform label prediction and output a class probability distribution, the classification head using a multi-layer perceptron;

[0061] Based on constructing a historical prior distribution and calculating the KL divergence between the category probability distribution and the historical prior distribution, the historical prior distribution represents the normalized frequency distribution of samples encountered in past training rounds in each category:

[0062] ;

[0063] in, Indicates the The KL divergence between the predicted distribution of a sample and the historical prior distribution, Indicates the The frequency ratio of the class in the historical prior distribution, represents the stability constant introduced to prevent the denominator from being zero, Indicates the The samples belong to The class probability of the class, Indicates the total number of categories of sonar targets;

[0064] Construct a class frequency offset to measure the difference between the label class and the ideal balanced distribution:

[0065] ;

[0066] in, Indicates the The category frequency offset of the sample prediction category, Indicates the main category index of the current prediction, Indicates the frequency ratio of the predicted main category in the historical prior distribution, Indicates the total number of categories;

[0067] The label credibility score is constructed by integrating category probability distribution, distribution deviation and category imbalance factors:

[0068] ;

[0069] in, Indicates the The label confidence score of each sample, represents the KL divergence penalty factor, represents the class frequency offset adjustment factor, Indicates the The class probability of the samples, represents the natural exponential function;

[0070] Set the label filtering threshold to filter label samples in descending order as the number of parameter update rounds decreases:

[0071] ;

[0072] in, Indicates the The label filtering threshold for round parameter update, represents the initial label screening threshold, Indicates the total number of rounds, Represents the exponential factor that controls the rate at which the threshold drops;

[0073] Filter all label samples that meet the label credibility score greater than the label screening threshold to build a label training set, and merge it with the real labeled samples to form a joint training set;

[0074] The parameters of the SonarResBlock model are updated based on the joint training set. The parameter update is performed by minimizing the cross entropy loss function of the SonarResBlock model and performing gradient update using the back propagation algorithm.

[0075] The beneficial effects of the present invention are:

[0076] This paper constructs a sonar target recognition system based on a residual network. Addressing existing issues such as insufficient dimensionality in sonar spectrum modeling, underutilization of physical properties, and variable training sample quality, this paper proposes several technical improvements, enhancing the system's recognition robustness and generalization capabilities in complex underwater acoustic environments. The system introduces three types of spectrum images: time-frequency diagrams, envelope spectrograms, and amplitude-phase diagrams. Furthermore, the system integrates the physical parameters of sonar propagation, such as propagation delay, frequency attenuation factor, and echo energy attenuation, to construct multi-channel spectrum images. This achieves high-dimensional modeling of sonar echo information and enhances the input data's ability to represent complex propagation processes. Incorporating spectrum enhancement strategies such as logarithmic spectrum compression, pseudo-color mapping, and local contrast enhancement effectively improves image quality and the separability of key target features.

[0077] On the other hand, the SonarResBlock model designed in the invention is based on the frequency-time bidirectional perception structure and the sonar propagation perception attention mechanism. It introduces acoustic physical information into the residual connection structure for channel response regulation, which significantly improves the feature extraction module's ability to capture target edges, contours and frequency attenuation features. On this basis, the system also constructs a label credibility scoring mechanism including KL divergence and category frequency offset factor, and combines the dynamic threshold function to complete the joint training data set construction, thereby effectively improving the stability of the model during semi-supervised training. Overall, the present invention forms a synergistic enhancement effect in the three aspects of input modeling, feature extraction and training mechanism, enabling the sonar target recognition system to achieve higher precision and stronger adaptability in an environment with incomplete data and fluctuating signal-to-noise ratio. Target discrimination. BRIEF DESCRIPTION OF THE DRAWINGS

[0078] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings:

[0079] Figure 1 This is a schematic diagram of the overall structure of a sonar target recognition system based on a residual network proposed in the present invention;

[0080] Figure 2 This is a structural diagram of the SonarResBlock model of the sonar target recognition system proposed in the present invention. DETAILED DESCRIPTION

[0081] The present invention will now be described in further detail with reference to the accompanying drawings, which are simplified schematic diagrams that illustrate the basic structure of the present invention in a schematic manner.

[0082] refer to Figure 1-2 , a sonar target recognition system based on residual network, including:

[0083] Sonar preprocessing module, used to collect raw sonar echo signals and perform preprocessing;

[0084] The spectrum construction module is used to convert the pre-processed sonar echo signal into multiple types of spectrum images, and perform weighted fusion on them in combination with the physical parameters of sonar propagation to generate multi-channel spectrum images;

[0085] A spectrum enhancement module is used to perform spectrum enhancement processing on the multi-channel spectrum image, wherein the spectrum enhancement includes logarithmic spectrum compression, pseudo color mapping and local contrast enhancement;

[0086] The residual feature extraction module is used to input the enhanced spectrum image into the SonarResBlock model to extract the sonar target feature representation;

[0087] The label filtering module is used to calculate the label credibility score based on the sonar target feature representation and construct a label training set for label samples with scores greater than the label filtering threshold;

[0088] The joint training module is used to merge the labeled training set with the real annotated samples to form a joint training set and perform iterative optimization of the SonarResBlock model;

[0089] The recognition output module is used to perform inference recognition using the updated model and output the category label of the sonar target.

[0090] The present invention constructs a complete sonar target recognition system to achieve a closed loop of the entire process from raw sonar signal acquisition, spectrum image construction, image enhancement, feature extraction, trusted label screening to model joint training and target recognition output. The system structurally integrates multi-module processing links and combines the residual network structure with sonar propagation characteristics to significantly improve recognition accuracy. It has good adaptability, especially for underwater weak echoes or multipath interference scenarios, and has strong robustness and scalability, solving the problem of insufficient generalization ability of existing sonar recognition systems.

[0091] In this embodiment, the modules are connected through the following methods:

[0092] Collecting original sonar echo signals and preprocessing the original sonar echo signals to obtain preprocessed sonar echo signals;

[0093] The spectral features of the pre-processed sonar echo signal are constructed to generate time-frequency diagrams, amplitude-phase diagrams, and envelope spectrograms. Combined with the physical parameters of sonar propagation, the three types of spectrum diagrams are fused according to channel weighting to construct a multi-channel spectrum image.

[0094] Performing spectrum enhancement processing on the multi-channel spectrum image to obtain an enhanced spectrum image;

[0095] The enhanced spectrum image is input into the SonarResBlock model to extract the key target features in the enhanced spectrum image and obtain the sonar target feature representation;

[0096] Based on the sonar target feature representation, a classification head is used for inference. The label confidence score is calculated based on the category probability distribution, the category KL divergence of the prediction results, and the category frequency offset. A label training set is constructed for label samples with label confidence scores greater than the label screening threshold. This is then merged with the real labeled samples to form a joint training set, and the SonarResBlock model is iteratively optimized.

[0097] The updated SonarResBlock model is used to perform recognition reasoning on the enhanced spectrum image, and the classification head is used to reason about the sonar target feature representation and output the target category label.

[0098] The present invention effectively ensures the process integrity and information consistency of sonar target recognition by sequentially connecting various functional modules in specific processing steps. Through the system process, it ensures the consistent transmission of various variables and data closure in the processing chain from raw sonar signals to category labels. In particular, through the introduction of a joint training optimization strategy, it enhances the learning ability of the model in weakly supervised scenarios while ensuring recognition accuracy, improves label utilization and the diversity of training samples, and enhances the generalization performance of the model.

[0099] In this embodiment, the preprocessing includes noise suppression, amplitude normalization and time window truncation.

[0100] The present invention introduces noise suppression, amplitude normalization and time window truncation technical means in the sonar preprocessing step, effectively eliminating the interference of non-structural noise in the original echo signal on spectrum construction, ensuring the temporal consistency of subsequent spectrum graphs through time window truncation, and improving the model's adaptability to signals of different intensities through normalization processing, thereby enhancing the stability and accuracy of subsequent spectrum image construction and recognition modules, providing a high-quality, structured input signal foundation for the entire recognition process.

[0101] In this embodiment, the sonar propagation physical parameters include propagation delay, frequency attenuation factor and echo energy attenuation.

[0102] The present invention proposes a multi-channel spectrum image construction method that combines the physical parameters of sonar propagation. It integrates three types of features: time-frequency diagram, envelope spectrum diagram and amplitude-phase diagram to form input data with enhanced physical information. By normalizing physical quantities such as propagation delay, frequency attenuation factor and echo energy attenuation and using them for channel weight adjustment, it achieves a fine fusion of the three types of spectrum diagrams, enhances the input image's ability to depict the sonar propagation mechanism, and improves the robustness and discriminability of the subsequent feature extraction module for complex underwater acoustic environments.

[0103] In this embodiment, the construction of a multi-channel spectrum image includes the following specific steps:

[0104] Perform short-time Fourier transform on the preprocessed sonar echo signal to generate the corresponding time-frequency diagram;

[0105] Perform Hilbert-Huang transform on the pre-processed sonar echo signal to construct the envelope spectrum;

[0106] The frequency domain complex spectrum is obtained by fast Fourier transform on the preprocessed sonar echo signal, and the corresponding amplitude spectrum and phase spectrum are extracted based on the frequency domain complex spectrum. The amplitude spectrum and phase spectrum are then channel-joined to construct the amplitude-phase diagram.

[0107] Normalize the sonar propagation physical parameters to unitless values. For the time-frequency diagram, envelope spectrum, and amplitude-phase diagram, adjust the channel fusion weights based on the normalized sonar propagation physical parameters:

[0108] ;

[0109] in, represents the weighting coefficient of the time-frequency graph channel, represents the weighting coefficient of the envelope spectrum channel, represents the weighting coefficient of the amplitude-phase image channel, Indicates the maximum attenuation reference value, represents the frequency attenuation factor, Indicates target distance and frequency The echo energy attenuation under Indicates frequency, represents the propagation delay;

[0110] The three types of spectrum images are weighted and combined according to the channel dimension to form a multi-channel spectrum image:

[0111] ;

[0112] in, Indicates the Multi-channel spectrum image of the frame, Represents the normalization operation, represents the channel-level concatenation operation, represents the time-frequency diagram, represents the envelope spectrum, Represents the amplitude phase diagram.

[0113] This paper improves the discriminability of multi-channel spectrum images in both the spatial and visual domains through a spectrum enhancement module. The introduction of logarithmic spectrum compression and pseudo-color mapping effectively enhances the contrast of detailed areas in the image. Furthermore, the local contrast enhancement function optimizes image texture features, making it easier for the model to perceive the boundary differences between the target and the background. The enhanced spectrum image not only improves the gradient response quality during model training but also significantly improves the ability to recognize small and weak targets during feature extraction.

[0114] In this embodiment, obtaining the enhanced spectrum image includes the following specific steps:

[0115] Perform channel-independent logarithmic spectrum compression on the multi-channel spectrum image to obtain a logarithmic compressed image:

[0116] ;

[0117] in, Indicates the Frame in channel , pixel position The logarithmic compression result on , Represents the multi-channel spectrum image Frame in channel , pixel position Pixel intensity on ;

[0118] Performing a pseudo color mapping operation on the logarithmic compressed image to map the grayscale value into a pseudo color image in the RGB color space to obtain a pseudo color spectrum image, wherein the pseudo color spectrum image is an RGB image with three channels;

[0119] Perform local contrast enhancement on the pseudo-color spectrum image using a weighted local statistical contrast enhancement function:

[0120] ;

[0121] in, Indicates the Frame at pixel location The contrast enhancement value on Represents the pseudo-color spectrum image Frame at pixel location The RGB value on Indicates pixel position The local neighborhood mean centered at represents the corresponding local neighborhood standard deviation, It represents the stability constant introduced to prevent the denominator from being zero;

[0122] Normalization is performed on the image after local contrast enhancement to obtain an enhanced spectrum image.

[0123] The SonarResBlock model designed in this paper integrates the frequency-time bidirectional perception structure and the sonar propagation perception attention mechanism. It uses the directional design of the convolution kernel size to extract frequency domain and time domain features, and dynamically adjusts the channel response through the propagation physical quantity to form a deep feature representation that is closely coupled with physical laws. In addition, the structure uses identity mapping to improve the efficiency of gradient propagation, maintain a stable flow of information, enhance the model's nonlinear expression ability and training stability, and has excellent target representation capabilities in complex underwater acoustic scenes.

[0124] In this embodiment, obtaining the sonar target feature representation includes the following specific steps:

[0125] The enhanced spectrum image is input into the SonarResBlock model for feature extraction, wherein the SonarResBlock model is composed of a sonar-adapted residual module;

[0126] The sonar-adaptive residual module includes a main branch and an identity mapping branch, wherein the main branch sequentially includes a first convolutional layer, a batch normalization layer, a ReLU activation function, a second convolutional layer, and a sonar propagation perception attention layer, and the identity mapping branch includes an identity connection or Convolutional connections:

[0127] When the sizes of the main branch output and the enhanced spectrum image are inconsistent, the identity mapping branch selects Convolutional connections;

[0128] When the size of the main branch output and the enhanced spectrum image are consistent, the identity mapping branch selects the identity connection;

[0129] The convolution structure of the main branch adopts a frequency-time bidirectional perception strategy, and the convolution kernel size of the first convolution layer is set to , used to extract contextual information in the frequency dimension, the convolution kernel size of the second convolution layer is set to , used to extract the local continuous change features of the time dimension. The frequency dimension and time dimension correspond to the height and width directions of the enhanced spectrum image respectively;

[0130] Output the feature map after the second convolution:

[0131] ;

[0132] ;

[0133] in, Represents the feature map after the first convolution process, represents the feature map after the second convolution, 、 Represent the first and second convolution operations respectively, represents the batch normalization function, represents the rectified linear unit activation function;

[0134] The feature map after the second convolution is input to the sonar propagation perception attention layer for channel response regulation. The sonar propagation perception attention layer uses the propagation delay, frequency attenuation factor and echo energy attenuation to construct the channel attention weight:

[0135] ;

[0136] in, Indicates the The attention weight of the channel, represents the Sigmoid activation function, represents the frequency attenuation factor, represents the propagation delay, Indicates target distance and frequency The echo energy attenuation under Indicates the maximum energy attenuation reference value, 、 、 Represents the weighting coefficient of each physical quantity;

[0137] The propagation delay, frequency attenuation factor and echo energy attenuation adopt normalized unitless values;

[0138] Apply the attention weight to the feature map after the second convolution to obtain the attention-weighted feature map:

[0139] ;

[0140] in, Indicates the Channel, pixel position The attention-weighted feature value on , Indicates the Channel, pixel position The eigenvalue after the second convolution;

[0141] The attention-weighted feature map is element-wise added to the output of the identity mapping branch to obtain the output feature map of the sonar-adaptive residual module, which serves as the input of the next round of sonar-adaptive residual module and finally outputs the sonar target feature representation.

[0142] The present invention realizes deep feature extraction of enhanced spectrum images by constructing a SonarResBlock model composed of sonar-adaptive residual modules. Structurally, a main branch and an identity mapping branch are designed in parallel, and a channel attention mechanism is constructed in combination with the physical parameters of sonar propagation. This enhances the robustness of feature representation to multi-scale variations of sonar signals. The convolution kernel size is set according to the frequency-time bidirectional perception strategy, which effectively improves the model's ability to capture local spectral structures and contextual associations, thereby improving the model's recognition accuracy for weak targets and complex backgrounds.

[0143] In this embodiment, the iterative optimization of the SonarResBlock model includes the following specific steps:

[0144] Inputting the sonar target feature representation into a classification head to perform label prediction and output a class probability distribution, the classification head using a multi-layer perceptron;

[0145] Based on constructing a historical prior distribution and calculating the KL divergence between the category probability distribution and the historical prior distribution, the historical prior distribution represents the normalized frequency distribution of samples encountered in past training rounds in each category:

[0146] ;

[0147] in, Indicates the The KL divergence between the predicted distribution of a sample and the historical prior distribution, Indicates the The frequency ratio of the class in the historical prior distribution, represents the stability constant introduced to prevent the denominator from being zero, Indicates the The samples belong to The class probability of the class, Indicates the total number of categories of sonar targets;

[0148] Construct a class frequency offset to measure the difference between the label class and the ideal balanced distribution:

[0149] ;

[0150] in, Indicates the The category frequency offset of the sample prediction category, Indicates the main category index of the current prediction, Indicates the frequency ratio of the predicted main category in the historical prior distribution, Indicates the total number of categories;

[0151] The label credibility score is constructed by integrating category probability distribution, distribution deviation and category imbalance factors:

[0152] ;

[0153] in, Indicates the The label confidence score of each sample, represents the KL divergence penalty factor, represents the class frequency offset adjustment factor, Indicates the The class probability of the samples, represents the natural exponential function;

[0154] Set the label filtering threshold to filter label samples in descending order as the number of parameter update rounds decreases:

[0155] ;

[0156] in, Indicates the The label filtering threshold for round parameter update, represents the initial label screening threshold, Indicates the total number of rounds, Represents the exponential factor that controls the rate at which the threshold drops;

[0157] Filter all label samples that meet the label credibility score greater than the label screening threshold to build a label training set, and merge it with the real labeled samples to form a joint training set;

[0158] The parameters of the SonarResBlock model are updated based on the joint training set. The parameter update is performed by minimizing the cross entropy loss function of the SonarResBlock model and performing gradient update using the back propagation algorithm.

[0159] The present invention introduces a multi-factor label credibility scoring mechanism. After label prediction, it comprehensively considers the category probability distribution, KL divergence and category frequency offset to construct a credibility scoring index for dynamically screening high-quality label samples. Combined with a label screening threshold function that decreases with the number of parameter update rounds, it avoids low-confidence labels from interfering with the initial training, gradually improves the quality of label samples and enhances model stability. In the joint training stage, the SonarResBlock model parameters are optimized by minimizing the cross-entropy loss function, achieving unified modeling and efficient training of real labels and screened labels, thereby improving the accuracy and robustness of sonar target recognition.

[0160] Example 1

[0161] In order to verify the feasibility of the present invention in implementation, the present invention is applied to underwater target recognition operations in a certain near-shore waters. The underwater environment in this sea area is complex, and sonar signals are easily interfered with by current disturbances, terrain reflections and background noise. Conventional sonar target recognition methods in such environments generally have problems such as insufficient target feature extraction, insufficient spectrum image expression and high label uncertainty, resulting in unstable recognition accuracy. The present invention proposes a sonar target recognition system based on a residual network, which combines spectrum graph fusion enhancement, residual feature extraction and label credibility screening mechanism to effectively address the above challenges.

[0162] During the actual deployment process, the sonar echo signals of underwater targets during continuous navigation are first collected. After preprocessing operations such as noise suppression, amplitude normalization and time window interception, three types of spectrum images are constructed, namely, time-frequency diagram, amplitude-phase diagram and envelope spectrum diagram. For the three types of images, the system introduces sonar physical property parameters such as propagation delay, frequency attenuation factor and energy attenuation as the basis for channel weighting, and fuses them to generate multi-channel spectrum images. In order to further improve image recognition, the multi-channel spectrum image is processed through logarithmic spectrum compression, pseudo-color mapping and local contrast enhancement to generate an enhanced spectrum image, thereby improving the separability of weak targets and fuzzy boundaries.

[0163] The enhanced spectrum image is input into the constructed SonarResBlock model for feature extraction. The network utilizes a parallel design of the main branch and the identity mapping branch. The main branch incorporates a frequency-time bidirectional perception strategy and a sonar propagation-aware attention mechanism, enhancing both the frequency dimension's contextual capture and the modeling of local continuous changes in time. By element-by-element addition of the identity mapping branch's output, the network achieves deep feature fusion that is more authentic to sonar physics, outputting a sonar target feature representation.

[0164] During the feature inference phase, the system uses a multi-layer perceptron classification head to perform probabilistic inference on the target feature representation. The system combines the KL divergence between the historical prior distribution and the predicted category distribution, while also constructing a category frequency offset to assess the label credibility of each sample. During continuous training, the screening threshold is dynamically adjusted based on a round-by-round reduction strategy, retaining only high-confidence samples to construct a pseudo-label training set. This is then combined with existing true-label samples to form a joint training set, enabling more robust iterative optimization of the model.

[0165] In order to verify the performance of the present invention in practice, a comparison was made with the traditional method.

[0166] Table 1 Comparison of specific performance indicators of sonar target recognition system and traditional methods

[0167]

[0168] It can be clearly seen from the detailed data in Table 1 that the method described in the present invention is superior to the traditional method in many key indicators.

[0169] From the most critical target recognition accuracy point of view, the accuracy of the method of the present invention reached 92.8%, which is 9.3 percentage points higher than the 83.5% of the traditional method. This shows that the method of the present invention has better recognition performance, especially suitable for underwater detection tasks with strict precision requirements. Further refinement to the recognition ability of low signal-to-noise ratio targets in actual scenarios, the method of the present invention achieved an accuracy of 88.7%, while the traditional method was only 62.4%. The significant improvement of this indicator shows that the present invention has improved the system's sensitivity and discrimination ability for weak echo signals in complex marine environments through the special structure of the residual network and spectrum enhancement technology.

[0170] Secondly, in terms of adaptability to category imbalance, the method of the present invention also performs outstandingly, with an adaptability of up to 87.2%, while the traditional method is only 65.3%, an improvement of 21.9 percentage points. The problem of category imbalance is widespread in actual underwater acoustic environments. Many target categories are scarce, which makes traditional methods easily overlooked or misjudged. The dynamic label credibility screening mechanism and joint training strategy adopted by the present invention effectively improve the recognition reliability of rare categories.

[0171] At the same time, the method of the present invention demonstrates significant advantages in terms of model training efficiency and stability. Comparing the number of training convergence rounds, the traditional method requires 120 rounds to reach a stable convergence state, while the method of the present invention requires only 65 rounds, indicating that the method of the present invention has almost doubled the training efficiency. Furthermore, the variance of the parameter update stability index during the training phase of the present invention is only 0.012, far lower than the 0.045 of the traditional method. This means that the training process of the method of the present invention is more stable, the impact of noise or abnormal data is significantly reduced, and the model parameter updates are more stable and reliable.

[0172] In terms of the reliability indicators of actual applications - false alarm rate and missed alarm rate, the false alarm rate of the method of the present invention is only 4.2%, and the missed alarm rate is only 3.1%. The traditional methods are as high as 15.7% and 12.3% respectively. Both indicators are greatly reduced, which reflects that the false positive rate of the method of the present invention in actual applications is significantly reduced, which has important practical significance for improving the overall system reliability and reducing the workload of subsequent manual review.

[0173] The recognition accuracy of the method of the present invention for small sample categories reaches 84.5%, far exceeding the 58.9% of the traditional method. This shows that the method of the present invention can still maintain efficient and stable recognition capabilities even in scenarios with limited training data and insufficient number of labeled samples. This improvement in small sample recognition capabilities enhances the generalization ability and actual deployment capability of the method of the present invention in real environments.

[0174] The above description is only a preferred specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any technician familiar with the technical field, within the technical scope disclosed by the present invention, who makes equivalent replacements or changes based on the technical solution and inventive concept of the present invention, should be covered by the scope of protection of the present invention.

Claims

1. A sonar target recognition system based on residual network, characterized in that: include: Sonar preprocessing module, used to collect raw sonar echo signals and perform preprocessing; The spectrum construction module is used to convert the pre-processed sonar echo signal into multiple types of spectrum images, and perform weighted fusion on them in combination with the physical parameters of sonar propagation to generate multi-channel spectrum images; A spectrum enhancement module is used to perform spectrum enhancement processing on the multi-channel spectrum image, wherein the spectrum enhancement includes logarithmic spectrum compression, pseudo color mapping and local contrast enhancement; The residual feature extraction module is used to input the enhanced spectrum image into the SonarResBlock model to extract the sonar target feature representation; The label filtering module is used to calculate the label credibility score based on the sonar target feature representation and construct a label training set for label samples with scores greater than the label filtering threshold; The joint training module is used to merge the labeled training set with the real annotated samples to form a joint training set and perform iterative optimization of the SonarResBlock model; The recognition output module is used to perform inference recognition using the updated model and output the category label of the sonar target.

2. The sonar target recognition system based on residual network according to claim 1, characterized in that: The modules are implemented as follows: Collecting original sonar echo signals and preprocessing the original sonar echo signals to obtain preprocessed sonar echo signals; The spectral features of the pre-processed sonar echo signal are constructed to generate time-frequency diagrams, amplitude-phase diagrams, and envelope spectrograms. Combined with the physical parameters of sonar propagation, the three types of spectrum diagrams are fused according to channel weighting to construct a multi-channel spectrum image. Performing spectrum enhancement processing on the multi-channel spectrum image to obtain an enhanced spectrum image; The enhanced spectrum image is input into the SonarResBlock model to extract the key target features in the enhanced spectrum image and obtain the sonar target feature representation; Based on the sonar target feature representation, a classification head is used for inference. The label confidence score is calculated based on the category probability distribution, the category KL divergence of the prediction results, and the category frequency offset. A label training set is constructed for label samples with label confidence scores greater than the label screening threshold. This is then merged with the real labeled samples to form a joint training set, and the SonarResBlock model is iteratively optimized. The updated SonarResBlock model is used to perform recognition reasoning on the enhanced spectrum image, and the classification head is used to reason about the sonar target feature representation and output the target category label.

3. The sonar target recognition system based on residual network according to claim 2, characterized in that: The preprocessing includes noise suppression, amplitude normalization and time window truncation.

4. The sonar target recognition system based on residual network according to claim 2, characterized in that: The sonar propagation physical parameters include propagation delay, frequency attenuation factor and echo energy attenuation.

5. The sonar target recognition system based on residual network according to claim 2, characterized in that: The construction of a multi-channel spectrum image comprises the following specific steps: Perform short-time Fourier transform on the preprocessed sonar echo signal to generate the corresponding time-frequency diagram; Perform Hilbert-Huang transform on the pre-processed sonar echo signal to construct the envelope spectrum; The frequency domain complex spectrum is obtained by fast Fourier transform on the preprocessed sonar echo signal, and the corresponding amplitude spectrum and phase spectrum are extracted based on the frequency domain complex spectrum. The amplitude spectrum and phase spectrum are then channel-joined to construct the amplitude-phase diagram. Normalize the sonar propagation physical parameters into unitless values, and adjust the channel fusion weights based on the normalized sonar propagation physical parameters for the time-frequency diagram, envelope spectrum diagram, and amplitude-phase diagram; The three types of spectrum images are weighted and combined according to the channel dimension to form a multi-channel spectrum image.

6. The sonar target recognition system based on residual network according to claim 2, characterized in that: Obtaining the enhanced spectrum image comprises the following specific steps: Perform channel-independent logarithmic spectrum compression on the multi-channel spectrum image to obtain a logarithmic compressed image; Performing a pseudo color mapping operation on the logarithmic compressed image to map the grayscale value into a pseudo color image in the RGB color space to obtain a pseudo color spectrum image, wherein the pseudo color spectrum image is an RGB image with three channels; Perform local contrast enhancement on pseudo-color spectrum images using weighted local statistical contrast enhancement function; Normalization is performed on the image after local contrast enhancement to obtain an enhanced spectrum image.

7. The sonar target recognition system based on residual network according to claim 2, characterized in that: The obtaining of the sonar target feature representation comprises the following specific steps: The enhanced spectrum image is input into the SonarResBlock model for feature extraction, wherein the SonarResBlock model is composed of a sonar-adapted residual module; The sonar-adaptive residual module includes a main branch and an identity mapping branch, wherein the main branch sequentially includes a first convolutional layer, a batch normalization layer, a ReLU activation function, a second convolutional layer, and a sonar propagation perception attention layer, and the identity mapping branch includes an identity connection or Convolutional connections: When the sizes of the main branch output and the enhanced spectrum image are inconsistent, the identity mapping branch selects Convolutional connections; When the size of the main branch output and the enhanced spectrum image are consistent, the identity mapping branch selects the identity connection; The convolution structure of the main branch adopts a frequency-time bidirectional perception strategy, and the convolution kernel size of the first convolution layer is set to , used to extract contextual information in the frequency dimension, the convolution kernel size of the second convolution layer is set to , used to extract the local continuous change features of the time dimension. The frequency dimension and time dimension correspond to the height and width directions of the enhanced spectrum image respectively; Output the feature map after the second convolution; The feature map after the second convolution is input to the sonar propagation perception attention layer for channel response regulation. The sonar propagation perception attention layer uses the propagation delay, frequency attenuation factor and echo energy attenuation to construct the channel attention weight; The propagation delay, frequency attenuation factor and echo energy attenuation adopt normalized unitless values; Apply the attention weight to the feature map after the second convolution to obtain the attention-weighted feature map; The attention-weighted feature map is element-wise added to the output of the identity mapping branch to obtain the output feature map of the sonar-adaptive residual module, which serves as the input of the next round of sonar-adaptive residual module and finally outputs the sonar target feature representation.

8. The sonar target recognition system based on residual network according to claim 2, characterized in that: The iterative optimization of the SonarResBlock model includes the following specific steps: Inputting the sonar target feature representation into a classification head to perform label prediction and output a class probability distribution, the classification head using a multi-layer perceptron; Based on constructing a historical prior distribution and calculating the KL divergence between the class probability distribution and the historical prior distribution, the historical prior distribution represents the normalized frequency distribution of samples encountered in past training rounds in each class; Construct a class frequency offset to measure the difference between the label class and the ideal balanced distribution; The label credibility score is constructed by integrating category probability distribution, distribution deviation and category imbalance factors: ; in, Indicates the The label confidence score of each sample, represents the KL divergence penalty factor, represents the class frequency offset adjustment factor, Indicates the The class probability of the samples, represents the natural exponential function; Set the label screening threshold to filter label samples in descending order with the number of parameter update rounds; Filter all label samples that meet the label credibility score greater than the label screening threshold to build a label training set, and merge it with the real labeled samples to form a joint training set; The parameters of the SonarResBlock model are updated based on the joint training set. The parameter update is performed by minimizing the cross entropy loss function of the SonarResBlock model and performing gradient update using the back propagation algorithm.

Citation Information

Patent Citations

  • Underwater acoustic target recognition method based on adversarial residual network

    CN113435276A

  • Underwater target identification method based on OfficientNet

    CN115204214A