Underwater sound target identification method based on BSCQT and DCAM

Through the combination of BSCQT and DCAM, high-precision water acoustic target recognition in complex underwater environments in the Yellow River is achieved, solving the problems of low identification accuracy and calculation efficiency in the prior art, and is suitable for underwater equipment with limited resources.

CN120263306APending Publication Date: 2025-07-04HENAN UNIVERSITY

Patent Information

Application Number
CN202510318613.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-18
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The existing water acoustic target recognition method has low recognition accuracy and calculation efficiency in complex underwater environments of the Yellow River, making it difficult to effectively suppress background noise and capture time-frequency dynamic characteristics. The complex network structure leads to large computing overhead and is difficult to deploy in resource-constrained devices.

Method used

The water acoustic target recognition method based on BSCQT and DCAM is adopted, and adaptive parameter filtering and weighting is performed through band-specific constant Q transformation, combined with a dynamic context-aware mask network, a lightweight DCAM dense TDNN network is built to realize multi-band feature extraction and noise suppression, and optimize the feature learning and recognition process.

Benefits of technology

It significantly improves the feature extraction accuracy and noise suppression effect of water acoustic signals, improves recognition accuracy, and reduces calculation complexity, making it suitable for efficient operation in resource-constrained underwater equipment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120263306A_ABST
    Figure CN120263306A_ABST
Patent Text Reader

Abstract

The invention provides an underwater acoustic target identification method based on BSCQT and DCAM, and the method comprises the steps: carrying out the label marking of a collected underwater acoustic signal, carrying out the non-overlapping sampling to generate a sample, and randomly dividing the signal into a training set and a test set; dividing the underwater acoustic signals in the training set and the test set into three sub-bands, performing adaptive parameter filtering and weighting on each sub-band by adopting frequency band specific constant Q transform, and connecting multi-band characteristics in series to obtain a BSCQT spectrogram; constructing a DCAM dense TDNN network based on a DCAM mechanism, and constructing a dynamic context awareness mask network; inputting the BSCQT spectrogram in the training set into a dynamic context awareness mask network for training, and storing model parameters to obtain an underwater acoustic target recognition model; and inputting the BSCQT spectrogram of the underwater acoustic signals of the test set into an underwater acoustic target recognition model to obtain an underwater acoustic target recognition result. According to the method, the feature extraction precision and the noise suppression effect of the underwater acoustic signal are remarkably improved, the feature distinction degree is enhanced, the recognition precision is improved, and redundant calculation is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of underwater safety monitoring in the Yellow River Basin, especially to the field of underwater acoustic target recognition for underwater safety monitoring in the Yellow River, and particularly to an underwater acoustic target recognition method integrating frequency band transformation and deep learning. Background Art

[0002] Underwater acoustic target recognition is a research direction of great strategic significance in underwater acoustic signal processing, and has important application value in fields such as dam defect detection, underwater monitoring, and ecological protection in the Yellow River Basin. The core goal of underwater acoustic target recognition is to use sensors to collect data and determine the type of underwater acoustic target through acoustic feature extraction and pattern recognition methods. With the rapid development of unmanned underwater platform technology, underwater acoustic target recognition methods are being more and more widely applied in autonomous underwater vehicles, remote sensor networks, and other intelligent underwater devices. However, to fully unleash the potential of these applications, there is an urgent need to improve the recognition accuracy and computational efficiency of the recognition algorithm.

[0003] There are two main challenges in the existing technologies: First, due to the complex underwater environment in the Yellow River introducing a large amount of background noise, and the acoustic signals having significant dynamic characteristics in different frequency bands, noise suppression and feature extraction are extremely challenging; Second, to obtain high recognition accuracy, it is often necessary to construct a complex network structure, resulting in a large number of network parameters and high computational overhead, making it difficult to deploy in resource-constrained practical systems. With the breakthrough progress made by deep learning algorithms in theory, new method innovations have emerged in recognition technologies.

[0004] The invention patent with the application number 202311244959.5 discloses an underwater acoustic signal recognition method based on a multi-branch backbone external attention network. First, the input data is passed through multiple parallel backbone network branches to extract feature information of different levels of the underwater acoustic signal; Second, channel and spatial attention modules are supplemented to weight the channel and spatial dimensions of the underwater acoustic signal respectively, adjusting the importance of different channels and spatial positions for feature representation; Finally, an external attention module is integrated to guide the feature extraction and prediction of the network with an external memory unit and additional calculations, thus significantly improving the recognition rate and robustness of the model. A large number of experiments have confirmed the superiority of the above-mentioned invention algorithm over the existing methods. However, the above-mentioned invention has obvious defects in recognition accuracy and computational efficiency: The multi-branch structure leads to feature redundancy and the traditional attention mechanism is difficult to capture the time-frequency dynamic characteristics; At the same time, the complex network structure results in large computational overhead and insufficient real-time performance, and it is difficult to deploy the external attention module in practical scenarios, unable to meet the requirements of underwater monitoring in the Yellow River. Summary of the Invention

[0005] Aiming at the technical problems of low recognition accuracy and low computational efficiency of existing underwater acoustic target recognition methods, the present invention proposes an underwater acoustic target recognition method based on BSCQT (Band-Specific Constant-Q Transform) and DCAM (Dynamic Context-Aware Mask). Through the multi-band adaptive weighting and feature concatenation of BSCQT, the feature extraction accuracy of underwater acoustic signals and the noise suppression effect are significantly improved, especially in the complex underwater environment of the Yellow River; at the same time, the DCAM mechanism enhances the feature discrimination by dynamically adjusting the receptive field and fusing multi-scale context information, further improving the recognition accuracy; in terms of computational efficiency, the lightweight design of DCAM and the frequency-adaptive pooling granularity effectively reduce the computational complexity, and DCAMNet optimizes feature learning through cascaded DCAM dense TDNN blocks, reducing redundant calculations and ensuring efficient operation in resource-constrained devices.

[0006] To achieve the above object, the technical solution of the present invention is implemented as follows: An underwater acoustic target recognition method based on BSCQT and DCAM, the steps are as follows:

[0007] Step 1: Label the collected underwater acoustic signals, perform non-overlapping sampling on the underwater acoustic signals of each category to generate samples, and randomly divide all samples into a training set and a test set;

[0008] Step 2: Divide the underwater acoustic signals in the training set and the test set into three sub-bands: low frequency, medium frequency, and high frequency. Use the band-specific constant-Q transform to perform adaptive parameter filtering and weighted extraction of multi-band features on each sub-band, and concatenate the multi-band features to obtain the BSCQT spectrogram of the underwater acoustic signal samples;

[0009] Step 3: Construct a DCAM dense TDNN network based on the DCAM mechanism, and use multiple DCAM dense TDNN networks to construct a dynamic context-aware mask network for deep learning; input the BSCQT spectrogram of the underwater acoustic signals in the training set into the dynamic context-aware mask network for training, save the trained model parameters, and obtain an underwater acoustic target recognition model;

[0010] Step 4: Input the BSCQT spectrogram of the underwater acoustic signals in the test set into the trained underwater acoustic target recognition model to obtain the underwater acoustic target recognition result.

[0011] Preferably, the method of performing non-overlapping sampling on the underwater acoustic signals of each category to generate samples is: continuously and non-overlappingly extract segments from continuous audio signals at a fixed duration to generate a series of independent left-position samples of underwater acoustic signals;

[0012] The frequency range of the low frequency of the sub-band is 30 - 1000 Hz, the frequency range of the medium frequency of the sub-band is 1000 - 2000 Hz, and the frequency range of the high frequency of the sub-band is 2000 - 5000 Hz.

[0013] Preferably, the method for performing adaptive parameter filtering on each sub-band by the band-specific constant Q transform is as follows: Determine the number of filters N of the constant Q transform filter in sub-band i bins,i ; Calculate the center frequency f of the k-th filter k ; According to the number of filters N bins,i and the center frequency f k , perform constant Q filtering on sub-band i using a window function to obtain CQT coefficients;

[0014] The implementation method of the weighting is as follows: Perform sub-band weighting on the CQT coefficients output by the constant Q filtering. The sub-band spectrograms of the weighted low-frequency, intermediate-frequency, and high-frequency are respectively: X low = W l × CQT low , X mid = W m × CQT mid , X high = W h × CQT high ; Among them, CQT low , CQT mid and CQT high respectively represent the CQT coefficients of the low-frequency, intermediate-frequency, and high-frequency sub-bands, and W l , W m and W h are learnable weight coefficients;

[0015] Obtain a multi-band feature representation by concatenating the sub-band spectrograms to obtain the BSCQT spectrogram X BSCQT = concat(X low , X mid , X high ); Among them, concat represents the spectrogram concatenation operation in the frequency dimension.

[0016] Preferably, the number of filters where, f max,i and f min,i respectively represent the maximum frequency and the minimum frequency of sub-band i, and n octave,i is the number of bins per octave of sub-band i, and bins represents the number of constant Q transform filters in sub-band i;

[0017] Of the k-th filter where, b represents the filter distribution density, k = 0, 1,..., K - 1, and K represents the number of filters in sub-band k, which is determined by the frequency range:

[0018] The calculation method of the CQT coefficients is as follows: Among them, CQT i (c,n) represents the CQT coefficient at the c-th filter and time n in sub-band i; x i (m) is the m-th signal value of the signal in sub-band i, w c,i (m) represents the window function of the c-th filter in sub-band i; H c,i represents the window length of the c-th filter in sub-band i, which is inversely proportional to the center frequency f k ; Q i is the quality factor of sub-band i, j is the imaginary unit, c = 0, 1, …, N bins,i -1, N bins,i is the total number of filters in sub-band i.

[0019] Preferably, the dynamic context-aware network includes a lightweight front-end convolutional network, multiple DCAM dense TDNN networks, and a classifier connected in sequence.

[0020] Preferably, the multiple DCAM dense TDNN networks are 3 cascaded DCAM dense TDNN networks for depth information compression;

[0021] The lightweight front-end convolutional network includes an initial 3×3 convolutional kernel and two lightweight residual blocks connected in sequence; the lightweight residual block adopts a residual structure design: the main branch convolution of the 3×3 convolutional kernel compresses local information, and the shortcut branch convolution of the 1×1 convolutional kernel transmits information;

[0022] Each DCAM dense TDNN network includes multiple dynamic context-aware mask mechanisms, a dense connection layer, and a 1×1 convolutional feature transformation layer connected in sequence;

[0023] The classifier includes a global average pooling layer, a fully connected layer, and a Softmax classification operation connected in sequence.

[0024] Preferably, the processing method of the dynamic context-aware network is as follows: Input the BSCQT spectrogram into the lightweight front-end convolutional network. First, perform local feature extraction through an initial 3×3 convolutional kernel to generate a preliminary feature map. Subsequently, the preliminary feature map is normalized through a batch normalization layer to accelerate training convergence and improve the generalization ability of the model. The normalized feature map is then non-linearly transformed through the H-Swish activation function to obtain an activated feature map. The input BSCQT spectrogram also undergoes information transmission through a 1×1 convolutional kernel to generate a feature map for the shortcut branch. The activated feature map obtained from the main branch is element-wise added to the feature map of the shortcut branch to complete the residual processing of local information compression and information transmission in the lightweight residual block, obtaining feature map I. Feature map I is input into the dynamic context-aware mask mechanism for processing to obtain an enhanced feature map. The enhanced feature map is processed through a dense connection layer to enhance the feature expression ability through feature reuse and gradient propagation. Then, it undergoes feature transformation through a 1×1 convolutional feature transformation layer to reduce the feature channel dimension and computational complexity, finally obtaining the output feature map F1. The output feature map F1 is successively processed through the second DCAM dense TDNN network and the third DCAM dense TDNN network. Each network repeats the processing flow of the dynamic context-aware mask mechanism, dense connection layer, and feature transformation layer, finally obtaining the output feature map F3. The output feature map F3 is input into the classifier, and through global average pooling operation, the time-domain features are compressed into a vector representation of a fixed dimension. The features are mapped to the target category dimension space through a fully connected layer. The output is converted into a category probability distribution through the Softmax function to obtain the classification result y.

[0025] Preferably, the calculation method of the lightweight residual block is as follows:

[0026] F LFCN = Hswish(BN(Conv 3×3 (X)))+Conv 1×1 (X)

[0027] where BN is the batch normalization layer and Hswish is the lightweight activation function. X represents the feature map input to the lightweight residual block. Conv 3×3 and Conv 1×1 represent the operations of the 3×3 convolutional kernel and the 1×1 convolutional kernel respectively.

[0028] The calculation method for each DCAM dense TDNN network to obtain the output feature map is as follows:

[0029] F = Transit(Dense(DCAM(Pool 5×1 (I)))), where I represents the feature map input to the DCAM dense TDNN network, Pool5×1 Represents a 5×1 time-domain downsampling operation; DCAM is a dynamic context-aware masking mechanism that enhances the spatial features of the input feature map; Dense is a dense connection layer for feature reuse and gradient propagation; Transit is a 1×1 convolutional feature transformation layer for reducing the feature channel dimension and computational complexity;

[0030] The classifier performs time-domain information aggregation on the feature map output by the last DCAM dense TDNN network and maps it to the class space through a fully connected layer, obtaining the classification result: y = Softmax(FC(GlobalPool(F3))); where GlobalPool is a global average pooling operation, FC is a fully connected layer, Softmax is used to generate the class probability distribution, and F3 is the output feature map of the last DCAM dense TDNN network.

[0031] Preferably, the method for the dynamic context-aware masking mechanism to enhance the spatial features of the input feature map is as follows: adaptively adjust the receptive field size through frequency analysis to obtain the dynamic pooling segmentation number; use the global average pooling operation to process the input feature map to obtain the global pooling result, use the dynamic pooling segmentation number to perform local pooling operation on the input feature map to obtain the local pooling result, combine the global pooling result and the local pooling result to capture multi-scale context features, generate an adaptive mask according to the multi-scale context features, and use the adaptive mask to enhance the spatial features of the input feature map to obtain the enhanced feature map.

[0032] Preferably, the calculation method of the dynamic pooling segmentation number is: seg cout = β×ω / (2π); where β is the base segmentation coefficient, ω = 2πf is the angular frequency, and f is the average frequency and where, f r is the center frequency of the r-th frequency bin of the frequency band corresponding to the input feature map x; E r is the average amplitude of the r-th frequency bin of the frequency band corresponding to the input feature map x, and F is the total number of frequency bins;

[0033] The enhanced feature map y' = y local ⊙mask; where the adaptive mask mask = σ(W2×ReLU(W1×context)), where W1 and W2 respectively represent one-dimensional convolutional operations for dimension reduction and dimension expansion, σ is the Sigmoid function, ReLU is the activation function, and ⊙ represents element-wise multiplication; y local = Conv local (x) is to extract the spatial features of the input feature map x through one-dimensional convolution of the local branch, where Conv local represents the local convolution operation;

[0034] The multi-scale context feature context = GlobalMean(x) + SegPool(x, seg cout ); where GlobalMean represents the global average pooling operation, and SegPool represents the pooling operation based on the number of dynamic pooling segments.

[0035] Compared with the prior art, the present invention has the following beneficial effects: First, the input underwater acoustic signal is divided into three sub-bands: low frequency, medium frequency, and high frequency. The band-specific constant Q transform (BSCQT) is used to perform adaptive parameter filtering and weighting on each sub-band. By finely processing the low-frequency signal and suppressing the high-frequency noise, the quality and distinctiveness of the features are significantly improved, thereby realizing multi-band feature extraction and noise suppression, providing high-quality feature input for subsequent recognition, which is particularly important in the complex underwater environment of the Yellow River. Second, a dynamic context-aware mask (DCAM) mechanism is constructed. By analyzing the frequency, the receptive field size is adaptively adjusted, enabling the network to better adapt to the time-frequency changes of the underwater acoustic signal, improving the flexibility and accuracy of feature extraction, so that the network can capture more accurate features in the case where the signal features vary greatly with time and frequency; combining global and local pooling to capture multi-scale context information can comprehensively consider the overall characteristics and local details of the signal, enhancing the expression ability of the features, enabling the network to more comprehensively understand the structure and pattern of the signal when processing complex underwater acoustic signals, and further improving the recognition accuracy; implementing a lightweight adaptive attention allocation mechanism to enhance the sensitivity of the network to key features while maintaining a low computational complexity, suitable for running on resource-constrained underwater devices. Finally, based on DCAM, a dynamic context-aware mask network (DCAMNet) is designed, which includes a lightweight front-end convolutional network and three cascaded DCAM dense TDNN networks. Target recognition is completed through global pooling and fully connected layers, optimizing feature learning and information transmission, reducing redundant calculations, improving computational efficiency, and ensuring high-precision recognition through global pooling and fully connected layers. This structural design realizes efficient feature extraction and target recognition, and is particularly suitable for complex underwater environments. The present invention first realizes noise suppression and feature extraction of underwater acoustic signals through the BSCQT method, secondly adopts the DCAM mechanism to realize a lightweight attention allocation mechanism, and finally constructs the DCAMNet based on DCAM to complete the high-precision recognition task of underwater acoustic targets in the Yellow River. It achieves high recognition accuracy while maintaining low computational complexity in the actual noisy underwater environment, showing a significant improvement compared with the existing advanced methods, and is particularly suitable for being deployed on resource-constrained underwater devices, providing a practical solution for the underwater acoustic target recognition task. Description of the Drawings

[0036] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0037] Figure 1 This is the overall flowchart of the present invention.

[0038] Figure 2 This is the calculation block diagram of BSCQT of the present invention.

[0039] Figure 3 This is the structural schematic diagram of the DCAMNet network of the present invention.

[0040] Figure 4 This is the calculation block diagram of DCAM of the present invention. Detailed implementation manners

[0041] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0042] As Figure 1 shown, an underwater acoustic target recognition method based on BSCQT and DCAM, the specific steps include:

[0043] Step 1: First, preprocess the original acoustic signal and divide the data set: including label marking the collected underwater acoustic signals, non-overlapping sampling of the underwater acoustic signals of each category to generate samples, and randomly dividing all samples into a training set and a test set according to a 7:3 ratio.

[0044] The original acoustic signal is the directly collected and untreated underwater audio data, which contains the target acoustic signal and background noise. The collected underwater acoustic target data is based on the original acoustic signal, and category labels (such as ship types) are assigned to each sample through label marking. Subsequently, non-overlapping sampling is performed on the underwater acoustic target data of each category, that is, fragments are continuously and non-overlappingly extracted from the continuous audio signal at a fixed duration (such as 3 seconds) to generate a series of independent samples. Non-overlapping sampling means that each sample does not overlap in time. This method ensures that there is no time repetition between samples and avoids information redundancy. Finally, the generated samples are randomly divided into a training set and a test set according to a 7:3 ratio, laying a foundation for model training and evaluation.

[0045] Step 2: Divide the underwater acoustic signals in the training set and the test set into three sub-bands: low frequency, medium frequency, and high frequency. Use the band-specific constant Q transform (BSCQT) to perform adaptive parameter filtering and weighting on each sub-band, and achieve noise suppression and multi-band feature extraction through concatenating multi-band features, obtaining the BSCQT spectrogram of the underwater acoustic signal samples.

[0046] Divide the underwater acoustic signals into three sub-bands: low frequency (30 - 1000 Hz), medium frequency (1000 - 2000 Hz), and high frequency (2000 - 5000 Hz). Such division enables the use of adaptive feature extraction parameters according to the signal characteristics of different frequency bands, thereby achieving precise feature extraction in the low-frequency band, effectively suppressing noise in the high-frequency band, while retaining the unique information of each frequency band and improving the distinguishability of features. Use the band-specific constant Q transform (BSCQT) to perform adaptive parameter filtering and weighting on each sub-band, and achieve noise suppression and multi-band feature extraction through concatenating multi-band features, obtaining the BSCQT feature map corresponding to each underwater acoustic signal sample in the training set and the test set.

[0047] As Figure 2 shown, the method for performing adaptive parameter filtering and weighting on each sub-band by BSCQT is as follows: First, determine the number of filters where, f max,i and f min,i respectively represent the maximum frequency and the minimum frequency of sub-band i, n octave,i is the number of bins per octave of sub-band i, i represents the index of the sub-band; octave represents the octave, which is a logarithmic scale unit of frequency used to describe the relative change of frequency; bins represents the number of CQT (constant Q transform) filters within sub-band i, that is, the number of filters. Then calculate the center frequency of the k-th filter where, b represents the filter distribution density, which is a constant and is determined according to specific application requirements and signal characteristics; in underwater acoustic signal processing, it is adjusted according to the frequency characteristics of the target signal, k = 0, 1, …, K - 1, K represents the number of filters in sub-band k, which is determined by the frequency range: The center frequency f k is used for the design of the window function in CQT filtering to ensure the resolution characteristics of the filter on the logarithmic frequency scale. After determining the filtering parameters of the sub-band, perform CQT filtering on the sub-band:

[0048]

[0049] where, CQT i (c, n) represents the CQT coefficient at the c-th filter and time n in sub-band i; x i (m) is the m-th signal value of the signal in sub-band i, which is extracted from the original signal through a band-pass filter; wc,i (m) represents the window function (such as the Hanning window) of the c-th filter in sub-band i, and its form is: H c,i represents the window length of the c-th filter in sub-band i, which is inversely proportional to the center frequency f k and the calculation formula is: where f s is the sampling rate (a fixed attribute of the input audio); Q i is the quality factor (Q value) of sub-band i, which determines the frequency resolution and can be adjusted according to the characteristics of the sub-band; j is the imaginary unit, c is the filter index, c = 0, 1,..., N bins,i -1, and N bins,i is the total number of filters in sub-band i.

[0050] Perform sub-band weighting on the CQT coefficients, and obtain X low = W l × CQT low , X mid = W m × CQT mid , X high = W h × CQT high ; where CQT low , CQT mid and CQT high respectively represent the CQT coefficients of the low, middle, and high frequency bands, and W l , W m and W h are the corresponding learnable weight coefficients, and X low , X mid and X high respectively represent the sub-band spectrograms of the weighted low, middle, and high frequencies. The weight coefficients, as the trainable parameters of the deep learning network, can be set to the initial value of 1 and automatically adjusted through backpropagation. Such weighted processing can highlight the important frequency bands, suppress the frequency bands with more noise, make the contributions of each sub-band to the final features more balanced, and thus improve the discrimination and recognition accuracy.

[0051] Finally, obtain the multi-band feature representation by concatenating the spectrograms: Concatenate to obtain the final BSCQT spectrogram:

[0052] X BSCQT = concat(X low , X mid , X high )

[0053] where concat represents the spectrogram concatenation operation in the frequency dimension, and X BSCQTIs the final multi-band feature representation. Integrate the features of the low, middle, and high frequency sub-bands into a unified feature representation, enabling the subsequent network to utilize multi-band information simultaneously. The concatenation operation preserves the characteristics of each sub-band, facilitating the network to learn the relationships and interactions between frequency bands.

[0054] Step 3: Calculate the dynamic pooling segments based on frequency analysis, construct a dynamic context-aware mask (DCAM) mechanism, enhance the input feature map, and obtain the enhanced feature map.

[0055] Adaptive adjustment of the receptive field size, i.e., the number of dynamic pooling segments seg, through frequency analysis cout , where frequency analysis includes average frequency calculation and dynamic segment length calculation. Combine global and local pooling to capture multi-scale context information, implement a lightweight adaptive attention allocation mechanism, and maintain low computational complexity while achieving high recognition accuracy.

[0056] The method for the DCAM mechanism to process data is as follows:

[0057] Step 1): Extract the spatial features of the input features through one-dimensional convolution of the local branch:

[0058] y local = Conv local (x)

[0059] where x is the input feature map, Conv local represents the local convolution operation, and y local is the extracted spatial feature.

[0060] Step 2): Calculate the frequency characteristics and the number of dynamic pooling segments corresponding to the frequency band of the input feature map x:

[0061]

[0062] ω = 2πf

[0063] seg cout = β × ω / (2π)

[0064] where f is the average frequency; f u is the center frequency of the u-th frequency bin; E u is the average amplitude of the u-th frequency bin, F is the total number of frequency bins, ω is the angular frequency, β is the base segmentation coefficient, a fixed value set manually, and seg cout is the number of dynamic pooling segments. In this way, the DCAM mechanism realizes the adaptive adjustment of the receptive field size, enabling the pooling operation to dynamically adapt to the frequency characteristics of the input feature map.

[0065] Step 3): Fuse the multi-scale context features:

[0066] context = GlobalMean(x) + SegPool(x, seg cout )

[0067] Among them, GlobalMean represents the global average pooling operation, and SegPool represents the pooling operation based on the number of dynamic pooling segments. By combining global pooling and segmented pooling, the model can capture the overall trend (global information) and local changes (local information) of the signal simultaneously, enhancing the robustness and discriminability of features.

[0068] Step 4): Generate an adaptive mask mask and the enhanced feature map y':

[0069] mask = σ(W2 × ReLU(W1 × context))

[0070] y' = y local ⊙ mask

[0071] Among them, W1 and W2 represent one-dimensional convolution operations for dimension reduction and dimension expansion respectively, σ is the Sigmoid function, ReLU is the activation function, and ⊙ represents element-wise multiplication. The feature map y' is input into the subsequent layers of the network for further feature processing and classification tasks.

[0072] The specific process implemented by the DCAM mechanism is as follows: The feature map x first calculates the frequency characteristics of the signal through the frequency analysis unit to determine the number of dynamic pooling segments seg cout ; then, the context pooling unit combines the global average pooling and the local pooling operation based on dynamic segmentation to capture multi-scale context information context; next, the mask generation unit generates an adaptive mask mask through a structure of dimension reduction - activation - dimension expansion; finally, the adaptive mask mask is multiplied element-wise with the spatial feature y local extracted by the local convolutional layer to obtain the enhanced feature map y'.

[0073] Step 4: Construct a DCAM dense TDNN network based on the DCAM mechanism, use multiple DCAM dense TDNN networks to construct a dynamic context-aware mask network for deep learning, input the BSCQT spectrogram of the underwater acoustic signal in the training set into the dynamic context-aware mask network for training, save the trained model parameters, and obtain an underwater acoustic target recognition model.

[0074] Such as Figure 3As shown in the figure, the dynamic context-aware network includes a lightweight front-end convolutional network, multiple cascaded DCAM dense TDNN (Time Delay Neural Network) networks, and a classifier. Setting three DCAM dense TDNN networks allows the model to gradually extract features through multi-level cascading, achieving hierarchical learning of features. At the same time, the design of the three DCAM dense TDNN networks strikes a balance between network depth and computational overhead, being able to learn complex feature representations while avoiding overfitting caused by overly deep networks or excessive computational resource requirements.

[0075] The structure of the lightweight front-end convolutional network is as follows: It includes an initial 3×3 convolutional kernel and two lightweight residual blocks connected in sequence.

[0076] The initial 3×3 convolutional kernel plays the roles of feature extraction, channel adjustment, and possibly spatial downsampling in the lightweight front-end convolutional network, extracting low-level features from the input data, adjusting the channel dimension, and providing appropriate input feature maps for the subsequent two lightweight residual blocks, thereby supporting the efficient operation of the entire network and the lightweight design goal.

[0077] The lightweight residual block of the lightweight front-end convolutional network adopts a residual structure design: local information compression is performed through the main branch convolution of the 3×3 convolutional kernel, and information transmission is performed through the shortcut branch convolution of the 1×1 convolutional kernel. Its mathematical expression is:

[0078] F LFCN =Hswish(BN(Conv 3×3 (X)))+Conv 1×1 (X)

[0079] where BN is the batch normalization layer, and Hswish is the lightweight activation function. X represents the feature map input to the lightweight residual block. Conv 3×3 、Conv 1×1 represent the operations of the 3×3 convolutional kernel and the 1×1 convolutional kernel respectively.

[0080] The structure of the DCAM dense TDNN network is as follows:

[0081] The DCAM dense TDNN network performs depth information compression through three cascaded DCAM dense TDNN networks. The mathematical expression of each network is F=Transit(Dense(DCAM(Pool 5×1 (I))))), where Pool 5×1Denotes a 5×1 time-domain downsampling operation, which is used to reduce feature redundancy; DCAM is a dynamic context-aware masking mechanism that enhances the spatial features of the input feature map; Dense is a dense connection layer for feature reuse and gradient propagation; Transit is a 1×1 convolutional feature transformation layer for reducing the feature channel dimension and computational complexity. I represents the feature map input to the DCAM dense TDNN network.

[0082] The classifier structure is as follows:

[0083] The classifier performs temporal information aggregation on the feature map output by the last DCAM dense TDNN network and maps it to the class space through a fully connected layer. Its mathematical expression is y = Softmax(FC(GlobalPool(F3))), where GlobalPool is the global average pooling operation, FC is the fully connected layer, and Softmax is used to generate the class probability distribution. Here, F3 is the output feature map of the third DCAM dense TDNN network.

[0084] The method of data processing in the DCAM dense TDNN network is as follows: the BSCQT spectrogram is input into the lightweight front-end convolutional network (LFCN), and firstly the local feature extraction is performed through the initial 3×3 convolution kernel to generate a preliminary feature map. Specifically, the 3×3 convolution kernel performs a convolution operation on the input BSCQT spectrogram to extract the local time-frequency features of the signal; then, the preliminary feature map is normalized through the batch normalization (BN) layer to accelerate the training convergence and improve the generalization ability of the model; the normalized feature map is then nonlinearly transformed through the H-Swish activation function to obtain the activated feature map. At the same time, the input BSCQT spectrogram is also transmitted through the 1×1 convolution kernel to generate the feature map of the shortcut branch. By adding the activated feature map of the main branch to the feature map of the shortcut branch element by element, the local information compression of the lightweight residual block and the residual processing of information transmission are completed to obtain the feature map I. The feature map I is then input to the dynamic context-aware mask (DCAM) mechanism for processing to obtain an enhanced feature map; the enhanced feature map y′ is then processed by a dense connection layer (Dense) to enhance the feature expression capability through feature reuse and gradient propagation; then the feature is transformed by a 1×1 convolutional feature transformation layer (Transit) to reduce the feature channel dimension and computational complexity, and finally the output feature map F1 is obtained. The output feature map F1 is processed by the second DCAM dense TDNN network and the third DCAM dense TDNN network in turn. Each network repeats the above DCAM mechanism, dense connection layer and Transit layer processing flow, and finally the output feature map F3 is obtained. The output feature map F3 is input to the classifier, firstly through the global average pooling operation to compress the time domain features into a fixed-dimensional vector representation; then the fully connected layer (FC) is used to map the features to the target category dimension space; finally, the output is converted into a category probability distribution through the Softmax function to obtain the classification result y.

[0085] The BSCQT spectrogram of the underwater acoustic signal samples in the training set is input into the dynamic context-aware network for training. During the training process, the mapping relationship from the BSCQT spectrogram to the target category is learned. The AdamW optimizer is used with a learning rate of 0.001 and a BatchSize of 64 for 50 rounds of training. After each round of training, the validation set is used for testing, and the parameters of the round with the highest accuracy of the validation set are saved to obtain a trained underwater acoustic target recognition model.

[0086] The AdamW optimizer is adopted, with a learning rate of 0.001, a BatchSize of 64, and 50 rounds of training are carried out. After each round of training, the validation set is used for evaluation, and the model parameters of the round with the highest accuracy rate on the validation set are saved to obtain the trained underwater acoustic target recognition model. The cross-entropy loss function is used to guide the model learning during training, and the mathematical expression of the cross-entropy loss function is where M is the number of samples, Category is the number of categories, and y index,category is the true label, is the probability predicted by the model.

[0087] Step 5: Input the BSCQT spectrogram of the underwater acoustic signal in the test set into the trained underwater acoustic target recognition model to obtain the underwater acoustic target recognition result.

[0088] Finally, input the test set into the trained dynamic context-aware mask network to obtain the underwater acoustic target recognition result. As shown in Table 1, the DeepShip and ShipsEar datasets are used to verify the performance of the present invention. The DeepShip dataset contains 47 hours of underwater acoustic signals of 4 types of ships. The ShipsEar dataset is combined into 5 types of underwater acoustic signals (background noise and 4 types of ships of different sizes).

[0089] Table 1 Simulation data

[0090]

[0091] As can be seen from Table 1, on the DeepShip dataset, the weighted average precision rate and recall rate of the method proposed in the present invention reach 99.23% and 99.23% respectively; among them, the precision rate of tugboats reaches up to 99.09%, and the recall rate of tugboats is best at 99.65%. This benefits from the synergistic effect of BSCQT and DCAM. The DCAM mechanism utilizes the frequency information of the BSCQT spectrogram and dynamically adjusts the pooling operation to adapt to different frequency characteristics; at the same time, a dynamic mask is generated through global context awareness to enhance key time-frequency characteristics. This synergistic effect improves the network's adaptive processing ability for the time-frequency characteristics of audio signals and enhances the classification performance. On the ShipsEar dataset, the recognition accuracy rate of Category A reaches 100%, the precision rates of other categories are all above 94%, and the weighted average F1 score reaches 0.9640. In addition, the floating-point operation counts (FLOPs) of this method on the DeepShip and ShipsEar datasets are only 0.55G and 0.93G respectively. These results show that the present invention has achieved balanced and excellent recognition performance in each category of different datasets.

[0092] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.

Claims

1. An underwater acoustic target recognition method based on BSCQT and DCAM, characterized in that The steps are as follows: Step 1: Perform label marking on the collected underwater acoustic signals, perform non-overlapping sampling on the underwater acoustic signals of each category to generate samples, and randomly divide all samples into a training set and a test set; Step 2: Divide the underwater acoustic signals in the training set and the test set into three sub-bands: low frequency, medium frequency, and high frequency. Use the band-specific constant Q transform to perform adaptive parameter filtering and weighted extraction of multi-band features for each sub-band, and concatenate the multi-band features to obtain the BSCQT spectrogram of the underwater acoustic signal samples; Step 3: Construct a DCAM dense TDNN network based on the DCAM mechanism, and use multiple DCAM dense TDNN networks to construct a dynamic context-aware mask network for deep learning; Input the BSCQT spectrogram of the underwater acoustic signals in the training set into the dynamic context-aware mask network for training, save the trained model parameters, and obtain an underwater acoustic target recognition model; Step 4: Input the BSCQT spectrogram of the underwater acoustic signals in the test set into the trained underwater acoustic target recognition model to obtain the underwater acoustic target recognition result.

2. The underwater acoustic target recognition method based on BSCQT and DCAM according to claim 1, wherein, The method of performing non-overlapping sampling on the underwater acoustic signals of each category to generate samples is: continuously and non-overlappingly extract segments from the continuous audio signal at a fixed duration to generate a series of independent left-position samples of underwater acoustic signals; The frequency range of the low frequency of the sub-band is 30 - 1000 Hz, the frequency range of the medium frequency of the sub-band is 1000 - 2000 Hz, and the frequency range of the high frequency of the sub-band is 2000 - 5000 Hz.

3. The underwater acoustic target recognition method based on BSCQT and DCAM according to claim 1 or 2, characterized in that, The method of performing adaptive parameter filtering on each subband by the band-specific constant-Q transform is as follows: Determine the number of filters N of the constant-Q transform filter within subband i. bins,i ; Calculate the center frequency f of the k-th filter. k ; According to the number of filters N bins,i and the center frequency f k , perform constant-Q filtering on subband i using a window function to obtain CQT coefficients. The implementation method of the weighting is as follows: perform sub-band weighting on the CQT coefficients output by the constant Q filtering. The sub-band spectrograms of the weighted low-frequency, intermediate-frequency, and high-frequency are respectively: X low = W l × CQT low , X mid = W m × CQT mid , X high = W h × CQT high ; where CQT low , CQT mid and CQT high respectively represent the CQT coefficients of the sub-bands of low-frequency, intermediate-frequency, and high-frequency, and W l , W m and W h are learnable weight coefficients; The multi-band feature representation is obtained by concatenating sub-band spectrograms, and the BSCQT spectrogram X is obtained. BSCQT = concat(X low , X mid , X high ); where concat represents the spectrogram concatenation operation in the frequency dimension.

4. The underwater acoustic target recognition method based on BSCQT and DCAM according to claim 3, characterized in that, The number of said filters wherein, f max,i and f min,i respectively represent the maximum frequency and the minimum frequency of sub-band i, n octave,i is the number of bins per octave of sub-band i, and bins represents the number of constant-Q transform filters within sub-band i; for the k-th filter where b represents the filter distribution density, k = 0, 1, …, K−1, and K represents the number of filters in sub-band k, which is determined by the frequency range: The calculation method of the CQT coefficient is as follows: Among them, CQT i (c,n) represents the CQT coefficient at the c-th filter and time n in sub-band i; x i (m) is the m-th signal value of the signal in sub-band i, w c,i (m) represents the window function of the c-th filter in sub-band i; H c,i represents the window length of the c-th filter in sub-band i, which is inversely proportional to the center frequency f k ; Q i is the quality factor of sub-band i, j is the imaginary unit, c = 0, 1, …, N bins,i -1, N bins,i is the total number of filters in sub-band i.

5. The underwater acoustic target recognition method based on BSCQT and DCAM according to claim 1 or 4, characterized in that The dynamic context-aware network includes a lightweight front-end convolutional network, multiple DCAM dense TDNN networks, and a classifier connected in sequence; 6. The underwater acoustic target recognition method based on BSCQT and DCAM according to claim 5, characterized in that, The multiple DCAM dense TDNN networks are 3 cascaded DCAM dense TDNN networks for depth information compression; The lightweight front-end convolutional network includes an initial 3×3 convolutional kernel and two lightweight residual blocks connected in sequence; the lightweight residual block adopts a residual structure design: the main branch convolution through the 3×3 convolutional kernel performs local information compression, and the shortcut branch convolution through the 1×1 convolutional kernel performs information transmission; Each DCAM dense TDNN network includes multiple dynamic context-aware mask mechanisms, a dense connection layer, and a 1×1 convolutional feature transformation layer connected in sequence; The classifier includes a global average pooling layer, a fully connected layer, and a Softmax classification operation connected in sequence; 7. The underwater acoustic target recognition method based on BSCQT and DCAM according to claim 6, wherein, The processing method of the dynamic context-aware network is: input the BSCQT spectrogram into the lightweight front-end convolutional network, first perform local feature extraction through the initial 3×3 convolutional kernel to generate a preliminary feature map; subsequently, the preliminary feature map is normalized through a batch normalization layer to accelerate training convergence and improve the generalization ability of the model; The normalized feature map is then non-linearly transformed through the H-Swish activation function to obtain the activated feature map; The input BSCQT spectrogram also undergoes information transmission through a 1×1 convolutional kernel to generate the feature map of the shortcut branch; the activated feature map obtained from the main branch is element-wise added to the feature map of the shortcut branch to complete the residual processing of local information compression and information transmission of the lightweight residual block, obtaining the feature map I; the feature map I is input into the dynamic context-aware masking mechanism for processing to obtain the enhanced feature map; the enhanced feature map undergoes processing through a dense connection layer to enhance the feature expression ability through feature reuse and gradient propagation; then, a 1×1 convolutional feature transformation layer is used for feature transformation to reduce the feature channel dimension and computational complexity, finally obtaining the output feature map F1; the output feature map F1 successively undergoes the processing of the second DCAM dense TDNN network and the third DCAM dense TDNN network, and each network repeats the above processing flow of the dynamic context-aware masking mechanism, dense connection layer, and feature transformation layer, finally obtaining the output feature map F3; the output feature map F3 is input into the classifier, and through global average pooling operation, the time-domain features are compressed into a vector representation of a fixed dimension; the features are mapped to the target category dimension space through a fully connected layer; the output is converted into a category probability distribution through the Softmax function to obtain the classification result y.

8. The underwater acoustic target recognition method based on BSCQT and DCAM according to claim 6 or 7, characterized in that The calculation method of the lightweight residual block is as follows: F LFCN = Hswish(BN(Conv 3×3 (X)))+Conv 1×1 (X) Among them, BN is the batch normalization layer, and Hswish is the lightweight activation function. X represents the feature map input to the lightweight residual block. Conv 3×3 and Conv 1×1 represent the operations of 3×3 convolutional kernels and 1×1 convolutional kernels respectively. The calculation method for each DCAM dense TDNN network to obtain the output feature map is as follows: F = Transit(Dense(DCAM(Pool 5×1 (I))), where I represents the feature map input to the DCAM dense TDNN network, and Pool 5×1 represents a 5×1 time-domain downsampling operation; DCAM is a dynamic context-aware masking mechanism that enhances the spatial features of the input feature map; Dense is a dense connection layer for feature reuse and gradient propagation; Transit is a 1×1 convolutional feature transformation layer for reducing the feature channel dimension and computational complexity; The classifier aggregates the time-domain information of the feature map output by the last DCAM dense TDNN network and maps it to the category space through a fully connected layer, and the classification result is: y = Softmax(FC(GlobalPool(F3))); where GlobalPool is the global average pooling operation, FC is the fully connected layer, Softmax is used to generate the category probability distribution, and F3 is the output feature map of the last DCAM dense TDNN network.

9. The underwater acoustic target recognition method based on BSCQT and DCAM according to claim 8, characterized in that, For the dynamic context-aware masking mechanism, the method for enhancing the spatial features of the input feature map is: adaptively adjusting the receptive field size through frequency analysis to obtain the number of dynamic pooling segments; using the global average pooling operation to process the input feature map to obtain the global pooling result, using the number of dynamic pooling segments to perform local pooling operation on the input feature map to obtain the local pooling result, combining the global pooling result and the local pooling result to capture multi-scale context features, generating an adaptive mask according to the multi-scale context features, and using the adaptive mask to enhance the spatial features of the input feature to obtain the enhanced feature map.

10. The underwater acoustic target recognition method based on BSCQT and DCAM according to claim 9, characterized in that, The calculation method of the number of dynamic pooling segments is: seg cout = β×ω / (2π); where β is the basic segmentation coefficient, ω = 2πf is the angular frequency, and f is the average frequency and where f r is the center frequency of the r-th frequency bin in the frequency band corresponding to the input feature map x; E r is the average amplitude of the r-th frequency bin in the frequency band corresponding to the input feature map x, and F is the total number of frequency bins; The enhanced feature map y′ = y local ⊙ mask; where the adaptive mask mask = σ(W2 × ReLU(W1 × context)), where W1 and W2 respectively represent one-dimensional convolutional operations for dimension reduction and dimension expansion, σ is the Sigmoid function, ReLU is the activation function, and ⊙ represents element-wise multiplication; y local = Conv local (x) extracts the spatial features of the input feature map x through one-dimensional convolution of the local branch, where Conv local represents the local convolution operation; The multi-scale context feature context = GlobalMean(x) + SegPool(x, seg cout ); where GlobalMean represents the global average pooling operation, and SegPool represents the pooling operation based on the number of segments of dynamic pooling.

Citation Information

Patent Citations

  • Underwater acoustic signal identification method based on multi-branch trunk external attention network

    CN117312946A

Cited By

  • Signal feature sensing method based on multi-dimensional index weighted convolution

    CN120763592A