An end-to-end lightweight underwater acoustic target recognition system and method for waveform representation learning
By introducing shallow feature extractor module and frequency separation module in the water acoustic target recognition system, combining the scale adaptive maximum pooling unit and frequency separation unit, the problem of time information loss in traditional methods is solved, and efficient water acoustic target recognition is achieved.
Patent Information
- Application Number
- CN202510175094.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-18
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-02-18
AI Technical Summary
The existing water acoustic target recognition method based on waveform characterization faces the challenge of excessive time dimension when processing high sampling rate data. The traditional downsampling method leads to loss of time information and affects identification performance.
An end-to-end lightweight water acoustic target recognition system is proposed, including a shallow feature extractor module and a frequency separation module. The recognition performance of the model is improved through scale adaptive maximum pooling unit and frequency separation unit.
In the case of small parameters, the accuracy of water acoustic target recognition is improved, the pre-processing step of extracting spectral features is avoided, the robustness of the model is enhanced, and it is suitable for complex and changeable underwater environments.
Smart Images

Figure CN119646670B_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the field of underwater acoustic target recognition, and specifically relates to an end-to-end lightweight underwater acoustic target recognition system and method for waveform representation learning. Background Art
[0002] Hydroacoustic targets include ships, submarines, marine life, etc., among which ship radiated noise is one of the main sources of marine environmental noise. Identifying the radiated noise signals generated by ships plays a vital role in the development of intelligent hydroacoustic applications, especially in research fields such as maritime traffic management and underwater environment monitoring systems. In recent years, a large amount of research work has been devoted to the development of robust ship radiated noise identification systems. However, the complex and changeable marine environment, the complex ship radiated noise generation mechanism, and the limited ship radiated noise database make ship radiated noise identification extremely complex and challenging.
[0003] Most underwater acoustic target recognition methods rely on audio feature extraction methods to convert high-dimensional raw waveforms into two-dimensional compact representations - spectrograms. Some of the most commonly used spectrogram features include STFT, wavelet packets, LOFAR (Low Frequency Analysis and Recording), CQT (Constant-Q Transform), GFCC (Gammatone Frequency Cepstral Coefficient), MFCC (Mel Frequency Cepstral Coefficients), Mel spectrum, etc. Although spectrogram features can well describe the characteristics of underwater targets, spectrogram-based recognition methods require manual feature design and selection of filter parameters based on a large amount of prior knowledge of the target. It is difficult to obtain sufficient prior knowledge for unknown targets and complex and changing underwater environments. Therefore, when faced with unknown complex ocean sound fields, spectrogram features are not robust, which makes such methods difficult to apply in complex and changing underwater environments.
[0004] In recent years, waveform-based end-to-end acoustic signal classification methods have attracted widespread attention due to their ability to learn features directly from raw time-domain waveforms. Such methods do not require preprocessing and manual feature parameter tuning, and they can be mainly divided into two categories: one is to directly use convolutional layers to learn filters without restricting the learned filter parameters; the other is to provide some prior knowledge to the learnable neural layer to prompt the network to learn specific types of filters, such as Gabor, Sinc, Gammatone filters, etc. However, waveform-based methods still have some shortcomings that need to be addressed. On the one hand, a deep understanding of the frequency characteristics of ship radiated noise is crucial for the effective identification of underwater acoustic targets. In addition, explicitly extracting various frequency features helps to improve network performance. Frequency separation strategies in feature space have been successfully applied in the field of computer vision, including image denoising and super-resolution. Similarly, in audio and speech processing related applications, some waveform-based studies have achieved functions similar to frequency separation characteristics by using pre-initialized learnable filter banks to learn directly from raw waveform data. However, almost all of these methods only focus on frequency characteristics, but ignore the frequency characteristics of deep feature space. On the other hand, waveform-based methods avoid the preprocessing steps required by spectrogram-based methods, but still face the challenge of too large a temporal dimension when processing high-sampling rate data. An early common solution was to use continuous downsampling layers to reduce the data dimension. However, this traditional downsampling method inevitably leads to the loss of temporal information, which affects the recognition performance and limits the effectiveness of the method.
[0005] Therefore, studying how to improve the accuracy of waveform-based underwater acoustic target recognition is of great significance to improving the performance of underwater sonar recognition systems. Summary of the invention
[0006] The purpose of the present invention is to overcome the shortcomings of the existing underwater acoustic target recognition technology based on waveform representation, and to propose an end-to-end lightweight underwater acoustic target recognition system and method for waveform representation learning, so as to improve the target recognition performance with a small number of parameters.
[0007] In order to achieve the above objectives, this application proposes an end-to-end lightweight underwater acoustic target recognition system for waveform representation learning, including:
[0008] A shallow feature extractor module, used to extract shallow features from the original hydroacoustic waveform;
[0009] Several frequency separation modules are used to extract components of different frequencies from shallow features and use global average pooling to aggregate feature vectors;
[0010] A classifier is used to give the type of the hydroacoustic target to the feature vector output by the frequency separation module.
[0011] As an improvement of the above system, the shallow feature extractor module includes:
[0012] 3 sequentially connected Conv1D-BN-ReLU blocks; each of the Conv1D-BN-ReLU blocks includes a one-dimensional convolution layer, a batch normalization, and an activation function;
[0013] The second and third Conv1D-BN-ReLU blocks also include a scale-adaptive max pooling unit;
[0014] The scale-adaptive maximum pooling unit includes a combination of several maximum pooling subunits and kernel selection subunits, and is finally connected to a downsampling operation subunit;
[0015] The processing process of the core selection subunit includes: firstly calculating two time series statistics of the characteristic signal of the underwater acoustic waveform, namely the mean and standard deviation of the characteristic signal; then concatenating the two time series statistics, and inputting the concatenated result into a subnetwork composed of two fully connected layers to obtain feature output, where Represents the number of maximum pooling sub-units; finally, a soft attention mechanism is used in the time dimension to decide whether to perform a dense maximum pooling operation.
[0016] As an improvement of the above system, the frequency separation module includes a frequency separation unit, a channel attention unit, a residual connection unit, a Conv1D-BN-ReLU block and a scale-adaptive maximum pooling unit connected in sequence;
[0017] The frequency separation unit uses one-dimensional depth-by-depth convolution and Softmax function to process shallow features to obtain the low-frequency part of the underwater sound waveform, and then concatenates the low-frequency part with the shallow features and processes them through one-dimensional point-by-point convolution to obtain a frequency-separated feature vector;
[0018] The channel attention unit includes compression and excitation stages; the compression stage includes a global average pooling operation; the excitation stage includes two fully connected layers; the first fully connected layer reduces the dimension of the input feature vector, and the second fully connected layer restores the feature vector to the original dimension, and uses the Sigmoid function to obtain the channel weight, which is then multiplied by the output feature of the frequency separation unit to obtain the channel attention weighted feature;
[0019] The residual connection unit is used to add the input features of the frequency separation module to the channel attention weighted features, and then input them into the Conv1D-BN-ReLU block and the scale-adaptive maximum pooling unit.
[0020] As an improvement of the above system, the system includes three frequency separation modules.
[0021] As an improvement of the above system, the training process of the system includes:
[0022] Step 1: Preprocess the original underwater acoustic signal, including resampling the original underwater acoustic signal data set, slicing the data set, and dividing it into a training set and a test set;
[0023] Step 2: Build an underwater acoustic target recognition system;
[0024] Step 3: Set the training hyperparameters, loss function and optimizer, and input the training set into the underwater acoustic target recognition system for training.
[0025] As an improvement of the above system, the slicing and dividing of the data set includes two methods, namely, slicing first and then dividing the data set and dividing the data set first and then slicing.
[0026] As an improvement on the above system, a sharpness-aware minimization optimizer was used during training, with an initial training learning rate of 0.001.
[0027] As an improvement of the above system, batch training is adopted during training, and the data set is input into the underwater acoustic target recognition system in batches; the training batch size is set to 256, and the number of training rounds is set to 100; the cross entropy loss is used as the optimization objective function.
[0028] The present application also provides an end-to-end lightweight underwater acoustic target recognition method for waveform representation learning, which is implemented based on the above system and includes:
[0029] The pre-processed underwater acoustic waveform is input into the trained underwater acoustic target recognition system to obtain the classification result of the underwater acoustic target.
[0030] Compared with the prior art, the advantages of this application are:
[0031] 1. The present invention takes advantage of the fine temporal structure in the original underwater acoustic signal and directly learns the representation from the original waveform, avoiding the preprocessing step of extracting spectrogram features;
[0032] 2. In view of the fact that existing methods only focus on shallow frequency characteristics and information loss caused by downsampling, this paper proposes two innovative units to improve the recognition performance of the model;
[0033] 3. The present invention can achieve good recognition accuracy with a small number of parameters, and achieves a good trade-off between recognition accuracy and model complexity;
[0034] 4. Aiming at practical application needs (such as underwater unmanned platforms), the present invention introduces lightweight design and optimizes the model parameters, and is expected to be deployed in resource-constrained environments. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] Figure 1 The figure shows the architecture of an end-to-end lightweight underwater acoustic target recognition system for waveform representation learning;
[0036] Figure 2 (a) shows a schematic diagram of a large-size maximum pooling structure;
[0037] Figure 2 (b) shows a schematic diagram of the small-size dense maximum pooling and downsampling structure;
[0038] Figure 2 (c) shows a schematic diagram of the scale-adaptive maximum pooling (AMP) unit structure;
[0039] Figure 2 (d) shows a schematic diagram of the nuclear selection (KS) subunit structure;
[0040] Figure 3 The figure shows the comparison of recognition performance and parameter quantities of different underwater acoustic target recognition methods;
[0041] Figure 4 Shown are the visualized t-SNE plots of all waveform-based methods on the two datasets;
[0042] Figure 5 Shown are the confusion matrix plots of some waveform-based and spectrogram-based methods on two data;
[0043] Figure 6 Shown is a visualization of the high- and low-frequency features of the frequency separation module;
[0044] Figure 7 Shown are feature output visualizations of traditional maximum pooling and the scale-adaptive maximum pooling module described in the present invention. DETAILED DESCRIPTION
[0045] The technical solution of the present application is described in detail below with reference to the accompanying drawings.
[0046] The present application provides an end-to-end lightweight underwater acoustic target recognition system for waveform representation learning, comprising: a shallow feature extractor module, a frequency separation module and a classifier.
[0047] The input of the underwater acoustic target recognition system is the original underwater acoustic waveform, and the shallow feature extractor module extracts low-level features from the original waveform. The shallow feature extractor module consists of three Conv1D-BN-ReLU blocks for extracting shallow features, where Conv1D-BN-ReLU represents a one-dimensional convolution layer accompanied by batch normalization (BN) and ReLU (Rectifier Linearity Unit) processes. In view of the non-stationary characteristics and time series length of the original waveform, downsampling operations are used in the network to expand the receptive field. In the first Conv1D-BN-ReLU block, one-dimensional convolution downsampling is selected (the step size is set to 4, and the commonly used step sizes include 1, 2, 3 or larger), while in the subsequent two Conv1D-BN-ReLU blocks, scale-adaptive maximum pooling units are used to retain important information. Assuming X is the original underwater acoustic waveform, the first Conv1D-BN-ReLU can be expressed as:
[0048]
[0049] Among them, Y is the output feature map, and It consists of BN and ReLU. The dimension is ( ) is a one-dimensional convolution kernel with a step length of .
[0050] The frequency separation module (FSBlock) includes a frequency separation (FS) unit based on frequency separation features, a channel attention (Squeeze and Excitation, SE) unit, a residual connection unit, a Conv1D-BN-ReLU block, and a scale adaptive max pooling (AMP) unit for retaining temporal information details. Among them, the frequency separation unit and the scale adaptive max pooling unit are two innovative units proposed by the present invention.
[0051] The overall structure of the frequency separation (FS) unit is as follows: Figure 1 As shown in the figure, it is mainly implemented through one-dimensional depth-wise convolution and one-dimensional point-wise convolution. First, the low-frequency component is obtained by constraining the frequency response of the depth-wise convolution kernel. The high-frequency component can be obtained by linearly combining the original signal with the low-frequency component. The specific implementation is to concatenate the low-frequency component with the original input feature and perform point-by-point operations to generate the final output feature. The detailed implementation process will be explained in conjunction with the formula later. Table 1 gives more detailed FS module parameter settings.
[0052] The frequency separation module can separate multiple frequency components. The unit uses one-dimensional depthwise convolution to obtain low-frequency components. , where the convolution weights are limited by the low-pass constraint. The process of the frequency separation unit is to first perform one-dimensional depth-wise convolution weights is randomly initialized and constrained by the Softmax function:
[0053]
[0054] in, and denote the low-pass constraint weights and the original convolution weights respectively. Represents the one-dimensional depth convolution kernel size. The Softmax function is used to ensure that the weights are positive and the sum is guaranteed to be one, thereby limiting the filter to a low-pass frequency response.
[0055] Then, the low-frequency part is transformed along the channel dimension and the original input features The concatenated features are processed by one-dimensional pointwise convolution to generate output features. :
[0056]
[0057] in, It is output channels, and Represents the weights of the point-wise convolution.
[0058] Considering the original input features Subtract low frequency components High frequency components can be obtained , this process can be expressed as follows:
[0059]
[0060] Obviously, the above formula is a special case of point-by-point convolution output. Therefore, the output of point-by-point convolution contains both high-frequency features and low-frequency features. and By performing random initialization and adjustment, an end-to-end frequency separation strategy can be implemented.
[0061] The channel attention (SE) unit can further refine the various frequency components learned from the frequency separation unit. The channel attention mechanism mainly includes two stages: compression and excitation. In the compression stage, the spatial dimension is compressed through global average pooling (pooling by channel), and each channel represents the global information of the channel feature. In the excitation stage, the global information of each channel is encoded and adjusted through two fully connected layers. The first fully connected layer reduces the dimension of the input feature vector, and the second fully connected layer restores it to the original dimension. Finally, the obtained channel attention weight is multiplied by the original input feature map channel by channel to obtain the output features of the strengthened important channels. The specific process is: given the output features from the frequency separation unit , the channel attention unit first compresses it into a one-dimensional vector through a global average pooling (GAP) operation Then, the channel weights Learning through two nonlinear fully connected layers and Sigmoid function:
[0062]
[0063] in, and are the weights of the two fully connected layers, Represents the Sigmoid function. Finally, by taking the input features of the channel attention unit With attention weight Multiply to get the channel attention weighted feature , is a refined feature that emphasizes the most relevant feature channels. This process can be expressed as:
[0064]
[0065] in Represents element-wise multiplication.
[0066] The residual connection unit adds the input features of the frequency separation module to the channel attention weighted features, and then inputs them into the Conv1D-BN-ReLU block and the scale-adaptive maximum pooling unit.
[0067] The scale-adaptive max pooling (AMP) unit can adaptively adjust the size of the receptive field as an alternative to the traditional max pooling operation. According to morphological decomposition: a single unit with a larger kernel size The maximum pooling layer can be constructed by stacking multiple layers with smaller kernel sizes. This process can be expressed as:
[0068]
[0069]
[0070] in, represents the input features, represents the maximum pooling operation, Represents a max pooling / dilation kernel. is the maximum number of pooling operations with smaller kernel sizes, each pooling kernel size is ( ).
[0071] The overall structure of the scale-maximum adaptive pooling unit is shown in Figure 2 (c). Its main goal is to choose whether to perform the maximum pooling operation through a dynamic decision mechanism. The unit uses the kernel selection (KS) subunit (as shown in Figure 2 (d)) to dynamically decide whether to perform the maximum pooling based on the input feature map. In Figure 2 (d), the KS subunit first calculates two temporal statistics of the input features respectively; then the calculated temporal statistics are used as the input of the gating mechanism to guide the selection of the adaptive pooling operation, which is specifically implemented through a subnetwork composed of two fully connected layers, and the subnetwork is used to generate the weights of the pooling decision; then the soft attention mechanism is applied to generate a weight for each pooling operation, which is implemented through the Softmax function; finally, the weight of the soft attention mechanism is used to decide whether to perform the maximum pooling operation. The AMP module contains three KS subunits. Each KS subunit performs temporal statistics calculation, decision generation guided by the gating mechanism, and weight generation of the soft attention mechanism on the input feature map. Finally, after all the maximum pooling operations are completed, downsampling operations are performed to obtain the final output feature map. The specific process of the AMP unit is to first calculate two time series statistics, namely the mean and standard deviation of the feature signal. The mean reflects the activation degree of the feature map, while the standard deviation represents the fluctuation or variability of the feature map; then, a gating mechanism is used to provide guidance for adaptive selection: the two time series statistics are concatenated and input into a subnetwork consisting of two fully connected layers, thereby obtaining feature output, where is the number of max pooling layers ( ); Next, a soft attention mechanism is used in the time dimension to decide whether to perform dense max pooling operations. The Softmax function is used to generate soft binary weights. After all max pooling operations, the nearest neighbor downsampling is performed to obtain the final output features. In other embodiments, the combination of large-size max pooling and kernel selection sub-units may be more than 3.
[0072] The frequency separation module can capture different frequency components in the feature space. The frequency separation module contains the residual connection between the frequency separation unit and the channel attention unit, accompanied by Conv-BN-ReLU and scale-adaptive maximum pooling units. The frequency separation module is defined as follows:
[0073]
[0074]
[0075] in, and are the input and output features, respectively, It is an intermediate feature. , and Represent frequency separation, channel attention, and scale-adaptive max pooling units, respectively. represents a one-dimensional convolution with a kernel size of 7. In other embodiments, the number of frequency separation modules may be greater than 3.
[0076] The classifier is used to classify the feature vector output by the frequency separation module and output the type of the underwater acoustic target. The classifier is a classifier of the existing neural network model.
[0077] The end-to-end lightweight underwater acoustic target recognition system for waveform representation learning provided in this application includes the following training process:
[0078] (1) Preprocess the original underwater acoustic signal, resample the original underwater acoustic data set, slice the data set, and divide it into training set and test set;
[0079] (2) Construct an end-to-end lightweight underwater acoustic target recognition system, which includes a shallow feature extractor module, a frequency separation unit, and a classifier;
[0080] (3) Set training hyperparameters, loss function, and optimizer, and input the training set into the recognition model for training;
[0081] (4) Use the trained model to identify targets in the test set and output the recognition accuracy.
[0082] The present application also provides an end-to-end lightweight underwater acoustic target recognition method for waveform representation learning, which is implemented based on the above system and includes:
[0083] The preprocessed underwater acoustic waveform is input into the trained underwater acoustic target recognition system to obtain the classification result of the underwater acoustic target.
[0084] The present invention proposes an end-to-end lightweight underwater acoustic target recognition system and method for waveform representation learning, which can achieve good recognition accuracy with a small number of parameters. Figure 1 The system as a whole consists of a shallow feature extractor module, three frequency separation modules, and a classifier consisting of two fully connected layers. The shallow feature extractor module extracts low-level features from the original waveform. These features are then fed into three cascaded frequency separation modules to further capture multiple frequency components, thereby obtaining a deeper and more discriminative representation. Specifically, the frequency separation module is mainly composed of two innovative units, namely a scale-adaptive maximum pooling unit for retaining temporal information details and a frequency separation unit for separating frequency features. For the features of the last frequency separation module, global average pooling is used to aggregate feature information to produce a single feature vector, which is then input into the classifier for category prediction.
[0085] The following is a description of the specific steps of an end-to-end lightweight underwater acoustic target recognition system and method for waveform representation learning in combination with the accompanying drawings and experimental results.
[0086] Step 001, the present invention uses two publicly available ship radiated noise datasets (ShipsEar and DeepShip) to verify the performance of the method. First, the two datasets are resampled with a sampling frequency of 16kHz. Then the datasets are sliced and divided into training sets and test sets. The present invention adopts two dataset division methods, which are as follows:
[0087] • Random segmentation: All original signals are segmented into 3-second segments with an overlap of 1.5 seconds, and then randomly divided into training and test sets in a ratio of 6:4. In this segmentation method, the segments in the training and test sets can come from the same signal.
[0088] • Unrelated segmentation: All original signals are first randomly divided into 8:2 ratios to generate training and test sets. For the network with waveform as input, each signal is initially divided into 3-second segments with an overlap of 1.5 seconds. During testing, the output probabilities of 10 3-second segments are averaged to produce the results of 30-second samples. For the model with spectrogram as input, each signal is divided into 30-second segments with an overlap of 15 seconds. Under this division method, the audio clips in the generated training and test sets do not belong to the same signal, which ensures that there is no overlap between the training and test sets, which is closer to the actual application scenario.
[0089] Step 002, construct an end-to-end lightweight underwater acoustic target recognition system, which includes a shallow feature extractor module, a frequency separation module, and a classifier. Table 1 provides the specific details of the recognition system. The frequency separation module is the basic component of the recognition system, and its details are shown in Table 2. Figure 2 (a)-Figure 2 (d) show the difference between the scale-adaptive maximum pooling module and other traditional maximum pooling layers, where Figure 2 (a) represents large-scale maximum pooling, Figure 2 (b) represents small-scale dense maximum pooling and downsampling, Figure 2 (c) represents the scale-adaptive maximum pooling unit, and Figure 2 (d) represents the kernel selection (KS) subunit, which decides whether to use the maximum pooling operation by selecting between the identity path and the maximum pooling path. Table 3 is the pseudo code of the scale-adaptive maximum pooling unit.
[0090] Table 1 Identification system parameters
[0091]
[0092] Table 2 Frequency separation module parameters
[0093]
[0094] Table 3 Pseudocode of scale-adaptive max pooling unit
[0095]
[0096] Step 003, set the training hyperparameters, optimizer and loss function, and use the training set to train the system and update the system parameters. Specifically, the sharpness-aware minimization (SAM) optimizer is used, and the initial training learning rate is 0.001. The training batch size is set to 256, and the number of training rounds is set to 100. In order to improve convergence, a step-by-step attenuation strategy is adopted for the learning rate. The learning rate is reduced by 0.1 times after the 0th and 60th rounds. The cross entropy loss is used as the optimization objective function.
[0097] Step 004: Use the trained system to predict the category of the test set, use the overall accuracy (Acc) to quantitatively evaluate the recognition performance, and use the floating-point operations (FLOPs) and network parameters to evaluate the model complexity. The Acc formula is as follows:
[0098]
[0099] in, is the number of correctly classified fragments, is the total number of fragments.
[0100] The recognition results are shown in Tables 4-7, where Tables 4 and 5 are the recognition performance comparison results of the random segmentation of the ShipsEar and DeepShip data sets, Tables 6 and 7 are the recognition performance comparison results of the unrelated segmentation of the ShipsEar and DeepShip data sets, and FSNet is the method proposed by the present invention. The values in bold indicate the best results, with the accuracy Acc in %, the FLOPs in G, and the parameter in M. As can be seen from Tables 4-7, the method of the present invention exhibits excellent recognition performance in all categories on the two data sets. At the same time, the method of the present invention has the lowest computational complexity among all waveform-based methods, and still has an advantage over the spectrogram-based ResNet. The method of the present invention has the least number of parameters among all methods, which proves that the method of the present invention has achieved the best balance between recognition performance and model complexity. In addition, the method of the present invention is the only waveform-based method that reaches or even exceeds the spectrogram-based method in performance, which not only demonstrates the potential of the waveform-based method, but also proves that the method proposed by the present invention can extract key discriminative information from the waveform. The method of the present invention performs well on both data sets, which means that the method is not very sensitive to the amount of data. Figure 3 (Left: random segmentation; right: irrelevant segmentation) is a comparison of the recognition performance and parameter quantity of different underwater acoustic target recognition methods on the DeepShip dataset. Obviously, the method of the present invention achieves higher accuracy and has lower model complexity.
[0101] Table 4 Comparison of recognition performance of ShipsEar and DeepShip datasets randomly split
[0102]
[0103] Table 5 Comparison of recognition performance of ShipsEar and DeepShip datasets randomly split
[0104]
[0105] Table 6 Comparison of recognition performance of unrelated segmentation of ShipsEar and DeepShip datasets
[0106]
[0107] Table 7 Comparison of recognition performance of unrelated segmentation of ShipsEar and DeepShip datasets
[0108]
[0109] The last layer feature output of all waveform-based methods on the two datasets is visualized using t-SNE (T-distribution Stochastic Neighbor Embedding), as shown in the figure below: Figure 4 The results show that the proposed method shows the most favorable clustering results on all datasets, which indicates that the proposed method can learn more discriminative features. In addition, the confusion matrices of the waveform-based and spectrogram-based methods are visualized on the two datasets. Figure 5 As shown. Figure 5 It can be seen that for the two data sets, the confusion matrix of the proposed method presents a clearer diagonal pattern, which indicates that it is significantly better than other compared methods in terms of recognition performance.
[0110] Figure 6 The feature output of the frequency separation unit is shown, highlighting the high-frequency and low-frequency features. This visualization clearly shows that the frequency separation unit explicitly separates the various frequency components, resulting in multiple frequency feature representations. Figure 7 The output of the scale-adaptive max pooling unit and the traditional max pooling are shown (pooling size is 3 and 7). It can be seen that when the input feature pattern is simpler, the scale-adaptive max pooling unit tends to use a smaller pooling size to preserve information (columns 2-3), while when the input feature pattern is more complex, the scale-adaptive max pooling unit tends to use a larger pooling size to smooth the noise (columns 4-5).
[0111] All components not specified in this embodiment can be implemented using existing technologies.
[0112] The present invention proposes an end-to-end lightweight underwater acoustic target recognition system and method for waveform representation learning, which is of great significance for marine traffic management, underwater environment monitoring systems, marine ecological protection, etc. The method of the present invention has high target recognition performance with a small number of parameters.
[0113] The present application may also provide a computer device, comprising: at least one processor, a memory, at least one network interface and a user interface. The various components in the device are coupled together through a bus system. It is understood that the bus system is used to achieve connection and communication between these components. In addition to the data bus, the bus system also includes a power bus, a control bus and a status signal bus.
[0114] The user interface may include a display, a keyboard or a pointing device, such as a mouse, a trackball, a touch pad or a touch screen.
[0115] It is understood that the memory in the disclosed embodiments of the present application may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct memory bus random access memory (DRRAM). The memories described herein are intended to include, but are not limited to, these and any other suitable types of memories.
[0116] In some embodiments, the memory stores the following elements, executable modules or data structures, or a subset thereof, or an extended set thereof: an operating system and applications.
[0117] The operating system includes various system programs, such as a framework layer, a core library layer, a driver layer, etc., which are used to implement various basic services and process hardware-based tasks. The application includes various application programs, such as a media player (Media Player), a browser (Browser), etc., which are used to implement various application services. The program for implementing the method of the embodiment of the present disclosure can be included in the application.
[0118] In the above embodiment, the processor may also call a program or instruction stored in the memory, specifically, a program or instruction stored in an application program, and is used to:
[0119] Execute the steps of the above method.
[0120] The above method can be applied to a processor or implemented by a processor. The processor may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by an integrated logic circuit of hardware in the processor or an instruction in the form of software. The above processor can be a general processor, a digital signal processor (Digital Signal Processor, DSP), an application specific integrated circuit (Application Specific Integrated Circuit, ASIC), a field programmable gate array (Field Programmable Gate Array, FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. The above-disclosed methods, steps and logic block diagrams can be implemented or executed. The general processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the above-disclosed method can be directly embodied as a hardware decoding processor to execute, or the hardware and software modules in the decoding processor can be combined to execute. The software module can be located in a mature storage medium in the field such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register, etc. The storage medium is located in the memory, and the processor reads the information in the memory and completes the steps of the above method in combination with its hardware.
[0121] It is understood that the embodiments described in the present application can be implemented by hardware, software, firmware, middleware, microcode or a combination thereof. For hardware implementation, the processing unit can be implemented in one or more application specific integrated circuits (ASIC), digital signal processors (DSP), digital signal processing devices (DSPD), programmable logic devices (PLD), field programmable gate arrays (FPGA), general-purpose processors, controllers, microcontrollers, microprocessors, other electronic units for performing the functions described in the present application or a combination thereof.
[0122] For software implementation, the technology of the present application can be implemented by executing the functional modules (such as procedures, functions, etc.) of the present application. The software code can be stored in a memory and executed by a processor. The memory can be implemented in the processor or outside the processor.
[0123] The present application may also provide a non-volatile storage medium for storing a computer program. When the computer program is executed by a processor, each step in the above method embodiment can be implemented.
[0124] Finally, it should be noted that the above embodiments are only used to illustrate the technical solution of the present application and are not intended to limit it. Although the present application is described in detail with reference to the embodiments, a person skilled in the art should understand that any modification or equivalent replacement of the technical solution of the present application does not depart from the spirit and scope of the technical solution of the present application and should be included in the scope of the claims of the present application.
Claims
1. An end-to-end lightweight underwater acoustic target recognition system for waveform representation learning, characterized in that: include: A shallow feature extractor module, used to extract shallow features from the original hydroacoustic waveform; Several frequency separation modules are used to extract components of different frequencies from shallow features and use global average pooling to aggregate feature vectors; A classifier, used to classify the feature vector output by the frequency separation module and output the type of the hydroacoustic target; The frequency separation module includes a frequency separation unit, a channel attention unit, a residual connection unit, a Conv1D-BN-ReLU block and a scale-adaptive maximum pooling unit connected in sequence; The frequency separation unit uses one-dimensional depth-by-depth convolution and Softmax function to process shallow features to obtain the low-frequency part of the underwater sound waveform, and then concatenates the low-frequency part with the shallow features and processes them through one-dimensional point-by-point convolution to obtain a frequency-separated feature vector; The channel attention unit includes compression and excitation stages; the compression stage includes a global average pooling operation; the excitation stage includes two fully connected layers; the first fully connected layer reduces the dimension of the input feature vector, and the second fully connected layer restores the feature vector to the original dimension, and uses the Sigmoid function to obtain the channel weight, which is then multiplied by the output feature of the frequency separation unit to obtain the channel attention weighted feature; The residual connection unit is used to add the input features of the frequency separation module to the channel attention weighted features, and then input them into the Conv1D-BN-ReLU block and the scale-adaptive maximum pooling unit.
2. The end-to-end lightweight underwater acoustic target recognition system for waveform representation learning according to claim 1 is characterized in that: The shallow feature extractor module includes: 3 sequentially connected Conv1D-BN-ReLU blocks; each of the Conv1D-BN-ReLU blocks includes a one-dimensional convolution layer, a batch normalization, and an activation function; The second and third Conv1D-BN-ReLU blocks also include a scale-adaptive max pooling unit; The scale-adaptive maximum pooling unit includes a combination of several maximum pooling subunits and kernel selection subunits, and is finally connected to a downsampling operation subunit; The processing process of the core selection subunit includes: firstly calculating two time series statistics of the characteristic signal of the underwater acoustic waveform, namely the mean and standard deviation of the characteristic signal; then concatenating the two time series statistics, and inputting the concatenated result into a subnetwork composed of two fully connected layers to obtain feature output, where Represents the number of maximum pooling sub-units; finally, a soft attention mechanism is used in the time dimension to decide whether to perform a dense maximum pooling operation.
3. The end-to-end lightweight underwater acoustic target recognition system for waveform representation learning according to claim 1 is characterized in that: The system comprises three frequency separation modules.
4. The end-to-end lightweight underwater acoustic target recognition system for waveform representation learning according to claim 1 is characterized in that: The training process of the system includes: Step 1: Preprocess the original underwater acoustic signal, including resampling the original underwater acoustic signal data set, slicing the data set, and dividing it into a training set and a test set; Step 2: Build an underwater acoustic target recognition system; Step 3: Set the training hyperparameters, loss function and optimizer, and input the training set into the underwater acoustic target recognition system for training.
5. The end-to-end lightweight underwater acoustic target recognition system for waveform representation learning according to claim 4 is characterized in that: The slicing and dividing of the data set includes two methods, namely, slicing first and then dividing the data set and dividing the data set first and then slicing.
6. The end-to-end lightweight underwater acoustic target recognition system for waveform representation learning according to claim 4 is characterized in that: The sharpness-aware minimization optimizer was used for training, and the initial training learning rate was 0.
001.
7. The end-to-end lightweight underwater acoustic target recognition system for waveform representation learning according to claim 4 is characterized in that: Batch training is adopted during training, and the data set is input into the underwater acoustic target recognition system in batches; the training batch size is set to 256, and the number of training rounds is set to 100; the cross entropy loss is used as the optimization objective function.
8. An end-to-end lightweight underwater acoustic target recognition method for waveform representation learning, implemented based on the system of any one of claims 1 to 7, comprising: The pre-processed underwater acoustic waveform is input into the trained underwater acoustic target recognition system to obtain the classification result of the underwater acoustic target.
Citation Information
Patent Citations
Ship target identification method and system based on underwater acoustic signals
CN117454240A
All-night snore detection method based on frequency domain attention and self-attention pooling double-fusion mechanism
CN119360902A