A Method for Extracting Features of Underwater Target Acoustic Signals Based on Deep Manifold Learning

Through the DMLAE method of deep manifold learning, combined with generator and discriminator, the noise interference and structural retention problems of target acoustic signal feature extraction in water in complex marine environments are solved, and more efficient feature extraction and recognition are achieved.

CN117216516BActive Publication Date: 2025-08-05TONG FANG ELECTRONICS SCI & TECH +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311109643.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-08-31
Publication Date
2025-08-05
Estimated Expiration
2043-08-31

AI Technical Summary

Technical Problem

When the prior art deals with the extraction of target acoustic signal characteristics in water in complex marine environments, there are problems such as high noise interference, high computational volume, susceptible to noise, and deep learning methods fail to effectively retain the internal structure of the target.

Method used

A deep manifold learning-based method is adopted to construct a deep manifold autoencoder (DMLAE), combined with generator and discriminator, maintain the nearest neighbor reconstruction weights through unsupervised learning and manifold learning, introduce adversarial regularization processing, and extract the global and local characteristics of the target acoustic signal in water.

Benefits of technology

The recognition accuracy of the extraction of the feature of target acoustic signals in water is improved, with an average increase of 14.96%, effectively retaining the internal structure and topological characteristics of the target acoustic signals.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117216516B_ABST
    Figure CN117216516B_ABST
Patent Text Reader

Abstract

This paper discloses a method for extracting underwater target acoustic signal features based on deep manifold learning. Specifically, the method involves: first, globally optimizing the original data through autoencoder reconstruction errors to identify a potential low-dimensional representation; then, leveraging manifold learning to maintain neighbor reconstruction weights, locally constraining the potential representation to preserve its intrinsic topological structure; finally, introducing a generative adversarial network architecture for regularization, forcing the potential representation to obey a specific distribution, thereby achieving a method that jointly maintains low-dimensionality in both local and global embedding. This method was experimentally validated on the DeepShip public dataset of deepwater vessels. Compared with existing deep learning and manifold learning feature extraction methods, the recognition accuracy was improved by an average of 14.96%.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of underwater target acoustic signal feature extraction, and in particular to an underwater target acoustic signal feature extraction method based on deep manifold learning. Background Art

[0002] In recent years, the continuous development and application of noise and vibration reduction technologies have made the characteristics of underwater target acoustic signals increasingly less distinct. Furthermore, in complex marine environments, the target sound field data acquired and stored by underwater acoustic equipment is modulated by numerous environmental parameters and affected by multi-channel transmission. This results in the target data being processed exhibiting numerous complex characteristics, undoubtedly posing a severe challenge to traditional underwater target acoustic signal feature extraction techniques.

[0003] From the perspective of traditional underwater target acoustic signal feature extraction, research methods mainly focus on the time domain, frequency domain, and nonlinear domain. Representative studies include feature extraction methods based on spectrum estimation, time-frequency analysis, and chaos theory. Reference 1 (Zhao Y, Zhang X, Huang J, et al. Techniques on ship radiated noise power spectrum feature extraction and realization [C] / / 2010 The 2nd International Conference on Computer and Automation Engineering (ICCAE). IEEE, 2010, 1: 554-558.) and Reference 2 (Huang J, Zhao J, Xie Y. Source classification using pole method of AR model [C] / / 1997 IEEE International Conference on Acoustics, Speech, and Signal Processing. IEEE, 1997, 1: 567-570) respectively use the traditional spectrum estimation Welch method and the modern spectrum estimation AR model to achieve classification and recognition, both achieving high recognition rates. Reference 3 (Wang Na, Chen Ke'an. Underwater noise timbre attribute regression model and its application in target recognition [J]. Acta Physica Sinica, 2010, 59(4): 2873-2881.) also achieved relatively ideal recognition results by constructing Mel-frequency cepstral coefficients and other feature quantities for target recognition. Reference 4 (Peng Yuan, Shen Liran, Li Xueyao, et al. Feature extraction and classification of underwater target radiated noise based on bispectrum [J]. Journal of Harbin Engineering University, 2003, 24(4): 390-394.) uses bispectral analysis to extract the amplitude of the target non-Gaussian vector to classify the target. The target sound field feature extraction method based on spectral estimation is simple, mature, and widely used. However, this method requires the assumption that the analysis signal is stationary, which is obviously not compatible with non-stationary target sound signals in complex environments.Reference 5 (Liu Y, Zhang X, Yu Y. Classification of vessel targets using wavelet statistical features [C] / / 2012 5th International Congress on Image and Signal Processing. IEEE, 2012: 1551-1555.) and reference 6 (Azimi-Sadjadi MR, Yao D, Huang Q, et al. Underwater target classification using wavelet packets and neural networks [J]. IEEE Transactions on neural networks, 2000, 11 (3): 784-794.) proposed using wavelet / wavelet packet transform to perform linear time-frequency decomposition of underwater target acoustic signals, and extracted the features of each wavelet frequency band of the sound field for target recognition, both achieving recognition rates of over 90%. Reference 7 (Wang S, Zeng X. Robust underwater noise targets classification using auditory inspired time-frequency analysis [J]. Applied Acoustics, 2014, 78: 68-76.) uses HHT transform to adaptively obtain multiple inherent modes of the target sound field signal, and extracts the weighted instantaneous spectrum centroid features of each inherent mode as the recognition feature vector. The SVM results show the excellence of its classification performance. Reference 8 (Chen J, Han B, Ma X, et al. Underwater target recognition based on multi-decision lofar spectrum enhancement: A deep-learning approach [J]. Future Internet, 2021, 13 (10): 265.) proposes to use the improved LOFAR map as the underwater target acoustic signal feature, and uses CNN neural network for classification and recognition, achieving more than 95% target recognition results. However, as a modern signal analysis method, the time-frequency analysis algorithm has many algorithmic problems that cannot be avoided, such as large computational complexity and susceptibility to noise interference.Reference 9 (Kantz H, Schreiber T. Nonlinear time series analysis [M]. Cambridge University Press, 2004.) proposed that the chaotic characteristic analysis of target radiation noise mainly focuses on calculating its Lyapunov exponent, fractal dimension, estimated entropy, Poincare cross section, natural measure, etc. to achieve the purpose of target recognition. However, this method cannot guarantee that the target acoustic signal still has chaotic characteristics under low signal-to-noise ratio conditions, which makes the algorithm's detection model invalid.

[0004] The above methods have achieved certain results in the feature extraction of underwater target acoustic signals, but they still have various limitations when processing large-scale, nonlinear noise data. At present, the feature extraction of underwater target radiation noise based on deep learning has gradually emerged. For target acoustic signals in complex ocean environments, deep learning can better extract feature information by taking advantage of its powerful nonlinear learning. Reference 10 (Kamal S, Mohammed SK, Pillai P RS, et al. Deep learning architectures for underwater target recognition [C] / / 2013 Ocean Electronics (SYMPOL). IEEE, 2013: 48-54.), Reference 11 (Chen Y, Xu X. Theresearch of underwater target recognition method based on deep learning [C] / / 2017 IEEE International Conference on Signal Processing, Communications and Computing (ICSPCC). IEEE, 2017: 1-5.), Reference 12 (Luo X, Feng Y. An underwater acoustic target recognition method based on restricted Boltzmann Machine [J]. Sensors, 2020, 20(18): 5399.) Three deep learning networks were applied to the feature extraction of underwater target radiated noise: restricted Boltzmann machine, deep belief network, and autoencoder. These deep learning algorithms are all based on the encoder network architecture. By extracting the frequency domain features of the acoustic target (power spectrum, LOFAR features, DEMON features, etc.) as network input, they can extract the features of underwater target acoustic signals and achieve good experimental results. However, these methods only make improvements to the neural network structure, input feature parameters, and training methods, without starting from the data itself and considering the structural characteristics between samples. As a result, the extracted features do not well preserve the intrinsic structure of the target.

[0005] Manifold learning, a nonlinear dimensionality reduction method, can discover and learn the intrinsic properties and internal structure of data. This advantage offsets the shortcomings of deep autoencoders. References 13 (Jain V, Saul LK. Exploratory analysis and visualization of speech and music by locally linear embedding [C] / / 2004 IEEE International Conference on Acoustics, Speech, and Signal Processing. IEEE, 2004, 3: iii-984.) and 14 (Errity A. Exploring the dimensionality of speech using manifold learning and dimensionality reduction methods [D]. Dublin City University, 2010.) proposed that the classification of acoustic signals is often determined by a few low-dimensional parameters. Reference 15 (Liu Hui, Yang Junan, Wang Yi. Research on acoustic target feature extraction method based on manifold learning [J]. Acta Physica Sinica, 2011, 60(7): 437-443.), Reference 16 (Liu Hui, Yang Junan, Wang Yi. Acoustic target feature extraction based on decorrelation neighborhood preserving discriminant projection [J]. Journal of Electronic Measurement and Instrumentation, 2010, 24(10): 905-910.), Reference 17 (Liu Hui, Yang Junan, Wang Yi, et al. Isometric mapping based on improved geodesic distance and its application in acoustic target feature extraction [J]. Acta Armamentarii, 2012, 33(10): 1178.) applied manifold learning to target acoustic signal feature extraction and achieved good classification results. Reference 18 (Lü Zhichao, Wang Haozhong, Bai Yiqi. Application of Manifold Learning in Shallow Water Acoustic Communication [J]. Journal of Electronics & Information Technology, 2021, 43:3) uses manifold learning to achieve dimensionality reduction for underwater acoustic signals. This reference demonstrates that manifold learning can also be used for feature extraction of underwater target acoustic signals. However, the manifold learning algorithm still has inherent flaws and performs poorly when processing large-scale, out-of-sample, and noisy data. Summary of the Invention

[0006] The purpose of the present invention is to provide a method for extracting features of underwater target acoustic signals based on deep manifold learning. This method obtains the intrinsic characteristics of the data through unsupervised learning, which not only solves the generalization learning and noise problems existing in manifold learning, but also enables the autoencoder to recover the potential representation that preserves the local geometric structure.

[0007] The technical solution to achieve the purpose of the present invention is: a method for extracting underwater target acoustic signal features based on deep manifold learning, comprising the following steps:

[0008] Step 1: Build a deep manifold autoencoder (DMLAE) model. The DMLAE model includes a generator and a discriminator, and initialize the model parameters.

[0009] Step 2: Input the data sample X into the generator input layer, optimize the objective function in the generator autoencoder network, and obtain the global feature Y of the target sound signal in the hidden layer;

[0010] Step 3: Calculate the nearest neighbor reconstruction weight matrix s of the target acoustic signal sample point;

[0011] Step 4: extract the local features of the underwater target acoustic signal and add local constraints to the global feature Y;

[0012] Step 5: Add L1 regularization term and rebuild the objective function;

[0013] Step 6: Perform adversarial regularization on the latent representation Y;

[0014] Step 7: Iteratively update the weights and biases of the deep manifold autoencoder network and obtain the optimal parameters to complete the DMLAE model training;

[0015] Step 8: Input the underwater target acoustic signal data into the trained DMLAE model, implement data feature extraction in the generator encoding network, and output the intrinsic feature matrix of the underwater target acoustic signal.

[0016] Furthermore, the DMLAE model includes a generator and a discriminator, as follows:

[0017] The generator is implemented by adding manifold regularization to the autoencoder, completing the global and local feature extraction of underwater target acoustic signal data. The autoencoder network in the generator extracts the global features of the data samples under unsupervised conditions. This global feature can characterize the overall properties of the target acoustic signal, but cannot reflect the spatial structure of the data. The idea of manifold learning is introduced into the autoencoder, and the neighbor reconstruction matrix between the original sample points is used to construct the local structural relationship between the hidden layer samples. This allows the extracted features to reflect the local properties of the target acoustic signal, thus achieving local feature extraction.

[0018] The discriminator distinguishes the true distribution of underwater target acoustic signal samples from the potential representation distribution generated by the generator, and continuously updates the generator through iteration to ensure that the distribution of the potential representation is consistent with the true distribution of the data.

[0019] Furthermore, in step 1, the initialization model parameters include: n underwater target acoustic signal data sample matrix X, X has D-dimensional features, output feature dimension q, maximum number of iterations iter, manifold learning local constraint balance parameter α, L1 regularization balance parameter λ, and reconstruction weight matrix neighbor coefficient k.

[0020] Furthermore, in step 2, the underwater target acoustic signal data sample matrix X=[x1, x2, ...x n ]∈R D×n , where x i is the feature vector of a single underwater acoustic sample, n represents the number of samples in the input network, and D is the feature dimension of a single sample;

[0021] The self-encoding network of the generator encodes the underwater target acoustic signal data sample matrix x and obtains the target acoustic signal global feature Y=[x1, y2, ...y n ]∈R d×n , where y i is the feature vector of the potential representation, d is the feature dimension of a single sample of the potential representation; by decoding Y in the hidden layer, the reconstructed underwater acoustic data X′ is obtained;

[0022] In the encoding and decoding process, by setting the hidden layer neurons, the underwater acoustic data is compressed into the potential representation Y, the unimportant features of the underwater target acoustic signal are removed, and the important features of the original target acoustic signal are restored to achieve feature extraction, which is expressed as

[0023] Y=f(W1X+b1) (1)

[0024] X′=h(W2Y+b2) (2)

[0025] Formula (1) is the encoding network mapping function, target acoustic signal - intrinsic features;

[0026] Formula (2) is the decoding network mapping function, intrinsic features - reconstructed target acoustic signal;

[0027] Parameters W1, W2 and b1, b2 represent the weights and biases of the deep manifold autoencoder network, respectively, and f and h are nonlinear activation functions;

[0028] The optimization is performed by minimizing the loss function, and the objective function is expressed as

[0029]

[0030] Where, L r is the reconstruction error loss function of the deep manifold autoencoder network.

[0031] Furthermore, in step 3, the neighbor reconstruction weight matrix of the target acoustic signal sample point is calculated as follows:

[0032] The idea of manifold learning to maintain neighbor weights is introduced into the hidden layer potential representation. For the underwater target acoustic signal dataset sample X = [x1, x2, ...x n ]∈R D×n , find each sample point x according to the Euclidean distance between these data points i k nearest neighbors N(x i ), the topological structure of the signal is determined by the local neighbor relationship of each target acoustic signal sample, N(x i ) is the set of neighboring points, k is the number of neighbors;

[0033] By the neighboring point N(x i ) construct a reconstruction weight vector s i , get the reconstruction weight matrix s=[s1,s2,...s n ]∈R n×n , the matrix s retains the structural relationship between the neighbors of each sample, and realizes the local feature extraction of the data by keeping the reconstruction weight unchanged in the low-dimensional space;

[0034] The neighbor reconstruction weight is calculated by minimizing the reconstruction error of formula (4):

[0035]

[0036] Where, the neighbor reconstruction weight s ij Reflects the sample point x i x j The size of the reconstruction contribution;

[0037] Simplifying formula (4) we can get:

[0038]

[0039] in

[0040] Q=(x i -x j )(x i -x j ) T ∈R k×k

[0041] S i =(s i1 , s i2 ,B,s ik )∈R k×1

[0042] From the above formula, we can see that for the rotation and scaling of the sample points, formula (4) can keep the reconstruction weight unchanged;

[0043] In order to achieve the same translation, for all x i Give constraints∑ j s ij =s i T l k =1, where l k ∈R k×1 is a unit vector, and the Lagrange multiplier method is used to solve it. ξ is the Lagrange multiplier as shown below:

[0044]

[0045] Obtain the reconstruction weight vector s i The closed-form solution of :

[0046]

[0047] Furthermore, in step 4, local features of the underwater target acoustic signal are extracted and local constraints are added to the global features, as follows:

[0048] The DMLAE model extracts local features of underwater target acoustic signals and requires a low-dimensional potential representation y i with y i The neighboring points of can reflect the reconstruction weight relationship of the original sound sample, that is, minimize the objective function of the following formula:

[0049]

[0050] In order to eliminate the translation, rotation and scaling factors of the low-dimensional embedding coordinate Y, the following constraints are added to Equation (8):

[0051] ①

[0052] ②

[0053] Simplified:

[0054] φ(Y)=||Y(IS T )|| 2 =tr(YMY T ) (9)

[0055] Where I is the identity matrix, S is the reconstruction weight matrix, tr() is the function for finding the matrix trace, and M = (Is) T (Is) is an n×n matrix.

[0056] Furthermore, the L1 regularization term is added as described in step 5, and the objective function is reconstructed as follows;

[0057] The overall optimization of the autoencoder is combined with the local optimization of manifold learning to achieve global and local feature extraction of underwater target acoustic signals. At the same time, the L1 regularization term is added to the optimization function. The L1 expression is shown in formula (10):

[0058]

[0059] Where W1 is the encoder weight, w i is the weight vector of each neuron in the encoding network in the model generator. According to the DMLAE model generator principle, the objective function is redefined as follows:

[0060]

[0061] where the parameters θ = (W1, b1, W2, b2), X′ is the signal reconstructed by the generator network from the original target acoustic signal X, α is the balance parameter of the local constraint for manifold learning, and λ is the balance parameter of the L1 regularization component of the encoder network;

[0062] For the solution of the objective function J(θ) in the above formula, the gradient descent method is used to obtain the global optimal solution of the convex function.

[0063] Furthermore, in step 7, the weights and biases W1, b1, W2, and b2 of the deep manifold autoencoder network are iteratively updated, and the optimal parameters θ = (W1, b1, W2, and b2) are obtained as follows:

[0064] The error terms δ1 and δ2 are introduced, corresponding to the output layer and hidden layer respectively, to measure the responsibility of the same node for the error. The calculation method is as follows:

[0065] δ1=(x′-x)·σ′(z1) (12)

[0066] δ2=(W2 T ·δ1 T )·σ′(z2) (13)

[0067] Among them, z1=W2y+b2, z2=W1x+b1, σ′ function is the derivative of the activation function in the text;

[0068] Solve for the gradient of θ:

[0069]

[0070]

[0071]

[0072]

[0073] Then use the gradient and learning rate α to update the parameters to obtain the optimal solution of the objective function:

[0074]

[0075]

[0076] Where, the learning rate α controls the network’s convergence to the local optimum and the convergence rate; W 1,2 Represents W1 and W2, b 1,2 Represents b1 and b2.

[0077] Furthermore, in step 6, the potential representation Y is subjected to adversarial regularization as follows:

[0078] The ship radiation noise is approximated as Gaussian noise, and the discriminator aims to restore the true distribution of the potential representation in the generator as much as possible;

[0079] The discriminator algorithm works by updating the discriminant network to distinguish between the true distribution and the potential distribution, and then updating the DMLAE generator network to optimize the potential distribution structure. For a given single sample data x of an underwater target acoustic signal, the network generator obtains the potential representation y, where p(y) is the true distribution of the acoustic signal and q(y) is the aggregated posterior distribution of the data.

[0080] When using neural networks for training and testing, an adversarial approach is introduced in the hidden layer of the generator network to optimize the potential representation. The discriminator in the hidden layer judges the real data after sampling and the fake data generated by the generator, so that the aggregated posterior distribution q(y) matches the prior distribution p(y) to complete the regularization of the potential representation.

[0081] The adversarial regularization objective function is as follows:

[0082]

[0083] The above formula is the objective function formula of the GAN network, where D(x) is the discriminator output and G(x) is the generator output, so that the hidden layer obtains the desired feature expression result.

[0084] Furthermore, the DMLAE model selects a double hidden layer network structure, the number of hidden layer network nodes is set to 1000, and the number of encoder output layer nodes is set to 4;

[0085] Use ReLU and Sigmoid functions as nonlinear activation functions of neural networks;

[0086] The output range of the Sigmoid function is 0 to 1, which gives the output value a probabilistic interpretation. It is suitable for tasks that require the output to be interpreted as a probability. The mathematical expression is as follows:

[0087]

[0088] When the input value is less than or equal to 0, the output of the ReLU function is 0, so that only some neurons in the network respond to the input, as shown in the following numerical expression:

[0089]

[0090] The DMLAE model uses the error loss function and the cross entropy loss function. The mathematical expressions are as follows:

[0091]

[0092] L bce =-(x logp(x′)+(1-x)log(1-p(x′))) (24)

[0093] Among them, formula (23) is the mean square error loss, which is used for regression tasks; formula (24) is the binary cross entropy loss, x is the input data, p(x′) is the predicted probability, which is used in classification tasks.

[0094] Compared with the prior art, the present invention has the following significant advantages:

[0095] (1) Leveraging the advantages of neural networks in nonlinear data processing, we can learn the low-dimensional data features of underwater acoustic signals;

[0096] (2) Through the local preservation embedding idea of manifold learning, the extracted latent representation retains the intrinsic structure of the ship radiated noise, achieving the purpose of extracting the characteristics of underwater target acoustic signals;

[0097] (3) Experiments and verification were conducted on the DeepShip deepwater vessel public dataset. Compared with existing deep learning and manifold learning feature extraction methods, the recognition accuracy was improved by an average of 14.96%. BRIEF DESCRIPTION OF THE DRAWINGS

[0098] Figure 1 This is a schematic diagram of the DMLAE model framework.

[0099] Figure 2 This is a schematic diagram of the autoencoder network model.

[0100] Figure 3 It is a schematic diagram of local linear embedding.

[0101] Figure 4 This is a schematic diagram of ship radiated noise.

[0102] Figure 5 This is a schematic diagram of the adversarial manifold regularized autoencoder model.

[0103] Figure 6 This is the S-curve dimensionality reduction result diagram.

[0104] Figure 7 This is a schematic diagram of the DMLAE network structure.

[0105] Figure 8 It is a schematic diagram of the activation function.

[0106] Figure 9 This is a diagram of the four types of ships.

[0107] Figure 10 It is the time domain diagram of the signal.

[0108] Figure 11 It is a schematic diagram of time domain and frequency domain characteristics.

[0109] Figure 12 This is a diagram of the confusion matrix.

[0110] Figure 13 It is a visualization diagram of time domain features.

[0111] Figure 14 It is a diagram showing the visualization of frequency domain features. DETAILED DESCRIPTION

[0112] Inspired by the complementary advantages of manifold learning and deep learning, this paper proposes a method for extracting underwater target acoustic signal features based on a deep manifold learning auto-encoder (DMLAE). The deep auto-encoder network is used to globally optimize the underwater target acoustic signal data. The idea of maintaining neighbor reconstruction weights in manifold learning is then used to locally constrain the potential representation of the data. At the same time, a generative adversarial network (GAN) architecture is introduced for regularization processing to reveal the inherent geometric structure and regularity hidden in the data, thereby realizing a deep manifold learning method that jointly maintains local and global embedding. This method obtains the intrinsic characteristics of the data through unsupervised learning, which not only solves the generalization learning and noise problems existing in manifold learning, but also enables the autoencoder to recover the potential representation that preserves the local geometric structure.

[0113] 1 Feature extraction method

[0114] The deep manifold learning autoencoder proposed in this paper is an unsupervised method for extracting underwater target acoustic signal features. It uses the advantages of neural networks in nonlinear data processing to learn the low-dimensional data features of underwater acoustic signals. And through the manifold learning local preservation embedding idea, the extracted latent representation retains the intrinsic structure of ship radiated noise, achieving the purpose of underwater target acoustic signal feature extraction. The DMLAE model consists of two parts: the generator and the discriminator. Figure 1 shown.

[0115] The generator is implemented by adding manifold regularization to an autoencoder, primarily to extract global and local features from underwater target acoustic signal data. The autoencoder network in the generator extracts global features from the data samples under unsupervised conditions. These features can characterize the overall properties of the target acoustic signal, but they fail to reflect the spatial structure of the data. To address this issue, we introduce manifold learning into the autoencoder, utilizing the nearest neighbor reconstruction matrix between the original sample points to construct local structural relationships between hidden layer samples. This allows the extracted features to also reflect the local properties of the target acoustic signal, thus achieving local feature extraction.

[0116] Due to the limitations of its network structure and functionality, the generator can only extract global and local features of underwater target acoustic signals and cannot constrain the distribution of these data features. To ensure that the features closely reflect the actual distribution of underwater target acoustic signals, the network discriminator can improve the performance of the network's latent representation by mapping the data's latent representation to a specific probability distribution space and implementing adversarial regularization.

[0117] The discriminator distinguishes the true distribution of underwater target acoustic signal samples from the distribution of potential representations generated by the network generator, and continuously updates the generator through iteration to achieve the goal of making the discriminator unable to distinguish between the two, thereby making the distribution of potential representations continuously close to the true distribution of the data, making the features extracted by the model more representative and better reflecting the intrinsic properties of the original data.

[0118] The following describes the three functions of the DMLAE model: global feature extraction, local feature extraction, and the principles of adversarial regularization processing.

[0119] 1.1 Global / local feature extraction

[0120] Due to the complex and ever-changing ocean environment, the collected datasets contain a large number of uncertain non-target signals, exhibiting high-dimensional, nonlinear, and unstructured characteristics. To ensure the integrity of the key features of underwater target acoustic signals, we directly used the signal's time and frequency domain information as input data for our experiments. The DMLAE model generator leverages its powerful self-learning, self-organizing, and adaptive capabilities to handle complex and uncertain data systems. It also exhibits strong robustness and fault tolerance, and performs well when analyzing and processing data primarily composed of noise.

[0121] First, define the underwater target acoustic signal X = [x1, x2, ...x n ]∈R D×n , where x i is the feature vector of a single underwater acoustic sample, n represents the number of samples in the input network, and D is the feature dimension of a single sample. The autoencoder network of the generator encodes the target acoustic signal data X to obtain the data representation Y of the hidden layer = [y1, y2, ...y n ]∈R d×n , where y i is the feature vector of the potential representation, and d is the feature dimension of a single sample of the potential representation. Then, by decoding the hidden layer feature Y, we can obtain the reconstructed underwater acoustic data X′, as shown in Figure 2 shown.

[0122] In the encoding and decoding process, by setting fewer hidden layer neurons, the underwater acoustic data is compressed into a potential representation Y, some unimportant features of the underwater target acoustic signal are removed, and the important features that can restore the original target acoustic signal are retained to achieve the purpose of feature extraction. The process can be expressed as

[0123] Y=f(W1X+b1) (1)

[0124] X′=h(W2Y+b2) (2)

[0125] Among them, formula (1) is the encoding network mapping function (target sound signal - intrinsic features), formula (2) is the decoding network mapping function (intrinsic features - reconstructed target sound signal), parameters W1, W2 and b1, b2 represent the weights and biases of the deep manifold autoencoder network respectively, and f and h are nonlinear activation functions. This part is a feedforward, fully connected neural network. The purpose is to make the output layer and the input layer data as close as possible, that is, the overall difference in the data is as small as possible, so that the potential representation of the hidden layer can reflect all the characteristics of the original target sound signal as realistically as possible. By minimizing the loss function for optimization, the objective function can be expressed as

[0126]

[0127] Where, L ris the reconstruction error loss function of the deep manifold autoencoder.

[0128] From the above analysis, it can be seen that the autoencoder network in the generator considers the global characteristics of the target acoustic signal data and only focuses on the reconstruction results of the output, without considering the spatial structural characteristics between the hidden layer potential representation samples. This is an inherent drawback of the autoencoder network, and the classic algorithm in manifold learning, Locally Linear Embedding (LLE), can solve this problem well. The LLE algorithm enables the manifold to be successfully embedded in low-dimensional coordinates by maintaining the local structural characteristics between the data. The local features of the data manifold are constructed through the linear combination of the data points and their nearest neighbors. LLE retains the reconstruction weights in the linear combination as much as possible. The basic steps of the algorithm are as follows: Figure 3 shown.

[0129] The idea of manifold learning to maintain neighbor weights is introduced into the hidden layer potential representation. For the underwater target acoustic signal dataset sample X = [x1, x2, ...x n ]∈R D×n , find each sample point x according to the Euclidean distance between these data points i k nearest neighbors N(x i ), the topological structure of the signal can be determined by the local neighbor relationship of each target acoustic signal sample. i ) construct a reconstruction weight vector s i , get the reconstruction weight matrix s=[s1,s2,...s n ]∈R n×n , which preserves the structural relationship between the neighbors of each sample and achieves local feature extraction of the data by keeping the reconstruction weights unchanged in the low-dimensional space. The neighbor reconstruction weights are calculated by minimizing the reconstruction error of formula (4).

[0130]

[0131] Where: Weight s ij Reflects the sample point x i x j The size of the reconstruction contribution.

[0132] Simplifying formula (4) we can get:

[0133]

[0134] in

[0135] Q i =(x i -x j )(x i -xj ) T ∈R k×k

[0136] S i =(s i1 , s i2 ,B,s ik )∈R k×1

[0137] From the above formula, we can see that for the rotation and scaling of the sample points, formula (4) can naturally keep the reconstruction weight unchanged. In order to achieve the same translation, for all x i Give constraints∑ j s ij =s i T l k =1, where l k ∈R k×1 , and then use the Lagrange multiplier method to solve it, as shown below:

[0138]

[0139] Obtain the reconstruction weight vector s i The closed-form solution of .

[0140]

[0141] The key to extracting local features of underwater target acoustic signals by DMLAE model is to require a low-dimensional potential representation y i The relationship between the original sound sample and its neighboring points can reflect the reconstruction weight relationship, that is, minimize the objective function of the following formula:

[0142]

[0143] In order to eliminate the translation, rotation and scaling factors of the low-dimensional embedding coordinate Y, the following constraints are added to Equation (8):

[0144] ①

[0145] ②

[0146] Simplifying, we can get:

[0147] φ(Y)=||Y(Is T )|| 2 =tr(YMY T ) (9)

[0148] Where I is the identity matrix, S is the reconstruction weight matrix, and M = (IS) T (IS) is an n×n matrix.

[0149] Based on the above analysis, we combine the overall optimization of the autoencoder with the local optimization of manifold learning to achieve global and local feature extraction of underwater target acoustic signals. At the same time, we add the L1 regularization term in the optimization function. L1 regularization can make the weights of some network nodes become 0. When the underwater target acoustic signal sample is input, feature selection can be performed, that is, a sparse effect is produced to prevent overfitting. The L1 expression is shown in formula (10):

[0150]

[0151] Where w i is the weight vector of each neuron in the encoding network in the model generator. According to the DMLAE model generator principle, the objective function is redefined as follows:

[0152]

[0153] where the parameters θ = (W1, b1, W2, b2), X and X′ are the original target acoustic signal and its reconstructed signal by the generator network, respectively, α is the balance parameter of the local constraint of manifold learning, and λ is the balance parameter of the L1 regularization component of the encoder network.

[0154] For solving the objective function J(θ) in the above formula, the gradient descent method is generally used to obtain the global optimal solution of the convex function. The gradient descent method (SGD, Adam, etc.) has the characteristics of simple implementation and fast convergence.

[0155] The error terms δ1 and δ2 are introduced, corresponding to the output layer and hidden layer respectively, to measure the responsibility of the same node for the error. The calculation method is as follows:

[0156] δ1=(x′-x)·σ′(z1) (12)

[0157] δ2=(W2 T ·δ T ) ·σ′(z2) (13)

[0158] Where z1 = W2y + b2, z2 = W1x + b1. Now we solve for the gradient of θ:

[0159]

[0160]

[0161]

[0162]

[0163] Then use the gradient and learning rate α to update the parameters to obtain the optimal solution of the objective function:

[0164]

[0165]

[0166] Where, the learning rate α controls the network's convergence to the local optimum and the convergence rate.

[0167] 1.2 Adversarial Regularization

[0168] The theoretical derivation of the DMLAE model combining global optimization with local constraints was previously analyzed. However, the generator only focuses on the reconstruction and representation of underwater target acoustic signals and cannot control the actual distribution of latent representation sample points. To address this issue, the network discriminator can effectively control the distribution of latent representations in the generator. The discriminator model is built based on the Generative Adversarial Network (GAN) framework. It performs feature extraction by matching the posterior probability distribution of the latent representation with the true distribution. It can also connect similar samples in the hidden layer, thereby preventing the manifold "breakage" problem commonly encountered in the embedding of deep learning networks, and better restores the embedded representation of the high-dimensional manifold structure of the underwater target acoustic signal in the low-dimensional space.

[0169] The underwater target acoustic signal data used in this invention includes ship radiated noise and ocean environmental background noise, both of which are unprocessed audio data collected and collated by hydrophones placed below the water surface. The ocean environmental background noise is unrelated to the ship radiated noise and generally appears as additive Gaussian white noise. In reality, the true distribution of ship radiated noise cannot be determined, but by analyzing the generation mechanism of underwater target acoustic signals, it can be seen that the ship radiated noise can be regarded as the superposition of line spectrum and continuous spectrum, and its amplitude distribution is Gaussian, such as Figure 4 Therefore, in this invention, we approximate the ship radiation noise as Gaussian noise, and the goal of the discriminator is to restore the true distribution (Gaussian distribution) of the potential representation in the generator as much as possible.

[0170] The principle of the discriminator algorithm is to distinguish the true distribution from the potential distribution by updating the discriminant network, and then update the DMLAE generator network to optimize the potential distribution structure. For a given single sample data x of the underwater target acoustic signal, its potential representation y is obtained through the network generator, p(y) is the true distribution of the acoustic signal, and q(y) is the aggregated posterior distribution of the data. The model framework is as follows Figure 5 shown.

[0171] When training and testing a neural network, we want each data feature of the underwater target acoustic signal to be independent and follow a known prior distribution. We introduce an adversarial approach to the hidden layer of the generator network to optimize the latent representation. The discriminator in the hidden layer compares the sampled real data (Gaussian noise) with the fake data generated by the generator (latent representation), ensuring that the aggregated posterior distribution q(y) matches the prior distribution p(y), thus regularizing the latent representation.

[0172] The adversarial regularization objective function is as follows:

[0173]

[0174] The above formula is the classic GAN network objective function formula, which can be applied to the model of the present invention. The difference is that the output of the encoder is operated on in the present invention so that the hidden layer obtains the desired feature expression result.

[0175] To verify that the model can preserve the intrinsic topological structure from the global and local aspects of the manifold, we conducted dimensionality reduction experiments on the classic manifold dataset S-curve and compared it with the LLE and AE algorithms. Figure 6 As shown in the figure, from the experimental results, the model can well unfold the S three-dimensional surface and basically restore the low-dimensional representation of the high-dimensional manifold, providing an experimental basis for the subsequent feature extraction of underwater target acoustic signals.

[0176] 1.3 Underwater target acoustic signal feature extraction algorithm

[0177] For the n time domain / frequency domain feature vectors X=[x1, x2, ...x n ]∈R D×n The deep manifold autoencoder extracts both global and local features. Global features capture the overall properties of the target acoustic signal, but the spatial geometric characteristics of the signal are lost. Local features, on the other hand, effectively preserve the signal's topological structure. The main steps of the underwater target acoustic signal feature extraction algorithm based on the deep manifold autoencoder are shown in Table 1.

[0178] Table 1 Main steps of feature extraction algorithm

[0179]

[0180] 1.4 Network Model Structure

[0181] The structural design of the neural network is crucial for the accuracy of feature extraction of underwater target acoustic signals. The number of network layers and the number of nodes in each layer should just meet the requirements. By setting different network structures and observing the experimental results, it was found that the model learning ability was strongest when the number of layers was 2, and other numbers of layers would produce underfitting and overfitting phenomena. This experimental model selected a double hidden layer network structure, the number of hidden layer network nodes was set to 1000, and the encoder output layer nodes were set to 4. The specific network structure is as follows Figure 7 shown.

[0182] The present invention uses ReLU and Sigmoid functions as nonlinear activation functions of neural networks. The output range of the Sigmoid function is 0 to 1, which makes the output value have a probabilistic interpretation and is suitable for tasks that require the output to be interpreted as probability or probability-like. Figure 8 The mathematical expression of (a) is as follows:

[0183]

[0184] When the input value is less than or equal to 0, the output of the ReLU function is 0, which makes some neurons in the network become inactive, thus achieving sparse representation, that is, only some neurons respond to the input. Figure 8 The numerical expression of (b) is as follows:

[0185]

[0186] The DMLAE model mainly uses the error loss function and the cross entropy loss function. The mathematical expressions are as follows:

[0187]

[0188] L bce =-(x logp(x′)+(1-x)log(1-p(x′))) (24)

[0189] Wherein, formula (23) is the mean square error loss, which is generally applicable to regression tasks. Formula (24) is the binary cross entropy loss, and p(x′) is the predicted probability, which is generally applicable to classification tasks. In the feature extraction method of the present invention, x is the original underwater target acoustic signal sample, and x′ is the reconstructed signal sample.

[0190] Adam was selected as the model optimizer. By adaptively adjusting the learning rate and using momentum, it helps the neural network converge faster and improves training performance. It excels at handling diverse parameter updates, controlling update amplitudes, and adapting to sparse gradients. It is widely used in various deep learning tasks and network structures. Table 2 shows the network parameter settings.

[0191] Table 2 Network parameter description

[0192]

[0193] 2 Experimental studies

[0194] 2.1 Dataset Analysis and Preprocessing

[0195] The dataset used in this paper comes from the Deepship deepwater vessel dataset published by Irfan et al. in 2021. The dataset consists of 265 real deepwater vessel underwater sounds, which are divided into four categories, including tankers, tugboats, passenger ships and cargo ships. Figure 9 As shown, Figure 9 (a) is a tugboat, (b) is a cargo ship, (c) is an oil tanker, and (d) is a passenger ship. The recording device is a hydrophone, which records and saves data at a sampling rate of 32kHz, lasting 47 hours and 4 minutes. This dataset includes records of various sea conditions and noise levels throughout the year. In addition to deepwater vessel data, it also records ocean background noise and noise generated by other human activities. The ocean background noise is added to the deepwater vessel dataset, lasting 7 hours and 52 minutes.

[0196] The original data set is an audio file in wav format. First, each audio file is resampled and output at a sampling frequency of 16 kHz. Pre-emphasis processing can enhance high-frequency content, improve the dynamic range of the signal, reduce the impact of noise and distortion, and provide better input features for subsequent processing algorithms, thereby improving signal quality and processing effects. The present invention uses a first-order FIR filter to perform pre-emphasis processing on the signal. The filter transfer function is as follows: H(z) = 1-az (25)

[0197] Where a is the pre-emphasis coefficient, which generally ranges from 0.9 to 1.0. In this experiment, it is set to 0.98.

[0198] Then, the pre-emphasized underwater acoustic signal is framed, so that the underwater target acoustic signal can be analyzed and processed as a steady signal within a short time range. In order to make the transition between frames smooth and maintain its continuity, the overlapping segmentation method is generally adopted. The present invention selects a frame length of 640 (lasting 0.04 seconds) and a frame shift of 320 (lasting 0.02 seconds). After the signal is framed, the frequency domain energy leaks due to truncation, and the window function can reduce the impact of truncation, keep the connection between each frame relatively smooth and continuous, and eliminate the signal discontinuity that may be caused at both ends of each frame. The present invention uses a Hamming window to process the frame signal. The mathematical expression of the window function is as follows:

[0199]

[0200] Analyze the characteristics of underwater target acoustic signals, and the time domain characteristics of the signals at each processing stage are as follows: Figure 10 As shown, Figure 10 (a) is the original signal, (b) is the pre-emphasized signal, (c) is the framed signal, and (d) is the windowed signal.

[0201] This experiment only extracts features from the time domain and frequency domain of underwater target sound signals. Considering the excellent performance of deep learning in processing images, the one-dimensional features of the signal in the time domain and frequency domain are converted into two-dimensional features, that is, input into the model in the form of grayscale images, as follows Figure 11 As shown, Figure 11 The left side of (a) and (b) in the figure is the original time and frequency domain data, and the right side is the corresponding two-dimensional image data.

[0202] The dataset is divided into five categories, namely four types of ships and ocean noise. To ensure fairness in model training, an equal number of samples are set for each type of data, and the data is divided into 80% training samples and 20% test samples. The specific division is shown in Table 3 below.

[0203] Table 3 Data Description

[0204]

[0205] 2.2 Dataset Analysis and Preprocessing

[0206] To verify the performance of the algorithm presented in this paper, a support vector machine (SVM) was used to evaluate the feature extraction results of underwater target acoustic signals. SVM generally solves data classification problems. When processing the results, it is difficult to accurately evaluate the effectiveness of the algorithm using only one evaluation metric. Four metrics are commonly used: Precision, Recall, F1-score, and Accuracy. The mathematical expressions are as follows:

[0207]

[0208]

[0209]

[0210]

[0211] TP represents the number of samples correctly identified as positive, TN represents the number of samples correctly identified as negative, FP represents the number of samples incorrectly identified as positive, and TN represents the number of samples incorrectly identified as negative. Higher values for these four evaluation metrics indicate better and more stable model performance.

[0212] 2.3 Experiment and Analysis

[0213] First, the proposed algorithm is used to extract features from the unprocessed time domain signal, and compared with the traditional acoustic signal feature extraction method. The classification accuracy is shown in Table 4 below.

[0214] Table 4 Comparative experimental analysis

[0215]

[0216] The experimental results show that the classification accuracy of underwater target acoustic signal feature extraction using the DMLAE algorithm is about 18.48% higher than that of other traditional methods. The difference is only 1.83% compared to the accuracy of identification using all features, indicating that the model can basically use the features with the minimum dimension (4 dimensions) to represent all features (640 dimensions) of the data sample. Figure 12 Shown is the confusion matrix of the experimental results.

[0217] from Figure 12 It can be seen that the model incorrectly predicts most tugboat samples as passenger ships. The experimental results show that the recall rate of tugboats is only 0.04%, and the F1 score is 0.07, resulting in a low recognition rate. In order to more intuitively reflect the distribution of various sample points in low-dimensional space, the t-distributed Stochastic Neighbor Embedding (t-SNE) algorithm is used to visualize the features extracted by the model. It is obvious that there is a serious aliasing phenomenon between cargo ships and passenger ships, while the classification effect of the other two types of ships is better, and the ocean noise signal can be basically completely separated, as shown below. Figure 13 shown.

[0218] The above experiments demonstrate the limitations of the DMLAE model for extracting time-domain signal features. This may be due to the smaller size of the tugboat and passenger ship compared to the other two types of vessels, making it difficult to separate them using the time-domain signal features. Next, we explore the model's performance and target signal classification accuracy from a frequency-domain perspective. First, the frequency-domain information of underwater target acoustic signals is input into the model, and the output features are classified and identified. The network's recognition performance is then analyzed using evaluation metrics, as shown in Table 5.

[0219] Table 5 Network performance analysis

[0220]

[0221] As shown in Table 5, the DMLAE model has the highest recognition accuracy for ocean noise signals, with all three evaluation indicators being 1.00 and the recognition rate reaching 100%. The recall rate for cargo ships reached 1.00, indicating that the model has achieved full coverage for cargo ship recognition, that is, all samples can be accurately identified. The accuracy of oil tankers is 0.99, indicating that samples in this category can basically be correctly identified. The two-dimensional visualization results are shown in Figure 5. Figure 14 The DMLAE model achieved an average accuracy of 96.89%, which is much higher than the model's time-domain feature extraction and demonstrates the significant superiority of deep learning over traditional feature extraction methods.

[0222] The above experiments show that using the frequency domain information of underwater target acoustic signals as network input improves the classification and recognition rate by 25.34%, demonstrating that the proposed algorithm clearly distinguishes these four types of ship and ocean background noise signals in the frequency domain. Because the DMLAE algorithm is based on the combination of manifold learning and deep learning, it was compared with other manifold learning and deep learning algorithms. To ensure experimental fairness, the comparative experiments were conducted under identical experimental data, network structure, parameter settings, and evaluation indicators. The classification accuracy of each algorithm is shown in Table 6 below.

[0223] Table 6 Comparative experimental analysis

[0224]

[0225] From the analysis in Table 6, it can be seen that DMLAE has the highest recognition rate compared with other methods, which is 14.96% higher on average. It is only 1.04% lower than the accuracy of recognition using all features, which can better reflect the intrinsic characteristics of the data. Among them, the convolutional autoencoder (CAE) is second only to DMLAE, with a recognition rate of 93.52%, indicating that convolutional neural networks have excellent performance for underwater target sound signals. EAER is also an autoencoder technology based on manifold regularization, with a recognition rate of 91.66%, indicating that the manifold regularization method has certain advantages in signal feature extraction. However, the manifold learning method LLE performed poorly. The data is ship radiation noise and ocean background noise data. Manifold learning has certain drawbacks in processing noise data, which results in the extracted features not being able to well represent its intrinsic nature.

[0226] 3 Conclusion

[0227] To address the problem of extracting features from underwater target acoustic signals, a Deep Manifold Autoencoder (DMLAE) technique was proposed. Experiments were conducted using a real-world Deepship dataset and ocean background noise. By introducing local constraints for manifold learning and L1 regularization terms to reconstruct the generator objective function, the hidden layer features were able to maintain the inherent structure of the data. The GAN network architecture was then introduced into the network model, and the discriminator performed adversarial regularization on the objective function, ensuring that the sample data followed a Gaussian distribution. Experimental verification demonstrated that the algorithm achieved a classification recognition rate of 96.89%, an average recognition rate improvement of 14.96% compared to other deep learning feature extraction algorithms, demonstrating the effectiveness of DMLAE for extracting features from underwater target acoustic signals.

Claims

1. A method for extracting underwater target acoustic signal features based on deep manifold learning, characterized in that: The following steps are involved: Step 1: Build a deep manifold autoencoder model and name it the DMLAE model; the DMLAE model includes a generator and a discriminator; initialize the DMLAE model parameters; Wherein, the generator is constructed by adding manifold regularization to the autoencoder; The autoencoder in the generator is used to extract global features of underwater target acoustic signal data samples under unsupervised conditions, where the global features represent the overall properties of the underwater target acoustic signal; The generator is further configured to construct a local structural relationship between hidden layer samples using a neighbor reconstruction matrix between original sample points to obtain local features, which reflect local properties of the underwater target acoustic signal; The discriminator is used to discriminate between the true distribution of underwater target acoustic signal samples and the potential representation distribution generated by the generator, and continuously update the generator through iteration; Step 2: Input the underwater target acoustic signal data sample matrix X into the generator input layer, and minimize the reconstruction error loss function L of the autoencoder network of the autoencoder r Optimize, that is , the global feature Y of the target acoustic signal is obtained in the hidden layer, where X′ is the reconstructed underwater acoustic data; Step 3: Calculate the nearest neighbor reconstruction weight matrix S of the target acoustic signal sample point; Step 4: Extract local features of underwater target acoustic signals and add local constraints to the global feature Y; the details are as follows: The DMLAE model extracts the local features of underwater target acoustic signals and minimizes formula (8) so that y i with y i The neighboring points of reflect the reconstruction weight relationship of the original underwater target sound signal data samples, ; Add the following constraints to equation (8): ① ; ② ; Simplified to: φ(Y) = ||Y(I - S T )|| 2 = tr(YMY T )(9) Where I is the identity matrix, S is the neighbor reconstruction weight matrix, tr() is the function for finding the matrix trace, and M = (IS) T (IS) is an n×n matrix; Step 5: Add the L1 regularization term and rebuild the objective function of the DMLAE model; the details are as follows: The L1 expression is shown in formula (10): (10) Where W1 is the encoder weight, is the weight vector of each neuron in the encoding network in the model generator; According to the DMLAE model generator principle, the objective function J(θ) of the DMLAE model is redefined as shown in formula (11): ; Where α is the balance parameter for local constraints in manifold learning, and λ is the balance parameter for the L1 regularization component of the encoder network; For the solution of the objective function J(θ) in the above formula, the gradient descent method is used to obtain the global optimal solution of the convex function; Step 6: Perform adversarial regularization on the global feature Y; Step 7: Iteratively update the weights and biases of the DMLAE model, and obtain the optimal parameters to complete the DMLAE model training; Step 8: Input the underwater target acoustic signal data into the trained DMLAE model, implement data feature extraction in the generator, and output the intrinsic feature matrix of the underwater target acoustic signal.

2. The underwater target acoustic signal feature extraction method based on deep manifold learning according to claim 1 is characterized in that: In step 1, the model parameters are initialized including: The number of samples n in the underwater target acoustic signal data sample matrix X; The feature dimension D of a single sample; Output feature dimension q; Maximum number of iterations iter; The balance parameter α of the local constraint of manifold learning; The balance parameter λ of L1 regularization; The neighbor coefficient k of the neighbor reconstruction weight matrix.

3. The underwater target acoustic signal feature extraction method based on deep manifold learning according to claim 2 is characterized in that: In step 2, Define the underwater target acoustic signal data sample matrix X=[x1,x2,...x n ]∈R D×n , where x i is the feature vector of a single underwater acoustic sample, n represents the number of samples in the input network, and D is the feature dimension of a single sample; The generator's autoencoder network encodes the underwater target acoustic signal data sample matrix X to obtain the target acoustic signal global feature Y=[y1,y2,...y n ]∈R d×n , where y i is the feature vector of the global feature Y, and d is the feature dimension of a single sample in the global feature Y; By decoding Y in the hidden layer, the reconstructed underwater acoustic data X′ is obtained; Y=f(W1X+b1)(1) X′=h(W2Y+b2)(2) Formula (1) is the encoding network mapping function; Formula (2) is the decoding network mapping function; Parameters W1, W2 and b1, b2 represent the weight and bias of the network of the deep DMLAE model, respectively, and f and h are nonlinear activation functions; By minimizing the reconstruction error loss function L of the autoencoder network of the DMLAE model r Optimize, L r Expressed as 。 4. The underwater target acoustic signal feature extraction method based on deep manifold learning according to claim 3 is characterized in that: In step 3, the neighbor reconstruction weight matrix of the target acoustic signal sample point is calculated as follows: The idea of manifold learning to maintain neighbor weights is introduced into the hidden layer potential representation. For the underwater target acoustic signal dataset, the sample matrix X = [x1, x2, ... x n ]∈R D×n , find each sample point x according to the Euclidean distance between sample points i0 The set of k nearest neighbor points N(x i0 ), the topological structure of the underwater target acoustic signal data is determined by the local neighbor relationship of each underwater target acoustic signal data sample, and k is the neighbor coefficient; By the neighboring point N(x i0 ) to construct a reconstruction weight vector S i0 , get the neighbor reconstruction weight matrix S=[S1,S2,...S n ]∈R n×n ; The neighbor reconstruction weight is calculated by minimizing the reconstruction error of formula (4): (4) Where, the neighbor reconstruction weight s i0j0 Reflects the sample point x i0 x j0 The size of the reconstruction contribution; Simplifying formula (4) we can get: (5) in, Q i0 =(x i0 -x j0 )(x i0 -x j0 ) T ∈R k×k S i0 =(s i1 ,s i2 ,…,s ik )∈R k×1 For the rotation and scaling of sample points, Equation (4) keeps the neighbor reconstruction weights unchanged; For all x i0 Give constraints∑ j0 s i0j0 =S i0 T l k =1, so that the neighbor reconstruction weight remains unchanged even when the sample point is shifted; where l k ∈R k×1 is a unit vector, and the Lagrange multiplier method is used to solve it. ξ is the Lagrange multiplier as shown below: ; Obtain the nearest neighbor reconstruction weight vector S i0 The closed-form solution of : (7)。 5. The underwater target acoustic signal feature extraction method based on deep manifold learning according to claim 4 is characterized in that: In step 7, the weights and biases W1, b1, W2, and b2 of the DMLAE model are iteratively updated, and the optimal parameters θ = (W1, b1, W2, and b2) are obtained as follows: The error terms δ1 and δ2 are introduced, corresponding to the output layer and hidden layer respectively, to measure the responsibility of the same node for the error. The calculation method is as follows: δ1=(x′-x)·σ′(z1)(12) δ2=(W2 T ·d1 T )·σ′(z2)(13) Among them, z1=W2y+b2, z2=W1x+b1, σ′ function is the derivative of the activation function; Solve for the gradient of θ: (14) (15) ; ; Then use the gradient and learning rate α' to update the parameters to obtain the optimal solution of the objective function: ; ; Where, the learning rate α' controls the network's convergence to the local optimum and the convergence rate; W 1,2 Represents W1 and W2, b 1,2 Represents b1 and b2.

6. The underwater target acoustic signal feature extraction method based on deep manifold learning according to claim 5, characterized in that: In step 6, the global feature Y is subjected to adversarial regularization as follows: The ship radiated noise is approximated as Gaussian noise. For a given single sample data x of the underwater target acoustic signal, the generator obtains y in the global feature Y, where p(y) is the true distribution of the acoustic signal and q(y) is the aggregated posterior distribution of the data. When using neural networks for training and testing, an adversarial approach is introduced in the hidden layer of the generator network to optimize global features. The discriminator in the hidden layer judges the real data after sampling and the fake data generated by the generator, so that the aggregated posterior distribution q(y) matches the prior distribution p(y) to complete the regularization of global features. The adversarial regularization objective function is as follows: ; The above formula is the objective function formula of the GAN network, where D(x) is the discriminator output and G(y) is the generator output, so that the hidden layer obtains the desired feature expression result.

7. The method for extracting underwater target acoustic signal features based on deep manifold learning according to claim 6, characterized in that: The DMLAE model selects a double hidden layer network structure, the number of hidden layer network nodes is set to 1000, and the number of encoder output layer nodes is set to 4; Use ReLU and Sigmoid functions as nonlinear activation functions of the network; The DMLAE model uses the mean squared error loss function and the binary cross entropy loss function. The mean squared error loss function is used for regression tasks, and the binary cross entropy loss is used for classification tasks.