CNN-LSTM model-based three-node passive acoustic array ship positioning method
Through the three-node passive acoustic array method based on the CNN-LSTM model, the inter-aural intensity difference and phase difference characteristics are used to locate the sound source, which solves the problem of high-precision positioning in complex underwater environments and realizes low-cost and efficient ship target positioning.
Patent Information
- Application Number
- CN202511333011.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-18
- Publication Date
- 2025-10-24
- Estimated Expiration
- 2045-09-18
AI Technical Summary
In a complex underwater acoustic environment, how to achieve high-precision ship target positioning under limited observation nodes and complex propagation conditions, especially when using passive acoustic networks to locate ship radiated noise, existing methods have problems such as insufficient positioning accuracy and excessive dependence on the environment.
A three-node passive acoustic array method based on the CNN-LSTM model is adopted. By collecting acoustic signals in real time, the inter-aural intensity difference and phase difference features are extracted. The spatial and temporal features are extracted by combining the CNN-LSTM model, and the azimuth and distance values of the sound source are output. This simplifies the dependence on the underwater acoustic environment and enhances the generalization ability of the model in complex environments.
It achieves high-precision positioning of ship radiated noise in complex underwater environments, reduces system deployment and maintenance costs, is suitable for distributed networks, can output azimuth and distance simultaneously, and improves the robustness and accuracy of the positioning system.
Smart Images

Figure CN120831631A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of underwater sound source positioning, in particular to a three-node passive acoustic array ship positioning method based on a CNN-LSTM model. BACKGROUND
[0002] Underwater sound source positioning technology is the basis for many marine application fields such as ocean exploration, national defense security, resource exploration and marine ecological monitoring. The underwater acoustic positioning technology can be generally divided into active positioning and passive positioning. In military reconnaissance and marine environment monitoring, low detectability, accurate and robust positioning of the target sound source is crucial for real-time tracking, environmental assessment and rapid emergency response. However, the underwater acoustic environment is extremely complex and highly dynamic, making the robust positioning of underwater sound sources a highly challenging task. In particular, when using a passive acoustic network (distributed hydrophone network) to locate the radiation noise of targets such as ships, it is a technical problem to be solved in the current underwater acoustic field to obtain high-precision azimuth and distance information under the condition of limited observation nodes and complex propagation conditions. SUMMARY
[0003] In view of the above prior art, the present application provides a three-node passive acoustic array ship positioning method based on a CNN-LSTM model, which realizes ship targets in a complex underwater acoustic environment.
[0004] To achieve the above purpose, the technical scheme of the embodiment of the present application is as follows: A three-node passive acoustic array ship positioning method based on a CNN-LSTM model, the method comprising the following steps: Step S1: Real-time acquisition of acoustic signals radiated by a target ship by deploying a three-node passive acoustic network under water; and pre-processing the acoustic signals of each node to obtain interaural intensity difference features and interaural phase difference features; Step S2: Constructing a CNN-LSTM model, inputting the combined feature map of the interaural intensity difference features and the interaural phase difference features into the spatial feature extraction module of the CNN-LSTM model for spatial feature extraction, and then outputting a spatial feature vector; Step S3: Inputting the spatial feature vector into the time sequence feature learning module of the CNN-LSTM model for time sequence feature extraction, and then outputting a latent feature vector; Step S4: Inputting the latent feature vector into the output layer of the CNN-LSTM model for regression prediction, and outputting the azimuth angle and distance value of the sound source.
[0005] As a preferred scheme of the present application, the pre-processing of the acoustic signals of each node in step S1 specifically comprises: cutting the original acoustic signals continuously collected by each node into fixed-length analysis frames and performing down-sampling; performing short-time Fourier transform on the down-sampled acoustic signals to obtain the time-frequency spectrum of the signals of each node, and then calculating the interaural intensity difference feature and the interaural phase difference feature; calculating the interaural intensity difference feature The formula of the interaural intensity difference feature
[0006] The formula of the interaural phase difference feature The formula of the interaural phase difference feature
[0007] wherein, and are the STFT complex matrices of the signals received by the two sensors representing the frequency, representing the STFT time frame index.
[0008] As a preferred scheme of the present application, the extracted interaural intensity difference feature and the interaural phase difference feature are respectively normalized to have zero mean and unit variance before being input into the CNN-LSTM model in step S2, and the specific formula is as follows:
[0009]
[0010]
[0011] wherein, represents the mean of a matrix of a sample, represents the standard deviation of the matrix, represents the value of the row and the column in the matrix, represents the number of rows of the matrix, represents the number of columns of the matrix, represents the normalized sample matrix.
[0012] As a preferred embodiment of the present invention, the spatial feature extraction module in step S2 is composed of two branches, each branch including a shallow convolution block, a middle convolution block and a deep convolution block. The shallow convolution block is used to preliminarily extract the frequency and intensity / phase difference information of the feature map of the combination of the interaural intensity difference feature and the interaural phase difference feature, and then the middle convolution block is used to further extract the deep features, and a batch normalization layer is added to the middle convolution block to enhance the model expression ability. Then, the deep convolution block is used to perform feature association, channel dimensionality reduction and feature reorganization, and finally the Tanh activation function is used to output the extracted spatial feature vector; In order to avoid the information loss problem caused by the pooling layer, the step size of the convolution layer is set in the entire spatial feature extraction module to control the downsampling of the feature map, without using the pooling layer. In order to speed up the convergence of the model and enhance the nonlinear expression ability of the model, the ELU activation function is used to introduce the non-saturation characteristics of the negative value interval before the entire convolution branch outputs the feature map, so that the output data maintains a certain gradient on the negative half axis. The calculation formula of the ELU activation function is as follows:
[0013] in, is the output of the activation function, is the input of the activation function, is the negative half-axis saturation coefficient.
[0014] As a preferred embodiment of the present invention, step S2 also includes dynamically masking the feature map of the extracted combination of interaural intensity difference features and interaural phase difference features, including randomly masking some continuous areas in the time domain and some frequency bands in the frequency domain, and jointly using time domain noise addition and feature masking during model training.
[0015] As a preferred solution of the present invention, in step S3, the spatial feature vectors output by each branch of the spatial feature extraction module are transformed into a one-dimensional feature vector, and then considering the temporal information of the two branches, the two one-dimensional feature vectors are spliced to form a temporal feature vector containing two frames, and the temporal feature vector is input into the subsequent temporal feature learning module; The temporal feature learning module consists of a two-layer LSTM network. The hidden layer state of the first LSTM layer is output to the input layer of the second LSTM layer. The LSTM network effectively learns and utilizes the long-range temporal dependencies in the feature sequence through its unique gating mechanism. To prevent overfitting, both layers of the LSTM network are subjected to L2 regularization constraints and output the final potential feature vector.
[0016] As a preferred scheme of the present application, the regression prediction in the two full connection linear layers of the CNN-LSTM model in the step S4 outputs the azimuth and distance values of the sound source, specifically including: the optimization target of the loss function of the CNN-LSTM model is to minimize the sum of the loss values between the predicted azimuth, the predicted distance and the true values, and the loss function is defined as follows:
[0017]
[0018]
[0019] wherein, is the total loss function, is the loss function of the predicted azimuth, is the loss function of the predicted distance, represents the number of samples, is the serial number of the sample, and respectively represent the true value and the predicted value of the azimuth, represents the normalized true distance value, represents the normalized distance value predicted by the model, and the distance normalization adopts the minimum-maximum scaling, and the formula is as follows:
[0020] wherein, and respectively represent the minimum value and the maximum value of the distance in the complete data set.
[0021] The present application has the following beneficial effects: (1) The present method uses the ILD+IPD combined features, combines the powerful spatial feature extraction capability of the CNN and the modeling advantage of the LSTM on the time sequence information, can learn the complex space-time pattern from the input features, directly learns the deep features related to the sound source position from the preprocessed ILD+IPD features, avoids the complex artificial feature design and selection process in the traditional method, and reduces the dependence on the accurate sound speed profile and other underwater acoustic environment prior knowledge, simplifies the complexity of the positioning system. At the same time, the data enhancement strategies of time domain noise superposition and time-frequency domain feature mask are adopted, which significantly improves the generalization ability and robustness of the model in the complex underwater noise and signal distortion environment, so as to realize the synchronous regression prediction of the azimuth and distance of the ship radiated noise.
[0022] (2) The method is suitable for a distributed network, does not need to rely on a large number of carefully designed dense arrays as a sensor collection network, has the characteristics of low cost and easy deployment, and effectively solves the problem of performance decline in a distributed network and the problem that existing deep learning methods generally rely on dense arrays through a three-node passive acoustic network design. Effective positioning can be achieved using a small number of hydrophones, effectively reducing the cost and complexity of system deployment and maintenance, and being easy to popularize and apply in practice.
[0023] (3) The CNN-LSTM model used in the method can output two key positioning parameters, the azimuth angle and the distance of the sound source, at the same time, providing more comprehensive target position information, which is superior to the traditional method which can only estimate the azimuth angle. BRIEF DESCRIPTION OF DRAWINGS
[0024] Figure 1 A step schematic diagram of the three-node passive acoustic array ship positioning method based on the CNN-LSTM model provided by the present application is provided. Figure 2 A CNN-LSTM model structure schematic diagram provided by the present application is provided. Figure 3 A passive acoustic node arrangement and target trajectory schematic diagram provided by the present application is provided. Figure 4 An azimuth angle error distribution diagram provided by the present application is provided. Figure 5 A distance error distribution diagram provided by the present application is provided. Figure 6 A CNN-LSTM model training stage loss curve schematic diagram provided by the present application is provided. Figure 7 A CNN-LSTM model verification stage loss curve schematic diagram provided by the present application is provided. DETAILED DESCRIPTION
[0025] The technical solutions of the present application are further described in detail below in combination with the accompanying drawings and specific embodiments. Unless otherwise defined, all technical and scientific terms used herein have the same meanings as understood by those skilled in the art to which the present application belongs. The terms used in the specification of the present application herein are only for the purpose of describing specific embodiments and are not intended to limit the present application. In the following description, the expression "some embodiments" describes a subset of all possible embodiments, but it should be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments and can be combined with each other without conflict.
[0026] In the following description, numerous specific details are provided to provide a more thorough understanding of the present invention. However, it will be apparent to those skilled in the art that the present invention may be practiced without one or more of these details. In other instances, certain technical features well known in the art are not described to avoid confusion with the present invention.
[0027] It should be understood that the present invention can be implemented in different forms and should not be interpreted as being limited to the embodiments proposed herein. On the contrary, providing these embodiments will make the disclosure thorough and complete, and will fully convey the scope of the present invention to those skilled in the art. And the purpose of the terms used herein is only to describe specific embodiments and is not intended to limit the present invention. When used herein, the singular forms "one", "an" and "said / the" are also intended to include plural forms, unless the context clearly indicates another way. It should also be understood that the terms "comprising" and / or "comprising" when used in this specification determine the presence of the features, integers, steps, operations, elements and / or parts, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, parts and / or groups. When used herein, the term "and / or" includes any and all combinations of the relevant listed items.
[0028] It should also be noted that when an element is referred to as being "fixed to" another element, it may be directly on the other element or there may be an intermediate element. When an element is referred to as being "connected to" another element, it may be directly connected to the other element or there may be an intermediate element. The terms "vertical," "horizontal," "inner," "outer," "left," "right," and similar expressions used herein are for illustrative purposes only and do not represent the only implementation methods.
[0029] In order to fully understand the present invention, a detailed structure will be provided in the following description to illustrate the technical solution proposed by the present invention. Optional embodiments of the present invention are described in detail below. However, in addition to these detailed descriptions, the present invention may also have other implementations.
[0030] Please refer to the attached Figure 1 and attached Figure 2 The present application provides a three-node passive acoustic array ship positioning method based on a CNN-LSTM model, the method comprising the following steps: Step S1: The acoustic signal radiated by the target ship is collected in real time through a three-node passive acoustic network deployed underwater; and the acoustic signal of each node is preprocessed to obtain the interaural intensity difference feature and the interaural phase difference feature; Step S2: constructing the CNN-LSTM model, inputting the feature map of the combined ear-to-ear intensity difference feature and ear-to-ear phase difference feature into a spatial feature extraction module of the CNN-LSTM model to perform spatial feature extraction, and then outputting a spatial feature vector; Step S3: inputting the spatial feature vector into a time sequence feature learning module of the CNN-LSTM model to perform time sequence feature extraction, and then outputting a latent feature vector; Step S4: inputting the latent feature vector into an output layer of the CNN-LSTM model to perform regression prediction, and outputting the azimuth angle and distance value of the sound source.
[0031] As a preferred scheme of the present application, the pre-processing of the acoustic signal of each node in step S1 specifically comprises: cutting the original acoustic signal continuously collected by each node into fixed-length analysis frames and performing down-sampling; performing short-time Fourier transform on the down-sampled acoustic signal to obtain a time-frequency spectrum of each node signal, and calculating the ear-to-ear intensity difference feature and the ear-to-ear phase difference feature; calculating the ear-to-ear intensity difference feature , which reflects the energy difference of the signal arriving at two sensors, and its formula is as follows:
[0032] calculating the ear-to-ear phase difference feature , which reflects the phase difference of the signal arriving at two sensors, and its formula is as follows:
[0033] wherein, and are the STFT complex matrices of the signals received by the two sensors representing frequency, representing the STFT time frame index.
[0034] In the present embodiment, each hydrophone node records waveform data of sound pressure changing over time, and the sampling rate needs to be greater than 2 kHz.
[0035] In this embodiment, the acoustic signal from each node is preprocessed. First, the raw acoustic signal continuously collected by each node is divided into analysis frames of fixed duration (fixed duration of 0.5 seconds). The audio is then downsampled to 2kHz to reduce redundant information and focus on the frequency band where ship noise primarily resides. A short-time Fourier transform (STFT) is then performed on the downsampled audio data to obtain the time-frequency spectrum of each node signal. The STFT parameters are selected with a sliding window of 32 milliseconds and a step size of 16 milliseconds. Interaural intensity difference (ILD) and interaural phase difference (IPD) are calculated based on the signals received by different hydrophone pairs in the three-node network. For the three nodes (a, b, c), after selecting a reference hydrophone, two hydrophone pairs can be formed (for example, if hydrophone b is selected as the reference hydrophone, the combinations ab and cb are formed). As a preferred embodiment of the present invention, in step S2, before being input into the CNN-LSTM model, the extracted interaural intensity difference feature and interaural phase difference feature are respectively standardized so that they have zero mean and unit variance. The specific formula is as follows:
[0036]
[0037]
[0038] in, represents the mean of a sample matrix, represents the standard deviation of the matrix, Indicates the first Rank The value of the column, represents the number of rows of the matrix, represents the number of columns of the matrix, Represents the standardized sample matrix.
[0039] As a preferred embodiment of the present invention, the spatial feature extraction module in step S2 is composed of two branches, each branch including a shallow convolution block, a middle convolution block and a deep convolution block. The shallow convolution block is used to preliminarily extract the frequency and intensity / phase difference information of the feature map of the combination of the interaural intensity difference feature and the interaural phase difference feature, and then the middle convolution block is used to further extract the deep features, and a batch normalization layer is added to the middle convolution block to enhance the model expression ability. Then, the deep convolution block is used to perform feature association, channel dimensionality reduction and feature reorganization, and finally the Tanh activation function is used to output the extracted spatial feature vector; In order to avoid the information loss problem caused by the pooling layer, the step of the convolution layer is set to control the down-sampling of the feature map in the whole spatial feature extraction module, and the pooling layer is not used, and in order to accelerate the convergence speed of the model and enhance the nonlinear expression ability of the model, the ELU activation function is used before the output feature map of the whole convolution branch, the non-saturation characteristic of the negative value interval is introduced, the output data maintains a certain gradient on the negative half axis, and the calculation formula of the ELU activation function is as follows:
[0040] wherein, is the output of the activation function, is the input of the activation function, is the saturation coefficient of the negative half axis, and is usually 1.
[0041] In the embodiment, the spatial feature extraction module adopts a multi-layer two-dimensional convolutional neural network (Conv2D) structure, wherein shallow convolutional blocks (1-3 layers): a larger convolution kernel (convolution kernel size 5x5) is used to obtain a larger receptive field and preliminarily extract frequency and intensity / phase difference information. Medium convolutional blocks (4-5 layers): a smaller convolution kernel (convolution kernel size 3x3) is used to further finely extract deep features. A batch normalization layer (Batch Normalization, BN) is added after the convolution layer to enhance the expression ability of the model. Deep convolutional blocks (6-9 layers): a larger convolution kernel (convolution kernel size 9x9) is used for feature association, and then 1x1 convolution is used for channel dimension reduction and feature reorganization. The last layer of the deep convolutional block also uses 1x1 convolution to further reorganize the features and output the extracted spatial feature vector using Tanh as the activation function.
[0042] In the embodiment, the step of the convolution layer is set to control the down-sampling of the feature map in the whole spatial feature extraction module, and the pooling layer is not used. Because the pooling layer does not have a learnable parameter, simply using the pooling layer for down-sampling may cause the loss of unwanted spatial information, which is not conducive to the task of accurately positioning the fine spatial clues in ILD and IPD. Therefore, by setting the step of the convolution layer for down-sampling, the size of the feature map can be reduced while the network learns how to retain the most important spatial features, thereby avoiding the information loss problem that may be caused by the pooling layer.
[0043] In the embodiment, the ELU activation function is used before the output feature map of the whole convolution branch, which is to introduce the non-saturation characteristic of the negative value interval, so that the output data maintains a certain gradient on the negative half axis, enhances the nonlinear expression ability of the model, and accelerates the convergence speed. In addition, the ELU activation function can also alleviate the problem of overfitting to a certain extent and improve the generalization ability of the model.
[0044] As a preferred scheme of the present application, the step S2 further comprises dynamic mask enhancement on the extracted feature map of the combined inter-aural intensity difference feature and inter-aural phase difference feature, including random masking of part of the continuous area in the time domain and part of the frequency band in the frequency domain. Time domain noise addition and feature mask are used jointly during training.
[0045] As a preferred scheme of the present application, in the step S3, the spatial feature vectors output by each branch of the spatial feature extraction module are converted into one-dimensional feature vectors through deformation operation, and then the two one-dimensional feature vectors are spliced to form a time sequence feature vector containing two frames, and the time sequence feature vector is input into the subsequent time sequence feature learning module. The time sequence feature learning module is composed of two layers of LSTM networks, and the hidden layer state output of the first layer of LSTM is output to the input layer of the second layer of LSTM. The LSTM network effectively learns and utilizes the long-range time dependence in the feature sequence through its unique gating mechanism (input gate, forget gate, output gate). In order to prevent overfitting, L2 regularization constraint is performed on both layers of LSTM networks, and finally the latent feature vector is output.
[0046] As a preferred scheme of the present application, in the step S4, regression prediction is performed through two fully connected linear layers of the CNN-LSTM model, and the output of the azimuth angle and distance value of the sound source specifically comprises: the optimization target of the loss function of the CNN-LSTM model is to minimize the sum of the loss values between the predicted azimuth angle, the predicted distance and the true value, and the loss function is defined as follows:
[0047]
[0048]
[0049] wherein, is the total loss function, is the loss function of the predicted azimuth angle, is the loss function of the predicted distance, represents the number of samples, is the serial number of the sample, and respectively represent the true value and the predicted value of the azimuth angle, represents the normalized true distance value, represents the normalized distance value predicted by the model, and the distance normalization adopts the minimum-maximum scaling, and the formula is as follows:
[0050] wherein, and respectively represent the minimum and maximum values of the distances in the complete dataset.
[0051] In this embodiment, the latent feature vector output by the timing feature learning module is fed into two fully connected linear layers, which directly regress the azimuth angle and distance value of the sound source. The output of the model azimuth angle is the predicted actual angle, and the output of the distance is the predicted actual distance after normalization, the interval is [0, 1], and the actual distance value predicted by the model needs to be obtained by normalizing the inverse transformation according to the distribution range of the distance in the actual.
[0052] In this embodiment, when the algorithm reference point on the sea level is set (the position of the reference point depends on the needs of data processing), the azimuth angle parameter determines the accurate angle of the target ship relative to the reference point on the plane. The distance parameter gives the straight-line distance between the target ship and the reference point on the plane. When the two parameters are determined at the same time, the two-dimensional plane position of the ship can be uniquely marked in the polar coordinate system with the reference point as the origin, and the accurate positioning of the target is realized.
[0053] For example, to verify the effectiveness of the present application and to illustrate the best embodiment, the experimental data, data processing flow, model training details and performance verification will be described in detail below.
[0054] First, the experimental data used in the present application is derived from a passive sound source positioning measurement experiment conducted in the shallow water area of the East China Sea. The target is a ship sailing at a speed of less than 10 knots along a predetermined S-shaped trajectory. As shown in the following FIG. 3, the blue curve is the actual sailing trajectory of the ship, and the three red pentagonal marks in the figure are the distributed passive acoustic nodes a, b, and c.
[0055] The trajectory covers a horizontal range of about 2.5 km x 2.5 km. During the experiment, the dynamic range of the target ship from the array center was 300 meters to 1600 meters. The accurate trajectory information (longitude and latitude, speed, etc.) of the target ship was recorded in real time through the automatic identification system (AIS) of the ship. A distributed three-node passive acoustic network was used. That is, three omnidirectional hydrophones (with a sensitivity of -168 dB re 1 V / μPa in the range of 20 Hz-20 kHz) were placed in the form of beacons on the seabed. The hydrophone array continuously collected sound pressure signals at a sampling rate of 128 kHz, with a duration of 1 hour, and the original data was stored with 24-bit quantization. Each node device has a high-precision time synchronization function.
[0056] Second, the collected raw data was segmented and synchronized. The raw audio data continuously collected by each hydrophone node was cut into analysis frames of fixed duration of 0.5 seconds. The ship's geographic coordinates (latitude and longitude) recorded by the AIS system were time-synchronized with each 0.5-second acoustic data frame at the millisecond level through linear interpolation. The ship's geographic coordinates were converted to azimuth and distance in a local coordinate system with a selected reference point as the origin. The maximum and minimum distance values were recorded and used to scale the true distance to the interval [0, 1] for model training. From one hour of synchronized data, a sample data set consisting of 7,199 audio segments time-aligned with the ship's GPS position was generated. These samples were divided into a training set (70%) and a validation set (30%) (2,160). To focus on the frequency bands where ship noise predominates and reduce computational effort, the raw audio data (128 kHz) was downsampled to 2 kHz. A short-time Fourier transform (STFT) was performed on the downsampled three-channel audio data, with the STFT parameters set to a 32 millisecond window length and a 16 millisecond frame shift (step size). After obtaining the time-frequency spectrum of each node signal, the signal of a hydrophone node was selected as the reference signal to form two pairs of hydrophones. The interaural intensity difference feature and interaural phase difference feature were calculated. Finally, the combined feature maps of the two sets of interaural intensity difference features and interaural phase difference features (i.e., ILD+IPD feature maps) were normalized and input into the model.
[0057] To improve the model’s generalization ability and robustness to complex underwater environments, the following joint data augmentation strategy is adopted during the model training phase: 1. Time domain enhancement: Gaussian white noise is injected into the original sound pressure signal for 0.5 seconds, and the signal-to-noise ratio (SNR) is randomly selected between -5 dB and 10 dB.
[0058] 2. Time-frequency domain feature enhancement: The extracted ILD+IPD feature map (training set samples) is dynamically masked, and local masking is performed along the time axis and frequency axis respectively: During the training process, time domain noise addition and time-frequency domain feature masking are used in cascade to generate diverse training samples.
[0059] Third, the present invention proposes that the CNN-LSTM model be trained using the Adam optimizer, with an initial learning rate of 0.0001, a batch size of 32, and 3000 training epochs.
[0060] In the experiment, the mean absolute error (MAE) was used to measure the model's predictive ability. The smaller the MAE, the closer the model's predicted value is to the true value. On the validation set, the model's predicted angle MAE reached 4.87°, and the distance MAE reached 186.60 meters.
[0061] Figure 4 and Figure 5 The error distribution diagrams of the angle and distance predicted by the model and the true value are shown in Table 1. It can be seen that the error distribution is approximately normal and concentrated near zero, indicating that the model's prediction deviation is small and relatively stable. The mean and standard deviation are shown in Table 1: Table 1 Mean and standard deviation of the model predictions
[0062] Table 1 shows that the average error in distance prediction is 80.35 meters, with a standard deviation of 289.84 meters. This indicates that the model can effectively estimate the distance of the target sound source with minimal systematic bias. The average error in angle prediction is 0.63 degrees, with a standard deviation of 8.47 degrees. These data demonstrate the effectiveness of the proposed method in simultaneously estimating the azimuth and distance of a sound source.
[0063] Fourth, if Figure 6 and Figure 7 The training and validation loss curves shown in the figure below illustrate the model's learning progress during training. The horizontal axis represents the number of training generations, and the vertical axis represents the calculated loss function. It can be seen that as the number of training generations increases, the model's training set loss and validation set loss steadily decrease, eventually reaching a stable state close to 0. The validation set loss also reaches a stable state close to 5 and does not increase as the training set loss decreases. This demonstrates that within the limited number of generations of the proposed method, the model does not overfit and can learn sufficient effective features from the data, achieving parameter convergence and thus achieving stable prediction of the target sound source location.
[0064] The model's average single-sample inference time was evaluated by performing a full inference test on the validation set for 10 rounds. The calculated average inference time for a feature sample generated from a single 0.5-second audio clip was (0.40 ± 0.033) milliseconds.
[0065] The above are merely specific embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any modifications or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in the present invention should be included in the scope of protection of the present invention. The scope of protection of the present invention should be based on the scope of protection of the claims.
Claims
1. A three-node passive acoustic array ship positioning method based on a CNN-LSTM model, characterized in that, The method comprises the following steps: Step S1: Real-time acquisition of acoustic signals radiated by a target ship through a three-node passive acoustic network deployed underwater, and preprocessing of the acoustic signals of each node to obtain interaural intensity difference features and interaural phase difference features; Step S2: Constructing a CNN-LSTM model, inputting the combined feature map of the interaural intensity difference features and the interaural phase difference features into a spatial feature extraction module of the CNN-LSTM model for spatial feature extraction, and then outputting a spatial feature vector; The spatial feature extraction module is composed of two branches, each branch including a shallow convolutional block, a middle convolutional block and a deep convolutional block, the shallow convolutional block is used to preliminarily extract frequency and intensity / phase difference information of the combined feature map of the interaural intensity difference features and the interaural phase difference features, the middle convolutional block is used to further extract deep features, a batch normalization layer is added in the middle convolutional block to enhance the expression ability of the model, the deep convolutional block is used for feature association, channel dimension reduction and feature reorganization, and finally a Tanh activation function is used to output the extracted spatial feature vector; In order to avoid information loss caused by the pooling layer, the step length of the convolutional layer is set to control the down-sampling of the feature map in the entire spatial feature extraction module without using the pooling layer, and in order to speed up the convergence of the model and enhance the nonlinear expression ability of the model, an ELU activation function is used to introduce the non-saturation characteristics of the negative value interval before the output feature map of the entire convolutional branch, so that the output data maintains a certain gradient on the negative half axis, and the calculation formula of the ELU activation function is as follows: wherein, is the output of the activation function, is the input of the activation function, is the negative half-axis saturation coefficient; Step S3: Inputting the spatial feature vector into a time sequence feature learning module of the CNN-LSTM model for time sequence feature extraction, and then outputting a latent feature vector; Step S4: Inputting the latent feature vector into an output layer of the CNN-LSTM model for regression prediction, and outputting the azimuth angle and distance value of the sound source.
2. The three-node passive acoustic array ship positioning method based on the CNN-LSTM model according to claim 1, characterized in that, The preprocessing of the acoustic signals of each node in the step S1 specifically includes: cutting the original acoustic signals continuously collected by each node into analysis frames with a fixed time length and performing down-sampling; performing short-time Fourier transform on the down-sampled acoustic signals to obtain a time-frequency spectrum of the signals of each node, and then calculating an interaural intensity difference feature and an interaural phase difference feature; calculating the interaural intensity difference feature The formula is as follows: Computing interaural phase difference features The formula is as follows: wherein, and are STFT complex matrices of the signals received by the two sensors respectively represents the frequency, represents the STFT time frame index.
3. The three-node passive acoustic array ship positioning method based on the CNN-LSTM model according to claim 2 is characterized in that: In step S2, the extracted interaural intensity difference features and interaural phase difference features are respectively standardized to have zero mean and unit variance before being input into the CNN-LSTM model, and the specific formula is as follows: wherein, denotes the mean of a matrix of samples, denotes the standard deviation of a matrix, denotes the value in the row and column of a matrix, denotes the number of rows of a matrix, denotes the number of columns of a matrix, denotes the normalized matrix of samples.
4. The three-node passive acoustic array ship positioning method based on the CNN-LSTM model according to claim 3, characterized in that, Step S2 also includes dynamic mask enhancement of the combined feature map of the extracted interaural intensity difference features and interaural phase difference features, including random masking of part of the continuous area in the time domain and part of the frequency band in the frequency domain, and joint use of time domain noise addition and feature mask during model training.
5. The three-node passive acoustic array ship positioning method based on the CNN-LSTM model according to claim 4, characterized in that, In step S3, the spatial feature vector output by each branch of the spatial feature extraction module is converted into a one-dimensional feature vector through a deformation operation, and then considering the time sequence information of the two branches, the two one-dimensional feature vectors are spliced to form a time sequence feature vector containing two frames, and the time sequence feature vector is input into the subsequent time sequence feature learning module; In step S3, the spatial feature vector output by each branch of the spatial feature extraction module is converted into a one-dimensional feature vector through a deformation operation, and then considering the time sequence information of the two branches, the two one-dimensional feature vectors are spliced to form a time sequence feature vector containing two frames, and the time sequence feature vector is input into the subsequent time sequence feature learning module; The timing feature learning module is composed of two layers of LSTM networks, the hidden layer state output of the first layer of LSTM is output to the input layer of the second layer of LSTM, the LSTM network effectively learns and utilizes the long-range time dependence in the feature sequence through its unique gating mechanism, in order to prevent overfitting, L2 regularization constraint is performed on the two layers of LSTM networks, and the final latent feature vector is output.
6. The three-node passive acoustic array ship positioning method based on the CNN-LSTM model according to claim 5, characterized in that, The regression prediction through the output layer of the CNN-LSTM model in the step S4 outputs the azimuth angle and distance value of the sound source, and specifically includes that: the optimization target of the loss function of the CNN-LSTM model is to minimize the sum of the loss values between the predicted azimuth angle, the predicted distance and the true value, and the loss function is defined as follows: where, is the total loss function, is the loss function for predicting azimuth, is the loss function for predicting distance, denotes the number of samples, is the sequence number of the sample, and denote the real value and the predicted value of the azimuth, respectively, denotes the normalized real distance value, denotes the normalized distance value predicted by the model, which uses the minimum-maximum scaling for distance normalization, and the formula is as follows: wherein, and respectively represent the minimum and maximum values of the distances in the complete dataset.
Citation Information
Patent Citations
Deep learning-based binaural sound source positioning method in digital hearing aid
CN108122559A
Underwater sound radiation noise identification method
CN120071959A
Filter and method for informed spatial filtering using multiple instantaneous direction-of-arrival estimates
US20150286459A1